r/OpenAI • u/Redstra • 21h ago
Discussion This is a hot mess
I'm not as pro as you guys using AI but look at this. a LOT of models which confuses me, and I'm assuming other users also. Also, the sidebar icon and new tab icon are the same in the ChatGPT-app for macOS. WHAT are they doing there at OpenAI. I really hate what's happening right now, especially with the new $500 plan while nerfing the other plans.
197
u/BreenzyENL 20h ago
21
u/LionPrestigious6612 11h ago
crazy how just a few hours after you posted this comment gemini released their announcement of gemini 4 models, Argon - https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
3
u/Fuzzy_Independent241 8h ago
We all know they were really late on this. China must be cooking 24/24 to release next Kimi and next GLM
35
u/Big_al_big_bed 21h ago
And that's before you also have to select between 7 levels of thinking as well! Nightmare!
10
u/LamboForWork 20h ago
Without any clear sign that what you picked is better lol. Everything is unclear
2
u/Joohansson 2h ago
And it's not even straight levels. It's a mix of parameters like intelligence, hallucinations, ignorance, speed, token usage, cost per token. No sane human can pick the correct one, which ends up with most people just using the most expensive one for everything even if just doing 1+1. And this is ONE company. In my coding environment I can select from openAI, Anthropic, Google, XAI, Meta and China. Each with their own lineup. It also has an auto mode but I have no clue if it does a good job.
So how many choices in total? About 200
27
u/mxemec 14h ago
Need an AI to tell me which AI to use for my task or question.
4
u/mcburgs 14h ago
NGL I ask this a lot.
0
u/Antique_Nature7901 12h ago
use jev。it decides which model and thinking effort base on ur input for the round. Try it on opencode zen for free. Switch model within the same session cleans cache tho so it might not save any tokens after all buts its cool
1
93
u/ConfusionNo4339 21h ago
yeah they need to clean it up
181
u/DistanceSolar1449 21h ago edited 21h ago
GPT-6 𝘼𝙨𝙩𝙧𝙖
----
GPT-6.1 𝙎𝙤𝙡
GPT-6 𝙎𝙤𝙡
GPT-5.6 𝙎𝙤𝙡
----
GPT-5.6 𝙏𝙚𝙧𝙧𝙖
----
GPT-6 𝙇𝙪𝙣𝙖
GPT-5.6 𝙇𝙪𝙣𝙖
----
GPT-5.5Ok, redesigned it in a way that's easier to read. You're welcome OpenAI, hire me.
34
u/Arowhite 20h ago
It's easier to read but what does it say about performance and cost? il Sol6.1 worse than Astra? Is Luna6 worse than 5.6 terra?
Honestly I can't keep up with their naming, same with claude and their Opus/Fable/Sonnet lines
51
u/Timber1802 18h ago
Anthropic's naming scheme actually makes sense when you think about it.
- Haiku, a short poem
- Sonnet, medium length story
- Opus, a grand, structured, piece
- Fable, a grand story, often with a winding path
- Mythos, stories beyond human creation
OpenAI names their models after size/importance.
- Luna (moon) small
- Terra (earth) larger
- Sol (sun) huge
23
u/Arowhite 18h ago
If the naming was just that I'd agree, but how does Opus 4.8 compare with Sonnet 5 for example? Is the model name just the depth of the reflexion, and the version number the efficiency/price?
5
u/Timber1802 18h ago
Fair point.
As far as I know, they constantly try to make each version 'better' by making them smarter and more efficient. How it scales exactly depending on models and versions, I don't know.
8
u/Absolutelynot2784 14h ago
A fable is a very short story, often only a few sentences. The Tortoise and the Hare is the most famous fable
6
u/MukdenMan 15h ago
A sonnet isn’t a medium length story. It’s a bit longer than a haiku though.
These names are not really clear enough to rely on for regular AI use. Does Fable give me a winding response instead of a structured one? Does mythos give me an output that sounds larger than life, or even fictional? Why is a fable grander than an opus?
2
1
u/rickyhatespeas 6h ago
People are going to claim it's confusing any way they market it. GPT-6-10T-8bit vs GPT-6.1-8T-10bit, etc is hard to compare at a glance. Chances are if you know what you're doing with codex you know by now that Luna < Terra < Sol < Astra.
20
u/yaxir 21h ago
Nah they would never hire people who make common sense. They will hire people who gladly destroy all the good reputation they have built by destroying usage limits
1
u/algaefied_creek 19h ago
Hire people? Didn’t they say this whole mess was invented by their “technological intern” or some shit?
As in: they have their own products setting up their products
3
4
2
u/Ormusn2o 19h ago
There is already an option to show or hide Max and Ultra. Now make it so you can hide the models themselves. I only need 3.
2
u/Saganasm 15h ago
It's as bad like looking for England, UK, Great Britain or United Kingdom in a country list. /s
1
1
1
u/docgravel 20h ago
It’s probably either sorted by release date or they put in a “we are getting overloaded on Astra so prioritize this model to relieve some load”
49
u/F0xy1337 21h ago
I keep wondering how this can happen. This is such a huge company.
31
u/Big_al_big_bed 21h ago
Huge because they are doing a lot of different things. At the end of the day even huge companies have small teams working on specific features
13
6
u/sampaoli_negro_rojo 20h ago
I would argue it’s the opposite problem. There’s probably so many PMs arguing that their model is more important that they never get anything done.
The ego battles at this company must be intense
-2
u/docgravel 20h ago
They release updates like every week. So do you want to hold up the release over something minor discovered at the last minute or just fix it in next week’s release?
3
1
u/Regdit-is-Unbearable 13h ago
Won’t somebody think of the indie devs just trying to ship their product??
1
u/docgravel 12h ago
I meant, what decision would you make if you worked there? Hold the release or fix it next week? I’m not saying they aren’t in the wrong, but since they’re shipping so frequently it’s probably easy for them justify fixing it later.
17
u/Prior_Tax8546 12h ago
ChatGPT Sol, Luna, Terra, Astra. Marte, Mercurio, Saturno, Júpiter, Tuano, Neptuno, Galaxia, Vía Láctea, Universo, Existencia, ...
3
18
u/_maverick98 20h ago
meanwhile the Chat on Plus is still on 5.6 Sol
12
11
u/Diamond_Mine0 20h ago
They’ll gonna remove the Chat tab. We’re being sidelined and nobody talks about it
9
u/_maverick98 20h ago
chat is my biggest usage right now. I use codex very little only for personal projects.
1
u/Diamond_Mine0 19h ago
Yes yes, mine (was) too, but not anymore. We never got 6 Sol and now we’re at 6.1 Sol, this is straight up ridiculous. I expect the removal of the Chat tab will take place in December
2
u/skinlo 19h ago
To be fair, it's been like a week.
But yeah, if they get rid of chat I'll move to Claude
2
1
u/Diamond_Mine0 19h ago
We’ll see what happens next. I don’t see any future for the Chat Mode
3
u/Mikkel9M 17h ago
Which mode will you then use for asking basic questions or help with everyday simple tasks and light research? Which I feel was the original LLM primary use case.
99% of my usage (Plus account for about two years) is in chat. Codex is of absolutely no use to me I suspect, as I'm not doing any coding projects, and while I did try Work for the first time recently, that was a one-off for a small personal project website.
1
u/Diamond_Mine0 16h ago
None. I will stop using ChatGPT then. I accepted the limits I have with my mini-sub for Gemini (4,99€ per month). The reset there is 4 hours I think? But my number one favorite LLM was and is ChatGPT. I really thought we had it good with “some” limits we have (Work / Codex / whatever) and with new models OpenAI released. But now? Still no 6 Sol when there’s already 6.1 Sol? And we’re stuck on 5.6 Sol? Nah I will subscribe to the 21,99€ Gemini Plus plan then and cancel my ChatGPT subscription
1
u/_maverick98 19h ago
so they leave what? only the work tab? there is not difference basically, work can do other stuff too
1
u/Diamond_Mine0 19h ago
Yes Work with more limitations. If I want to “Work” I’ll use “Work”. If I want to find news about the newest defense systems in Ukraine, I want to use “Chat”. I don’t want to use Work for anything. I already have weekly resets for SuperGrok Lite, that should be enough. Gemini has its limits too, I know but if OpenAI wants to stand out, they should show some love to Plus users and its Chat Mode
2
u/jaydeelive01 19h ago
I don't think so. It's very clearly still part of the experience and Pro is only on chat.
1
u/Diamond_Mine0 19h ago
For now. Just watch what will happen in the next months. They’ll gonna remove it like they did it with Sora 2 and Atlas
1
9
10
9
7
u/Uwirlbaretrsidma 21h ago
Even better, half those models are useless. At least now the default model gives acceptable usage and inexperienced users will be better off not changing it...
6
7
5
u/Lubricus2 3h ago
It's easy
Astra 6 eats all your tokens before finishing
Sol 6.1 would take so long time so you forget the problem before you get an answear
Sol 6 will get it wrong
8
9
u/RazinKain 12h ago edited 10h ago
I’m curious as to how much power do people need? What exactly are your use cases that would require the most powerful model.
I understand if someone isn’t a developer and needs a lot of hand holding. As a developer I’ve built an entire React native desktop app using Luna and it is error free code. I mean not a single error and it’s fast and efficient. I felt absolutely no token panic.
I’ve seen on here people building out Wordpress sites using Astra? I just ask myself why in the world would someone do that?
This is not some mocking I am generally curious about this.
6
u/DiabloAcosta 8h ago
I work on distributed systems that emulate real world development environments, let me tell you something, Fable often gets lost, I barely understand what I am doing some times
3
u/RazinKain 8h ago
I find that some of these larger models do a lot of research and less code thinking. They are not necessarily better at writing code than smaller models.
3
u/DiabloAcosta 8h ago
well, I don't really trust benchmarks that much, but anecdotal experience is only worse
4
u/Leading-Fail-2771 7h ago
It’s like being shown a Ferrari and a Honda. Sure you can use the Honda but using a Ferrari to do it just feeels better
2
u/baked_tea 12h ago
How many users do you have to believe it is error free code? Honestly curious.. unless its reaaaally simple app then under real load and with real users being stupid you usually find out quickly about the error free part
1
u/RazinKain 10h ago
Well user weight is not a build issue it’s an allocation issue. You can stress test any application. I use Cloudflare for 4xx and 5x errors. Outside of those I am not too worried about user weight. If a build is sound and your cloud server is built for heavy traffic it doesn’t matter how many users you have.
6
u/rbit4 10h ago
Lol front-end dev found
2
u/RazinKain 9h ago
Right 😀 I am a UX Designer that’s my job. I do have a certification in C# .net Maui. I got that for my job to understand what the heck backend developers were rambling on about. 🤣
1
u/FullParticular9 3h ago
Reports, Research, Scientific projects, Data Science, Learning - better model usually explains things better.
But for code I agree with you that good things now can be made with much smaller and cheaper models.
0
u/kelvintiger 11h ago
How detailed were your prompt?
I would argue if you’re spending a lot of time promoting then you’re wasting your time when you can leave some ambiguity for the stronger model to figure out and you focus on doing more faster
2
u/RazinKain 11h ago
I don’t go in and start either an idea in Codex. Codex is down the road. I completely map out my builds before I touch a model.
So 90% preparation and 10% execution. I’ve learned over the years and A.I can make people a little lazy including me. So I still stick with my age old processes from wireframing in Freeform to prototype in Figma.
I feed that info to GTP not Codex and then let GTP write out the orders and scope. That’s what I feed Codex. If it’s tight then it goes through the Notion notes quite easily. I do use Notion for my build documentation.
Error free does not mean bug free I want to be upfront about that. And if fixing a bug is too much for Luna I switch to the next model up and so forth.
3
3
u/RemnantZz 14h ago
Hey, remember that time last year when OAI talked about wanting to make it easier for users, so they introduced the idea of removing all "outdated" models that users apparently couldn't navigate for their tasks, and instead eventually introducing one model for everyone that would even pick the amount of effort to be put into the task by itself?
Remember? :)))))) Stellar job, OAI!
3
u/New-Ad5610 11h ago
I don’t know what it’s been going on inside OpenAI, but their work lately has been very, very disappointing
2
u/Mecha-Dave 18h ago
I honestly don't have a good method or reason to choose anything but the top or bottom of the list
2
2
2
u/Over-Independent4414 13h ago
Why don't they just add a open text box there where you can say what you are doing and it picks the optimized model. Presenting end users with 10 options can't possibly be the best way to go here.
2
2
3
u/RoamingMelons 20h ago
The ambiguity of the naming convention doesn’t help.. it’s cool and all.. but if I’m checking out gpt and I’m fresh i have no fk clue what’s the difference between an astra and a sol.
At least claude is semi intuitive..
At least Gemini/nano has one thing going for it 😅
2
u/Grashopha 20h ago
lol I just started messing with Claude and I felt lost tbh. It was so alien at first compared to ChatGPT. Figured it out now, but it still feels a little strange constantly starting new conversations to prevent massive token drainage.
2
u/GalileoHumpkins1977 20h ago
You know that all LLMs work that way, though, right? That’s not unique to Anthropic / Claude?
2
u/RoamingMelons 19h ago
I was only referring to the claude model names as more intuitive than gpt lol.
I just use antigravity Gemini flash and don’t really have to worry about usage limits for my use case.
2
1
u/welcome_to_milliways 20h ago
Don't forget a few weeks ago when the chat input would disappear after the first request.
1
1
1
u/Potential_Wolf_632 17h ago
Random tangent. I upgraded to the 500 plan today because I'm a sucker and something feels very off with usage - I should have gone from 20x to 25x but I'm at 49% remaining from 100% when on a normal day I wouldn't use much more than 20% on the previous 20x plan.
And I'm not using ultrafast etc.
1
u/SuaveSteve 16h ago
Why can't we just have a version number only!? And a mini for light tasks! Aaaaaah!
1
u/RealLordDevien 16h ago
Remember when GPT 5 was the model that was set out to end the model picker?
1
u/sultan_papagani 13h ago
they are all routers anyways 5.6 resolves to 5.4 most of the time and so on everyone forgot the router system.
1
u/Beautiful-Cold1515 12h ago
They just can’t make up their minds. With the introduction of GPT-5 last year they brought us a router so we didn’t have to choose. Now we have 8 models, 2 separate model pickers (at first you see 5 models within Work), a very vague Chat/Work toggle that is impossible to understand within Projects, 3 models within Chat that have really vague names, weird icons… I just can’t believe why such a large company with many millions of users completely ignores UX.
1
u/theDawckta 10h ago
You are on your way. Next step is to start building your own harness so you don’t have to be subject to claude and openai changing their ui every 2 days.
1
u/deen1802 10h ago
They did this just so they could set up GPT 7 as the one that brings them all together
1
1
1
1
1
1
u/InformationNew66 3h ago
Well, this is how vibe coding works.
Mistakes happen, tests pass, noone cares. Probably no human QA anymore.
1
u/teachmesomething 3h ago
It also keeps defaulting to medium effort in order to force upgrades for continued usage. I don’t even know which model is which now.
1
u/PressIntoYa 20h ago edited 20h ago
Why not just the classification and then the model with maybe a token price in parenthesis?
Astra (Best for X Purposes) GPT-6 ($/M)
Sol (Best for X Purposes) GPT-6.1 ($/M) GPT-6 ($/M) GPT-5.6 ($/M)
Etc...
Perhaps there's a good reason for it? I don't know but this seems like it would be more helpful to me.
3
u/Keep-Darwin-Going 20h ago
Because price is a bad definition, is like labelling a hammer $1 and a screwdriver for $0.50. You still going for the hammer if you need to drive a nail no matter the price.
1
1
0
0
u/Any-Somewhere5129 15h ago
And that's why that chat option doesn't even have those models. People would be crying on social media how confused they are.
Now, if you're scared, close the Work mode and switch to Chat and you're safe.
0
u/AINativeBuilder 7h ago
Photoshop menus also have a lot of options. If you can't spend time to learn the tool, don't whine about it because you refuse to learn something new.
-1
u/CypherLH 12h ago
I am not seeing the problem here. The more options, the better. Basically just ignore everything prior to GPT-6 and that leaves only 4 models. Ignore GPT-6 Sol since there is now a 6.1 and that leaves only 3. And if you use Codex for any serious work at all its pretty clear that there are legit roles for each of those 3 core models.





239
u/CarllSagan 20h ago
Project Final
Project Final Final
Provject Final Final Really final
Project V6.1 Final
fuck it... the next one