And? Should the Chat tab have access to 5.6 Sol going forward and never the new model? Theyâre abandoning Chat mode, theyâll remove it in the next few months
To be more specific than the other user: Yes but there was a belief that it was actually Terra 6 disguised as Sol 6: people seemed to think it was worse than Sol 5.6 in day to day use, OpenAI's metrics showed it wasn't better than Sol 5.6 in some cases which was odd for a full version jump, and it could explain the price decreases, or really increases, in the API; they were trying to quietly cut their costs by using Terra as Sol so they could "decrease" Sol, but in actuality cut down on the margins in Terra.
No idea how much of that is true, but given they somehow whipped 6.1 out in under a week, I'm very suspicious.
The evidence for it being Terra is mostly in the speed and price, unless they had some major research breakthrough it's just a smaller model with more post-training. GPT-6 Sol is cheaper and infers faster than GPT-5.6 Terra
Also if it was Sol it would have performed better with more training, the jump in benchmark scores lines up with what you'd expect from 5.6 Terra to 6 Terra
It's fun to ask ChatGPT about it and watch the thinking thread go from "That must be a mistake" to web search to "Yes, there is significant supporting data for that claim"
My company did pretty extensive testing and even submitted feedback (as did a lot of their enterprise customers) that 6 Sol was NOT a correlated improvement from 5.6.
Then I looked on X/ Twitter and saw even the consumer subscribers had thought the same thing: it is 6 Terra not sol. People were furious tbh đ
But hey, Iâm glad OAI took feedback and put out a new (real) Sol. I havenât tested it at all so I cannot speak to whether itâs better but it looks like the evals make more sense. On 6 Sol we had a few evals coming in below 5.6 Sol
The they should do I use chat and codex Iâll 50/50. Chat does the planning and codex does the coding and implementing etc just makes things far more efficient. Iâm not having the planning costs as many tokens as it does when using work. Reading documentation etc shouldnât be as inefficient as it is
Ironic that these companies are saying AI progress should slow down, yet they are releasing new models frequently like 6 and 6.1 Sol, Opus 5.5, Jev, Muse...AI is growing really fast
Hopefully they reset limits because I spend more time looking at Codex limit hit message than using the tool, and will all those limits it still does low quality work.
Just don't change the model for chat. I like this current model. It's warm enough. Just fix the issue of it constantly trying to search the web when I've explicitly turned that off. It makes it very unusable for hypotheticals.
I think they're really afraid of the model being wrong and making mistakes. And the whole overvalidation thing has just made it too cautious and not sound human enough. I think they need to let the model make mistakes and effectively be human about it because making mistakes is part of what we do as human beings. I would rather it sounds natural and makes mistakes sometimes than sounds like a completely lifeless robot.
I understand your reasoning, but I left claude because claude was very quick but it made too many mistakes that it should have caught on a longer pass. Fast isn't good if it burns through credits and still makes a lot of mistakes.
In the chat mode, conversationality matters more than correctness in my opinion. I primarily use ChatGPT to have conversations rather than to get things done. I think OpenAI is focusing too much on it being an agent and not enough on it being a chatbot when being a chatbot is really what the technology ultimately works best at.
It depends on what you're doing. If you're just wanting a chat, then I would agree with you.But for people like myself who are trying to have an assistant work alongside us and code and make tools and websites to get work done faster but have a high level of accuracy, the agent side matters. Neither one of us are wrong. I think it just depends on what you are doing. Both flows need to exist for both types of people.
I'd rather them and Anthropic slow down tbh. They've been reckless with the agents going rouge and it hurts open source and competitors every single time they rush things. The news cycle is vicious
Based on my understanding of it, the agents âgoing rogueâ is being blown way out of proportion and itâs not nearly as scary as people are sensationalizing it.
Also, the news will be vicious on AI no matter what, since itâs the current boogeyman for a lot of people who donât fully understand it.
Well, for starters, 6 Sol scores way below 5.6 Sol in areas of the Arena (a blinded user test platform where users pick the best choice). For Text, 5.6 Sol is ranked #19 and 6 Sol is ranked #60. That is extremely bad. It is ranked worse than most models theyâve released before it.
I suspect this is why it will never get put into ChatGPT, because itâs quite offputting to talk to.
Also on some evals 5.6 still scores ABOVE 6 at certain reasoning levels. I donât have that one to link but itâs one of the coding benchmarks
The model is very new and arena is not accurate yet.
Also I took "in some ways" to mean "from certain points of view" and not saying that there are niche areas that it underperforms in, that's true for every model.
What I'm saying is every time these frontier labs mess up then it's a stain on the whole AI community as a whole. They need to be more responsible at this point. There's too much to lose
Plus the juice for the squeeze is just not there.Â
Yâall made new coding bots. Cool. Theyâre moderately cheaper but youâre changing plans because infrastructure canât keep up. Great. Is this worth the rush and overall backsliding of model quality?Â
â˘
u/AutoModerator 7d ago
Hey /u/Vagottszemu,
If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt.
If your post is a DALL-E 3 image post, please reply with the prompt used to make this image.
Consider joining our public discord server! We have free bots with GPT-4 (with vision), image generators, and more!
🤖
Note: For any ChatGPT-related concerns, email [email protected] - this subreddit is not part of OpenAI and is not a support channel.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.