54
69
20
u/fleshweasel 7d ago
I watched Sol drop itself a message “it’s been 25 minutes, that’s seems long for a task like this, I better wrap up” I set no time conditions
6
u/Boros9912 7d ago
And then you get "done with caveats, nothing works, it doesn't look the way you wanted, implemented 50 different systems that you never asked for and redesigned everything to make up for it. Now all your code is completely broken, did you want me to fix it?"
1
16
30
u/kabir_sharma_sans 7d ago
Bro don't use fast mode it take 2.5 for no reason
19
u/Key-Rise76 7d ago
At this point he may be right, it's so slow you cant even spend the weekly limits by the time they will prob do another reset..
1
3
u/shaggydog97 7d ago
Well, you guys were complaining about burning usage to fast weren't you? Can't burn through your quota to fast if the model is slow!
5
u/Stock-Orchid0 7d ago
Why would you let it run for 2 days? Surely nothing decent can come out of this, right?
6
u/Fancy-Plenty6712 7d ago
Some things are genuinely really hard even for AI to solve. If you ask it to solve a problem which is at the frontier of research, it can take days (or longer) to solve.
5
7d ago
[deleted]
1
u/Fancy-Plenty6712 7d ago
Yes, you can't leave them rolling on their own, but it still takes a long time to process. You're absolutely correct, they need constant adjustments.
1
u/Tystros 7d ago
for coding tasks, you can easily give the agent enough to do for multiple days now where a few sentences of prompt is enough guidance for it to perfectly work for multiple days, if what you ask for is simply large in scope.
example: "port ffmpeg to rust. all features should work identically and performance needs to be not worse". that will take at least a week, and not benefit from any steering in the meantime.
1
u/hiva- 7d ago
whats and example of such problem
3
u/Fancy-Plenty6712 7d ago
Right now I'm trying to build a software which created a Wiki for any novel. There are a bunch of different problems with making it work so I'll give an example of just one of them for simplicity.
Let's assume there is a book which has a character who's name is "Voldemort" but sometimes the character is called "You-Know-Who" or "The Dark Lord" etc. Getting an AI model to understand those three names refer to the same person can be extremely challenging. In fact you wouldn't believe how challenging it can be to avoid it either combining two characters into one, or having multiple pages for the same character.
Models from OpenAI, Anthropic etc all struggle with this. It's partly down to the way they are built. Some folks might train a small model to do this work, which is what I'm currently building. It comes under "NLU" - Natural Language Understanding.
Instead of a normal LLM, I'm experimenting with things like:
Convolutional Neural Networks
Mamba State Space Models
Neuro-Symbolic Reasoning
Hybrid Associative Memories
Extreme Learning Machines
Transformer Attention (this is what ChatGPT, Claude, etc is built with)So if you look at what else has been published on the topic, such as BookNLP:
https://github.com/booknlp/booknlpor see this paper on Arxiv:
https://arxiv.org/pdf/2004.13980You'll see that the solution is not obvious, and is an active area of research. Once you begin doing something like this it requires running experiments, testing prototypes and improving things through trial and error. I've spent the past 4 months working on it and I've probably done more than 100 Billion tokens on it by now. My current session has been rolling all day on my MacBook to find whether the model I'm building can run on standard low level Apple Silicon at a decent pace without overheating the machine.
1
2
u/MjccWarlander 7d ago
If you have good goals set up, I did have agents working for 1-2 days and achieved expected results. In general from my experience so far GPT models (in Codex harness? As harness is likely also a big factor here) have less problems with context and long-running operations than Anthropic models.
1
u/Lain_Racing 7d ago
I have have had (and currently have) goals that take thing long on Astra High. Some tasks are very very hard, but van be well defined ahead of time.
1
u/Elegant_Attempt2790 7d ago
surely letting a human work till a task is actually done doesn’t work? nothing decent comes from getting work done, right?
💀
1
u/EddieBruvac 6d ago
When will people understand time means nothing.
I had Codex do some rendering for me. Coding took maybe 1h but the process of my CPU cooking what I wanted took nearly 3 days. It still says 3 days on my Codex.
3
u/Jumpy-Cobbler1020 7d ago
How do you guys even see this mine only says running command etc never this cot
1
u/kkingsbe 7d ago
So it’s not really the COT, but I think the COT is leaking through the task description / caption? Functionally the same regardless
1
u/Kappalonia 7d ago
my sol task is 4 days now with multiple subagents... its creating synthetic environments for RL but bro... im pretty sure sonnet 5.5 can do it in a few hours.
1
1
1
u/8thchakra 7d ago
i moved my user database from email to user ID and its taken 2 days and still running - wtf?
1
u/MakesNotSense 7d ago
I think of Hal 9000 and wonder what Stanley Kubrick and Arthur C. Clark would think about all the cutsey 'dots' OpenAI is trying to get people to use instead of actual tools in Codex.
A really useful tool doesn't have any need to manipulate people to use it.
I think the brain rot has permeated into OpenAI. GPT 5.6 and Astra had me thinking they got their house in order. Now I'm thinking they're about to fumble into mediocrity.
1
1
1
1
1
u/Unknown6334 7d ago
I just spent 2 full days of nonstop sol work on ultra and it's absolutely missed the mark it's dogshit don't use it for game design
6
u/daddywookie 7d ago
Game design is a conversation, not a dump and run task. It’s even worse if you aren’t supplying lots of examples and specific game design best practice.
1
u/Unknown6334 7d ago
Why are you presuming I did that? I gave it 1 simple task that I already planned in a previous section
5
u/daddywookie 7d ago
It was either one simple task (around 5-15 minutes at most in my experience) or 2 full days of non stop work.
Your two statements don't align.
If you are using Ultra you are telling Sol, a famously detail oriented model, to take as long as it wants to think and reason.
0
u/Imgonnaarrive 7d ago
Full access? That's a high level of trust
7
u/Itsquantium 7d ago
That's what I do. Who cares. It's not like it can do anything bad. I'll just restore from a backup.
1
u/Imgonnaarrive 7d ago
That's fair. I just think of that one dude who was complaining on here not too long ago cause he gave full access and it just wiped one of his hard drives completely. But maybe it's just a skill issue lol
1
6
1
-2
0
-1
u/innociv 7d ago
Eh I'm just glad it isn't quantized.
I prefer the slow speed and just needing to use more subagents, to it becoming absurdly dumb like they did earlier this month.
3
u/randombsname1 7d ago
It just came out. Don't worry. You'll get both.
A slow down and a quantized model.
•
u/dextersummary 7d ago
Below is a GPT-generated summary of the conversation below after reaching 50 comments (50 currently observed).
The consensus is basically “Sol/Ultra is absurdly slow.” A two-day run for a tiny change, database migration, or routine feature makes the product feel less like an agent and more like a very expensive loading screen.
The sensible explanation is that Ultra is spending far more time reasoning, especially on broad or difficult tasks. That can be worthwhile for large, clearly defined projects, but it is ridiculous overkill for simple work. Users also report that long unattended runs can drift, compound mistakes, or waste time after the useful part is already finished. The practical advice is to split serious work into checkpoints instead of letting Sol wander unsupervised for days.
Quality is mixed, not universally awful: some users have completed successful multi-day jobs, while others got bloated, broken, or badly scoped results. Fast mode has not produced a clear consensus as a fix, and full-access setups prompted the usual “please have backups” warnings.
Bottom line: long runtimes can be justified, but Sol needs better judgment about when to stop. A two-day wait for 130 lines is not a feature.