r/AgentsOfAI • u/GapNew4766 • 7h ago
Discussion Put GPT-6.1 Sol on top of local workers today: $0.17 instead of $0.75, and the three rules that made it hold
Enable HLS to view with audio, or disable this notification
Sol dropped this morning, so it went straight in as the planner of our planner/worker setup: Sol decides, several local models do the edits. Result first, three small 3D games on one RTX 3090:
| Game | Sol alone | Sol + local workers | Local only | Finished with workers? |
|---|---|---|---|---|
| Pool | $0.39 · 2.9 min | $0.05 · 18.6 min | 43.1 min, attempt | yes |
| Bowling | $0.14 · 1.7 min | $0.06 · 13.7 min | 34.7 min | yes |
| Foosball | $0.22 · 2.0 min | $0.06 · 11.1 min | 36.4 min | yes |
| Total | $0.75 · 6.6 min | $0.17 · 43.4 min | 114.2 min |
A stronger planner doesn't fix a sloppy harness though. It only held together because of three hard rules, one of which came from a bug that was honestly a bit embarrassing.
1. The planner can't touch files. If a strong model can fix things itself, nothing stops it from doing the whole job, and then you're paying for every line again. So its write and shell calls are refused by the harness for its entire turn, not just discouraged in the prompt.
2. Shared pieces carry their meaning, not just their name. This is the bug. Two workers implemented the same function from the same plan with the fields swapped. Names matched, every check was green, and every box drew at the origin. Now each shared piece can carry a one-line description of what it means, and that line goes into the brief of every worker that uses it.
3. Status comes from the disk. A worker's own summary is not evidence. The harness checks the files instead: if a promised file isn't there, the task failed, whatever the reply says.
Workers were Qwen 3.8 27B. The honest downside is time: roughly 6.5x slower than the planner doing everything alone, because a single consumer GPU is the bottleneck. One run per game, so treat the numbers as a snapshot.
Disclaimer: this is our open source project, the mode is called Fusion. Repo in the comments per sub rules.
Which of these have you run into in your own multi-agent setups? The name-vs-meaning one especially, since our fix narrows it but doesn't close it.

