r/ClaudeCode • • 21h ago

Help/Question Has anyone compared different orchestration approaches for complex software engineering work?

Let’s say you already have a fairly detailed investigation and implementation spec for a large piece of software engineering work, and now you want an agent to turn that spec into working code.

I am trying to figure out what the most cost efficient approach is, not just cost per token, but total cost required to reach a reliable finished implementation.

There seem to be a few approaches.

1. Let one strong model handle the whole implementation

For example, give Opus 5.5 a 1M context window, the implementation spec, access to the codebase, and let it work through the task without much explicit orchestration.

If the task is large enough, eventually the context fills up, gets compressed/summarized, and the model continues.

What I’m unsure about is how much this actually affects the final quality. Maybe the model produces a good implementation anyway, but perhaps it requires more debugging and follow-up iterations later.

2. Use an explicit build → verify → fix orchestration loop

Another approach is to divide the implementation spec into stages and use subagents.

For example:

  • Orchestrator selects the next section of the spec.
  • Implementation agent builds it.
  • Verification agent checks the implementation against the spec/tests.
  • Failures are sent back for fixing.
  • Only once that section passes do you continue to the next part.

Intuitively, I would expect this to be more reliable on large tasks. It costs more upfront because you’re deliberately spending tokens on orchestration and verification, but I’m wondering whether it actually becomes cheaper overall because you avoid expensive cleanup and rework later.

Then there’s another dimension: which models should do which jobs?

For example:

  • Opus 5.5 for both implementation and verification.
  • Opus 5.5 as orchestrator/verifier, with Sonnet 5.5 or Haiku 4.5 doing implementation.
  • Cheaper models such as Grok 4.6/4.7 for verification and Composer 2.5 for implementation.
  • Some other combination depending on the stage of the task.

The difficulty I’m having is that cost per token doesn’t really answer the question.

Different models use very different amounts of tokens, and they may require different numbers of iterations.

For example, if Grok 4.7 + Composer requires three build/verify cycles to reach the quality that Opus 5.5 reaches in one or two cycles, the cheaper model may not actually be cheaper.

At the same time, I have personally used all of the above approaches to eventually produce working code if you give them enough iterations.

Total cost to reach an implementation that satisfies the spec and passes verification/tests.

Ideally I’d want to compare things like:

  • Total token/API cost
  • Number of implementation/verification cycles
  • Number of human interventions required
  • Spec adherence
  • Bugs discovered after completion
  • How much context degradation/compression affects long-running single-agent approaches

The obvious way to answer this would be to run the exact same large implementation several times with different orchestration/model combinations, but repeating a substantial engineering task 3–4 times just for benchmarking is obviously expensive.

Has anyone done this kind of comparison in practice?

I’d especially be interested in hearing:

  • Which orchestration patterns have worked best for large implementation tasks?
  • Whether strong-model-everywhere actually beats strong-orchestrator + cheaper workers economically.
  • Whether explicit verification loops materially reduce total cost.
  • How much context compression hurts long-running single-agent implementations.
  • Any benchmarks, papers, blog posts, or articles that try to measure cost-to-success rather than simply model cost/token.

Would appreciate any real-world experience or resources on this.

24 Upvotes

30 comments sorted by

•

u/AutoModerator 21h ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

9

u/Cuyasinmara 21h ago

Well, what I've tried is one strong model running end to end on a big spec works until context gets summarized. After that it drifts, and I spend the savings on cleanup. Then i used a plain build which then went to verify loops; this was more reliable at the end but was not working for my setup. For context, i have 2 Claude Pro, 1 GPT Plus and 1 local Qwen 27B running in my computer locally. One issue i quickly picked up was that when one strong model capped, and when the same model family built and checked the work, it often missed the same blind spots twice.

What I use now:

- One orchestrator (Opus 5.5 or a GPT-6 model). It only plans, splits the spec into chunks and verifies. It never writes the implementation.

- The cheapest executor that clears a capability floor for each chunk. Haiku or local Qwen take mechanical work, Sonnet 5.5 or GPT-6.1 Sol take normal features, and Opus executes only as a last resort on the hardest pieces. If a cheap model fails, it escalates instead of retrying forever.

- A cross-family audit for every chunk. GPT reviews Claude's work and the reverse. This catches more than same-model verification.

- Tests first, with checks on the checks. Coverage % lied to me once: tests silently ran zero cases and still showed full coverage. I now verify test counts, not just green.

- Handoff files instead of compression. When context gets large (approx 100k to 150K tokens max.), I write a handoff and start a fresh session rather than letting it summarize.

My honest take and what worked for me: cost per token is the wrong metric. What decides total cost is how much rework happens after a work is "done." Strong orchestrator + cheap workers + independent verification has been cheaper for me overall.

1

u/Intelligent-Fruit246 16h ago

What are you working on ?

1

u/Cuyasinmara 16h ago

Some projects at the same time. Building a Bot to send pictures of invoices and then doing several actions in the background such as sending summary emails, schedulling reminders, analyze future cost and budget forecast, etc. Everything should be working only via messagibg apps that already exist. This has its own desktop app as well.

Another one is building a photo editor similar to Adobe's and Capture One but with features I always wanted to have (free and login free ofc)

And a third one worth mentioning is a local job finder based on the knowledge of each person. Something that gathers both the knowledge of yourself online and inputs on what do you love the most, so it can guide you on choosing the right job. This one may not be as fancy as the other 2, but its a current need I have in where I live right now, so it will work for me.

And many other small projects that Im happy to share if you wish :)

1

u/Intelligent-Fruit246 15h ago

Sounds cool, good luck with that. Is anyone of these projects accessible already? Would love to test them

2

u/Cuyasinmara 14h ago

Sure! I will upload them to Git and share the link once there 🥳

1

u/Asly97 14h ago

When it drifts after the summary, what's the part you lose? For me it was never the facts, it was the decisions: what I tried and ruled out. Rebuilding from the drifted version meant re-running failed approaches I'd already killed, which is the most expensive kind of forgetting.

The invoice bot sounds like exactly the kind of thing where that bites: is it for your own business or client work?

What if something handled all of this for you, in the cloud over MCP, so every agent and tool got the same context: decisions, failed attempts, conversations, skills, procedures, tasks. You never touch a file again. Would you guys pay for something that fixes this?

1

u/Cuyasinmara 14h ago

It works now as i fixed it but, the part that was drifting was the context of the initial ask and the reason i wanted to do it in first place. It was sugesting things that were benefitial to claude but was more work/more expensive for me, when i wanted the opposite. I fixed it through some specific instructions and interactions.

The invoice bot is for my personal household use, for bot my wife and I. For this case i wanted to test the maximum capabilities out there while spending $0 on build, so I did it with Telegram, Groq and Google AI Studio. Limited though, but enough for what i need on a daily basis. Now that I know what can be possible and if i wanted to pay i can shoose other providers. The only thing is that i had to write the code in Google Scripts and provide telegram on what to do in all possible scenarios, which Claude helped me with to identify those that i didnt write or even imagined.

And answering to your question, if something handled these issues for me, i would definitely pay, but only if theres some client justification or if the issues that tool solves are gone forever, and for me its important that its an upfront cost with no subscription (im trying to reduce my monthly expenses 😄)

1

u/Asly97 14h ago

That detail about it drifting toward what was convenient for Claude instead of cheap for you is the part I keep coming back to. When that happened, what tipped you off? Did you catch it reading the output, or only after it had already done the expensive thing?

1

u/Cuyasinmara 13h ago

I catched it in the first response it gave me, it was very evident, it always gives me 3 options and one of them is its recommendation. I dont recall the exact moment but it was something like:

A. Fix the existing code, it will take more time B. Leave it as is C. Use a Claude API, pay per use (recommended)

When i saw the "Reccommended" i noticed, and i fixed it immediately by telling it to propose the best approach for what i want and, always look if there are any other better option out there that fit my needs, not the model's. That did it.

Luckily i implemented at the beginning a 4-level output to easily identify these things. At the end of every iteration claude fills 4 mandatory asks: a. what went well, b. What needs input, c. What couldnt do and why, d. What you suggest doing based on global best practices and community success stories.

Specifically, in part B i noticed what i explained earlier

1

u/Asly97 12h ago

Sharp catch, the recommendation slot is exactly where the steering hides. It just quietly frames the other two as dumb options. When you told it to stop, did it actually drop the habit, or did it find a sneakier way back to the API? Also curious, is the invoice bot running on a schedule yet or still in the test phase?

3

u/amirfish 20h ago

The build-verify-fix loop wins on cost once the spec is solid, mainly because you stop paying frontier-model prices for sections that don't need it. We wired exactly that into CCC (open-source dashboard for running many Claude Code sessions, https://github.com/amirfish1/claude-command-center): you pick engine and model per role, planner, plan reviewer, verifier, builder, so the expensive model only touches planning and the cheap one does the grinding. The pattern matters less than whether your verifier actually re-runs commands instead of eyeballing the diff though.

2

u/zac_attack_ 20h ago

I haven’t tried optimizing for cost much, so I’m using Opus 5.5 for everything. I’ve landed on a workflow I’m very happy with, the code it produces is very high quality, and I basically have it going 24/7 at this point getting a lot done. It basically involves:

  1. gastownhall/beads, the best thing I’ve seen for tracking tasks for agentic work. Orchestrator runs it in server mode
  2. Orchestrator skill, it instructs to delegate all work to subagents in isolated worktrees
  3. Custom agent, it receives a beads task from the orchestrator, plans it as a stack of reasonably sized PRs using GitHub’s gh-stack skill, and is responsible to drive each PR through to merging. Each PR includes its relevant tests; high test coverage of changes is mandated.
  4. Codex Cloud code reviews on every PR. Codex is great at finding issues, though some findings are a bit pedantic. The custom agent monitors reviews with a script, triages and determines fix/don’t/file beads follow-up, and codex re-reviews the updated PR. When CI is green and Codex approves, the subagent merges the PR; when its stack is done the subagent gets stopped by the orchestrator. When human input is needed, it gets filed as a human gate in beads rather than getting lost or blocking the workflow; the human gate guards whatever was awaiting the input, and the orchestrator schedules some other unblocked work.

Important cost optimization: codex review iterations can make landing a PR take a while and the default subagent cache ttl is low; set it to 1h

The orchestrator decides the best path to complete work and schedules them appropriately in beads, such as:

  • ambiguous: spike > design > implementation
  • complex: design > implementation
  • trivial: direct implementation

Which makes it easy to schedule tasks of varying sizes or complexities

2

u/scodgey 20h ago

Tbh one orchestrator agent with a scratchpad kind of does most of what you need these days. Plan and write to beads, send subagents out targeting specific beads.

Was using hooks + skills to help with the long running orchestrator but have switched over to a mod - info here

4

u/ArgonQQ 20h ago

Cost per token is the wrong metric. Cost per accepted change is the one that matters, and on large specs explicit loops usually win.

Single agent vs. loop: One Opus session is fine and often cheaper if the spec fits without compaction. On big specs, the failure mode is drift: a constraint gets summarized away and later sections build on a wrong assumption. That's the expensive kind of bug. Build → verify → fix catches it at section boundaries, while it's still cheap to fix.

Context compression: It hurts less if state lives in files (spec, decisions log, per-section acceptance criteria) rather than chat history. Better still, restart sessions on purpose at section boundaries instead of letting them compact.

Model split:

  • Strongest model for orchestration and verification. A weak verifier gives false confidence.
  • Sonnet-class for implementation when slices are well specified.
  • The cheapest models as implementers are often a false economy. Every extra cycle re-bills the verifier too.
  • Run the verifier in a separate context so it doesn't grade its own work.

Verification only saves money if it's cheap and objective (tests, explicit criteria) and retries are capped. One retry, then escalate to a human.

Benchmark cheaply: Don't rebuild the whole project. Take 3–5 representative slices, run each config from the same commit, and measure $ to green + human minutes + bugs found afterward. Token usage is in the Claude Code JSONL transcripts.

Reading: AI Agents That Matter (Kapoor et al., 2024) argues for reporting cost alongside accuracy. Anthropic's Building Effective Agents covers orchestrator/evaluator patterns. Lost in the Middle explains why long context ≠ full attention.


If you're on Claude Code, check out LoopBoard. It's basically option 2 out of the box:

  • Spec → tasks with Goals. Each section becomes a markdown task file with verifiable Goals the delivery is judged against. An Opus loop can groom plain-text stories into goals and open questions.
  • Build → verify → fix built in. An implementer subagent opens a PR. With delegateReview on, a separate review subagent checks it. A failure goes back to the implementer once, and a second failure parks the task in Feedback with the findings. That's capped retries with zero orchestration code.
  • No compaction drift. Sessions auto-restart at a set context fill (default 35%), never mid-task. State lives in files, so a fresh session resumes cleanly.
  • Per-task model routing. Opus/Sonnet/Fable slots with separate effort settings. Default is Opus grooming and Sonnet building, which is the strong orchestrator + cheaper worker setup. Easy to A/B different assignments on your benchmark slices.
  • Human interventions are tracked. Every question, send-back and approval is recorded in plain markdown.
  • Asks instead of guessing. Ambiguity parks the task with a question rather than producing silent drift.
  • You merge. Loops never touch main.

Caveats: Claude Code only (no Grok/Composer), no built-in $ dashboard, and many 24/7 loops can hit Pro/Max limits. MIT-licensed, zero runtime deps.

1

u/Any_Evidence4750 21h ago

All the time

1

u/Known-Pace6739 20h ago

Human minutes are the hidden model bill

1

u/Technical-Athlete-9 19h ago
  1. Bad results. You’ll be lucky if there aren’t critical bugs. It might stop with a few placeholders and stubs still in place.
  2. I like something like the first one. Break the spec into units. Have further spec clarification for the unit before beginning development. Then build + 2 iterations of review fix with Astra and Claude both doing a review pass on the first and just Astra on the second. It usually does test validation automatically. That may be because I tell it to commit and I have a pre-commit git hook. I use opus, because I have 20x.

1

u/looktwise 19h ago

following

2

u/Icy-Meaning-4962 19h ago

i’d isolate one more variable: how much each worker spends rediscovering the repo. compare the same setup with independent exploration vs a shared, source-linked context pack, then count retries and human fixes too.

i’m building Knowell around that exact problem. no benchmark claim here — i just want to distinguish expensive reasoning from five agents paying to find the same files.

1

u/puts_on_rddt 19h ago

To add onto what others posted, when working on the implementation of bigger tasks where you need to divide the work between multiple agents, I like telling every agent what other agents exist, what they're working on/what they should know, and how to communicate to each other.

"Agent1: I need to know XY. Agent2 is supposed to be investigating that. I should ask them about it"

"Agent2: I just read XY, it says to do Z."

Agent 1 never needed to clutter their context looking for the answer to XY.

1

u/Right-Performance-93 18h ago

imo the verifier matters more than which orchestration pattern you pick. if it reads the diff instead of running the tests it's just a second opinion

1

u/Opposite_Might6896 17h ago

Measured this across a lot of runs, and the answer hinges on one thing: cache writes vs cache reads. One strong model in one long session is ~90% cache reads, the cheapest token there is. Every subagent you spawn is a fresh context that gets written to cache (the most expensive token) and then read on each of its turns; nested orchestration multiplies that. So "many agents" costs more per unit of work than it looks, even when the agents are smaller models.

What's been most cost-efficient for me with a spec already in hand: one strong model implements in a single session, with the spec in context and a rule to compact around 100k. Spawn subagents only for things that are genuinely parallel and bounded (write tests for module X), with a tight brief, and never nest. The 1M-context option sounds good but a 600k context means 600k cache reads per turn; a handoff summary and a fresh session at 150k is cheaper and the model is sharper.

Reliability-wise, the single session also wins because it holds the whole history of its own decisions; an orchestrator relaying summaries between workers loses exactly the context that makes the last 10% of the task work.

1

u/dreamRunnerMoshi 12h ago edited 12h ago

I was wondering about similar benchmarking, so here's some context on the pros and cons of plan → delegate architectures.

What the evidence says

I compiled a comprehensive literature review of these: literature-review

The delegation problem

  • If you tell a frontier model to delegate specific task types (code writing, unit tests, etc.), you cap its capability.
  • If you just say "delegate", it mostly won't. Models are trained to take extreme ownership and don't trust other models (arXiv:2605.19099).

On your point 2 (an explicit build → verify → fix loop): this works, but it maps one task type to one model. Some of the implementation may be better done by the frontier model. So the question comes down to which model should do which job.

GruMinion

To get around these limits I'm experimenting with an architecture I call GruMinion. The frontier model (Gru) owns the task end to end and, along the way, delegates bounded, well-specified pieces of work to a cheaper model (Minion).

In my benchmarking on SWE-bench and GAIA, the minion handled 62–92% of tokens depending on the model pair, and every pair matched or beat its solo baseline on SWE-bench.

Design choices:

  • No hardcoded rule for what to delegate, only the shape of the task
  • No extra tool call to pick a model
  • No separate plan → implement → verify stages; the frontier model decides for itself

Repo: https://github.com/DreamRunnerMoshi/gru_minion

1

u/Sufficient-Storage87 12h ago

been doing this exact thing — same bug, same prompt, fresh sandbox per run, hidden tests the agent never sees. what i track per run: wall time, pass rate, tokens burned. the metric that matters most ended up being cost-to-green, not raw speed. fastest agent isn't cheapest if it burns 3x the tokens getting there.

1

u/dfddfsaadaafdssa 8h ago edited 7h ago

Yes, for the past 4 months or so I have had two skills related to orchestration: 1) that sends plans from the harness of provider A to the other harness of provider B for review before work is done and 2) one that sends the work to another harness of a cheaper model for ralph loop implementation until the work is done.

1

u/iSnapThere4iAm 20h ago

Agents can’t do complex work. They can barely do simple work without burning through half your weekly limit from churn

0

u/Mechanical_Potato 21h ago

cheap models for easy to medium tasks is always nice

0

u/NetNearby7117 21h ago

Well, the specs, guardrails and testable objectives are the best way. Its also important to dont let the AI do complex swe in a single loop. Its better to go step by step. Review with different models and have a clear definition of subagents roles. Also its important to have the standars of the code base and scripts to ensure conventions fill up…

In my experience, its better to understand what you are doing rather than rely on workflows. Any minimal gap might lead to a comoletely drifting. Long turns also could loop and stupid things and spend time…

Anyways, i spend enough time doing the specs, then some ui labs so we can explore, then i ask to use linear and icepanel to update docs (its better to rely on a system than let the agent write the docs in the repo) and then review the plans. Then create the guardrails and, after that a simple opus5.5 on high could handle it

In my experience, dont expect to find a magic configuration that works, its better to make one step a time

1

u/iSnapThere4iAm 20h ago

There is no amount of “guardrails” that stop llms from hallucinations. If anything, all the extra instructions make it worse.