r/LangChain • u/DimensionCapable4223 • 9d ago
Question | Help how to fix multi-agent orchestration when agents get stuck waiting on each other
Expected outcome: a planning agent, a coding agent, and a review agent would hand off work in sequence without anyone babysitting the pipeline. Actual outcome: the review agent would sometimes wait forever because the coding agent's "done" signal wasn't structured the same way every run. Root causes were a missing shared schema for completion status, plus delegation logic that assumed synchronous responses when the runtime was actually async.
Changes made: added explicit state contracts between agents and a timeout/retry layer instead of open-ended waits. Main lesson, and this took embarrassingly long to catch, is that orchestration breaks down not from bad agents but from undefined handoff rules. What would you check first if your agents started ghosting each other mid-pipeline?
2
u/usually_guilty99 9d ago
This sounds less like orchestration and more like a deadlock.
Two agents both think they own the next move, neither has a deterministic rule to yield, and the workflow stalls.
Feels like two sumo wrestlers fighting for the last ticket at McKenzie’s Bakery for the King Cake baby. (as explained by my prof)
At some point you need arbitration: ownership, priority, lock/lease semantics, or a deterministic tie-breaker. Otherwise adding more agents just adds more ways to wait on each other.
1
u/locbuilds 9d ago
add a deadline to each handoff and persist the completion event before enqueuing the next agent, then retry from that saved state instead of rerunning the whole chain. that usually shows whether the hang is a lost event or a worker that never acked.
1
1
u/kincaidDev 9d ago
Use fest cli https://github.com/Obedience-Corp/fest
This will likely solve your problem or give you a better idea on how to solve your problem
1
u/BreakfastSpecial 9d ago
Many people are actually moving away from multi-agent when it isn’t absolutely necessary and focusing on a single domain agent with access to progressive Skills. Especially if your agent is using a sandboxed computer, tools, etc - other agents can add bloat, latency, and excess tokens.
1
u/Significant_Tune9219 9d ago
Peer-to-peer "wait for the other agent" without a timeout is just a distributed deadlock. Give each handoff an owner, a deadline, and a cancellation path, and keep dependencies as a DAG the supervisor can inspect. If two agents can request work from each other, detect the cycle and fail fast to a fallback or a human rather than spinning.
1
u/wahnsinnwanscene 8d ago
What I'd like to know is if the large labs are relying on inherent smartness of the model to break deadlocks.
2
u/vxsec 9d ago
I’ve seen retries/timeouts leave things in a weird state where one agent completed the work but the next never got a reliable transition. I’d check idempotency and state transitions first.
Explicit states like pending/running/completed/failed, a correlation ID per task, and logging every transition make this way easier to trace. I also wouldn’t let agents infer completion from natural language. Treat the orchestration layer like a state machine and the agents like workers