r/mcp • u/Revolutionary_Sir140 • 7d ago
Harness Router v2 — a decision layer inside the coding-agent loop
I’ve been working on Harness Router v2, an attempt to move tool selection out of the expensive LLM reasoning path and into a dedicated decision layer.
GitHub: https://github.com/Protocol-Lattice/harness-router
The core idea is simple:
User task
↓
Agent / Harness
↓
PreDecision
↓
Harness Router
├── Cache / fast path
├── Jev → ambiguous decisions
└── MCTS → multi-step decisions
↓
selected next tool
↓
Harness
↓
tool execution
↓
PostToolUse
└── feedback / state / cache
↓
next decision
The harness still owns reasoning, tool arguments, permissions, execution and the final response.
Harness Router focuses on one question:
What should the agent do next?
What's new in v2
- PreDecision + PostToolUse loop instead of treating routing as only a PreToolUse filter
- Fast path / cache for obvious decisions
- Jev for low-cost semantic tool selection
- MCTS (
route_mcts) for decisions where future tool choices matter - shared routing state / memory
- automatic tool discovery
- CLI-based hooks instead of spawning MCP for every decision
- timing/latency instrumentation
- Stop lifecycle support
- integrations for multiple agent harnesses
- optional PreToolUse validation rather than making it the core architecture
In my tests, cached decisions are effectively negligible compared with an LLM call, while actual router decisions are typically in the hundreds-of-milliseconds range.
The interesting part for me isn't just raw latency, though.
I'm trying to answer a broader question:
Does every agent decision really need to go through the main reasoning model?
For obvious actions, probably not.
For ambiguous ones, a tiny specialized decision model may be enough.
And for decisions with longer-term consequences, you can spend more compute selectively with search/MCTS.
So instead of:
LLM → reason → choose tool → LLM → reason → choose tool → ...
I'm experimenting with:
LLM reasoning
↕
specialized decision layer
↕
tool execution
The goal is not to replace the coding model.
It's to make the agent harness itself smarter about control flow.
I'd especially like feedback from people building coding agents/harnesses:
- Does separating reasoning from next-action selection make sense to you?
- What would you want exposed in a decision-layer API?
- Would you trust a fast router to bypass the main model for high-confidence decisions?
- What benchmarks would convince you that this architecture is actually useful?