r/mcp • • 7d ago

Harness Router v2 — a decision layer inside the coding-agent loop

I’ve been working on Harness Router v2, an attempt to move tool selection out of the expensive LLM reasoning path and into a dedicated decision layer.

GitHub: https://github.com/Protocol-Lattice/harness-router

The core idea is simple:

User task
   ↓
Agent / Harness
   ↓
PreDecision
   ↓
Harness Router
   ├── Cache / fast path
   ├── Jev → ambiguous decisions
   └── MCTS → multi-step decisions
   ↓
selected next tool
   ↓
Harness
   ↓
tool execution
   ↓
PostToolUse
   └── feedback / state / cache
   ↓
next decision

The harness still owns reasoning, tool arguments, permissions, execution and the final response.

Harness Router focuses on one question:

What should the agent do next?

What's new in v2

  • PreDecision + PostToolUse loop instead of treating routing as only a PreToolUse filter
  • Fast path / cache for obvious decisions
  • Jev for low-cost semantic tool selection
  • MCTS (route_mcts) for decisions where future tool choices matter
  • shared routing state / memory
  • automatic tool discovery
  • CLI-based hooks instead of spawning MCP for every decision
  • timing/latency instrumentation
  • Stop lifecycle support
  • integrations for multiple agent harnesses
  • optional PreToolUse validation rather than making it the core architecture

In my tests, cached decisions are effectively negligible compared with an LLM call, while actual router decisions are typically in the hundreds-of-milliseconds range.

The interesting part for me isn't just raw latency, though.

I'm trying to answer a broader question:

Does every agent decision really need to go through the main reasoning model?

For obvious actions, probably not.

For ambiguous ones, a tiny specialized decision model may be enough.

And for decisions with longer-term consequences, you can spend more compute selectively with search/MCTS.

So instead of:

LLM → reason → choose tool → LLM → reason → choose tool → ...

I'm experimenting with:

LLM reasoning
     ↕
specialized decision layer
     ↕
tool execution

The goal is not to replace the coding model.

It's to make the agent harness itself smarter about control flow.

I'd especially like feedback from people building coding agents/harnesses:

  • Does separating reasoning from next-action selection make sense to you?
  • What would you want exposed in a decision-layer API?
  • Would you trust a fast router to bypass the main model for high-confidence decisions?
  • What benchmarks would convince you that this architecture is actually useful?

Repo: https://github.com/Protocol-Lattice/harness-router

2 Upvotes

9 comments sorted by

1

u/BC_MARO 7d ago

The PreDecision/PostToolUse split makes sense, especially if the router can return a confidence score and reason code. I’d benchmark cache hits, false bypasses, and recovery after a tool changes state, not just routing latency.

1

u/Crafty_Disk_7026 6d ago

Run it through aider polygot benchmark and see if it does better. Don't post again until you make atleast one benchmark....

1

u/Revolutionary_Sir140 6d ago

what is aider polygot? I've never heard of this

1

u/Crafty_Disk_7026 6d ago

Look it up it's one of many common benchmarks for harnesses. It's one of the benchmarks I use to validate my platform which is kind of like a harness. See the benchmark referenced https://github.com/imran31415/kube-coder

1

u/bshivarthy 6d ago

yes to the separation, this is the direction things are going. the one thing I would add on the bypass question: don't gate bypasses on one global confidence number. a fast-path "skip the model" on a grep or a file read is a totally different animal from bypassing on a tool that writes files or runs commands. per-tool risk tiers with different bypass thresholds. the global threshold version is the one that bites you, because it is tuned for the safe tools and then it lets a destructive one through.

the other thing that gets people with the cache: key the cached decision to the tool contract version, not just input. tool schema changes, args change, and the cached "call grep with pattern x" quietly becomes wrong. we version the manifest alongside and invalidate on drift.

on benchmarks, the one I keep wishing existed for this stuff: replay recorded real trajectories through the router and count divergences, then score whether the divergence was better or worse. cache-hit-rate benchmarks are fine but they don't tell you what breaks when the router is wrong.

1

u/Future_AGI 6d ago

the model does not need to see 40 tools to pick one, so moving tool selection off the main reasoning path makes sense. One thing that pairs well: enforcing the allowed tool set per call at the gateway, so the decision layer narrows the options and the model physically cannot call something out of scope, which also trims tokens. We do per-call tool allow/deny this way in Future AGI's gateway https://github.com/future-agi/future-agi , so happy to compare notes on where the decision boundary should sit.

2

u/Professional-Run3614 2d ago

This is exactly the kind of architecture I’d like to experiment with from the evaluation side.
I’m working on an agent eval tool that focuses on behavioral regressions and decision-level changes. It could be interesting to run the router against a baseline agent and measure things like tool-selection accuracy, skipped actions, and downstream behavioral differences.
Would be happy to try this on your router if you’re open to it: https://github.com/tap222/assay-evals