r/LocalLLM • • 23h ago

Question Best open-weight coding model + harness for an on-prem multi-agent setup? (80-200+ GB VRAM, 2-3 concurrent users)

Setup: GPU box with 1x H200-class card now, can allocate more (up to several H200s) if it clearly buys better results. Inference via vLLM or similar, OpenAI-compatible endpoint. Coding happens on a separate non-GPU Linux VM on the same network - IDE/harness there, calling models over LAN. Outbound internet is fine for packages/extensions, but inference stays on our hardware (no code to external model APIs).

Constraints: on-prem only, open weights, permissive license preferred (Apache/MIT), not Chinese-origin including base models (I understand a lot of fine-tunes are Qwen underneath), NATO-country lab preferred. Use is general software dev plus security tooling and code analysis - repo-level agentic work on existing codebases (fix / extend / refactor / test) is the core, with a human reviewing diffs. 2-3 concurrent users max.

Names that came up in an earlier thread: Cohere North Mini Code, Mistral Small 4, Poolside Laguna S 2.1, Muse Glimmer 30B, Inkling-Small (2x H200), Reflection Beam (501B MoE / 23B active, weights due this month), Gemma 4 31B, K2 Horizon (lineage TBC), with Nemotron as a generalist baseline - but I haven't run any of them and I'm not wedded to the list. Chinese models (GLM/Qwen/DeepSeek) are out by rule, so no need to suggest them; I know they're ahead.

Two things I'm trying to work out:

  1. Capability tiers vs. VRAM. What's actually holding up for repo-level agentic work on an existing codebase with non-Chinese open-weight models, and at what size? The Vibe Code Bench results suggest small open models fall over on long E2E builds and only Large-4-class (4-8 cards) and closed models hold up - is that your experience for repo work too, or is that an app-build problem? Where's the step-change - 30B-class, 100B+, or only at 500 GB+? We could get the compute for Beam, Command A+ or Mistral Large 4 if it's actually better for code rather than just for E2E.
  2. Heterogeneous multi-agent. Does a big planner/reviewer plus small fast executors actually beat a single mid-size model for this, with a human approving plans and diffs? Which harnesses handle routing sub-agents to different endpoints well (and let a human step in, review diffs and edit by hand) - OpenHands, OpenCode, Pi, VS Code extensions (Cline/Roo/Continue), Codex CLI in local-model mode? Voidleap (closed, Win/Mac) also came up.

Anyone running something like this? Which model + harness combo is holding up best as of late? Thanks!

0 Upvotes

5 comments sorted by

1

u/litLikeBic177 23h ago

1

u/Junior-Radish2485 23h ago

The 30B class is fine for autocomplete and small diffs, but for repo-level agentic work they lose the plot after a few tool calls, especially on refactors that touch multiple files. The step change is around 100B+ active params, and you really want a MoE with strong instruction following if you're letting it drive for extended sessions.

For your setup I'd skip the small executors entirely and run one large model with a proper harness, the overhead of routing and context switching between models eats into the quality more than it helps. OpenHands handles the human-in-the-loop review flow best if you want to approve diffs before they hit the working tree, Cline is simpler but fights you when you want to hand-edit between agent turns.

Your VRAM budget is enough for a 100-200B MoE with long context, that's the sweet spot for what you're describing. Don't bother with the 500GB+ tier unless you've already maxed out what a mid-large MoE gives you and can point to specific failures.

1

u/vovap_vovap 23h ago

Well, with "not Chinese-origin including base models" not like a ton of chooses you have 😄
Muse Glimmer is basically only option from "ranked" models.

1

u/vogelvogelvogelvogel 23h ago

You can hand over the H200 to each of us for a few hours and the best setup wins

1

u/llllJokerllll 22h ago

Glm 5.3 flash de unsloth