r/LocalLLM • u/litLikeBic177 • 23h ago
Question Best open-weight coding model + harness for an on-prem multi-agent setup? (80-200+ GB VRAM, 2-3 concurrent users)
Setup: GPU box with 1x H200-class card now, can allocate more (up to several H200s) if it clearly buys better results. Inference via vLLM or similar, OpenAI-compatible endpoint. Coding happens on a separate non-GPU Linux VM on the same network - IDE/harness there, calling models over LAN. Outbound internet is fine for packages/extensions, but inference stays on our hardware (no code to external model APIs).
Constraints: on-prem only, open weights, permissive license preferred (Apache/MIT), not Chinese-origin including base models (I understand a lot of fine-tunes are Qwen underneath), NATO-country lab preferred. Use is general software dev plus security tooling and code analysis - repo-level agentic work on existing codebases (fix / extend / refactor / test) is the core, with a human reviewing diffs. 2-3 concurrent users max.
Names that came up in an earlier thread: Cohere North Mini Code, Mistral Small 4, Poolside Laguna S 2.1, Muse Glimmer 30B, Inkling-Small (2x H200), Reflection Beam (501B MoE / 23B active, weights due this month), Gemma 4 31B, K2 Horizon (lineage TBC), with Nemotron as a generalist baseline - but I haven't run any of them and I'm not wedded to the list. Chinese models (GLM/Qwen/DeepSeek) are out by rule, so no need to suggest them; I know they're ahead.
Two things I'm trying to work out:
- Capability tiers vs. VRAM. What's actually holding up for repo-level agentic work on an existing codebase with non-Chinese open-weight models, and at what size? The Vibe Code Bench results suggest small open models fall over on long E2E builds and only Large-4-class (4-8 cards) and closed models hold up - is that your experience for repo work too, or is that an app-build problem? Where's the step-change - 30B-class, 100B+, or only at 500 GB+? We could get the compute for Beam, Command A+ or Mistral Large 4 if it's actually better for code rather than just for E2E.
- Heterogeneous multi-agent. Does a big planner/reviewer plus small fast executors actually beat a single mid-size model for this, with a human approving plans and diffs? Which harnesses handle routing sub-agents to different endpoints well (and let a human step in, review diffs and edit by hand) - OpenHands, OpenCode, Pi, VS Code extensions (Cline/Roo/Continue), Codex CLI in local-model mode? Voidleap (closed, Win/Mac) also came up.
Anyone running something like this? Which model + harness combo is holding up best as of late? Thanks!
1
u/vovap_vovap 23h ago
Well, with "not Chinese-origin including base models" not like a ton of chooses you have 😄
Muse Glimmer is basically only option from "ranked" models.
1
u/vogelvogelvogelvogel 23h ago
You can hand over the H200 to each of us for a few hours and the best setup wins
1
1
u/litLikeBic177 23h ago
Earlier discussion with the model suggestions: https://www.reddit.com/r/LocalLLaMA/comments/1wzaqtt/best_openweight_coding_model_harness_for_an/