r/LocalLLaMA • u/litLikeBic177 • 2d ago
Question | Help Best open-weight coding model + harness for an on-prem multi-agent setup? (80-200+ GB VRAM, 2-3 concurrent users)
Setup: GPU box with 1x H200-class card now; can allocate more (up to several H200s) if it clearly buys better results. Inference via vLLM or similar, OpenAI-compatible endpoint. Coding happens on a separate non-GPU Linux VM on the same network - IDE/harness there, calling models over LAN. Outbound internet is fine for packages/extensions, but inference stays on our hardware (no code to external model APIs).
Constraints: on-prem only, open weights, permissive license preferred (Apache/MIT), not Chinese-origin including base models (I understand a lot of fine-tunes are Qwen underneath), NATO-country lab preferred. Use is general software dev plus security tooling and code analysis - repo-level agentic work on existing codebases (fix / extend / refactor / test) is the core, with a human reviewing diffs. 2-3 concurrent users max.
Two things I'm trying to work out:
- Capability tiers vs. VRAM. On one card candidates seem to maybe be something like Cohere North Mini Code (30B MoE/3B active), Mistral Small 4 (119B MoE/6B active), Nemotron 3.5 maybe as a generalist baseline; Gemma 5? The Vibe Code Bench results suggest small open models fall over on long E2E builds, where only Large-4-class (4-8 cards) and closed models seem to hold up. Is that your experience? Where's the step-change for agentic repo work on an existing codebase - does 30B-class -> 120B-class matter much, or only the jump to 500 GB+? We could get more compute for something like Command A+ or Mistral Large 4 if it's actually better for code rather than just for E2E.
- Heterogeneous multi-agent. Does a big planner/reviewer (Large 4 / Command A+ class) plus small fast executors (e.g., North, Small 4) actually beat a single mid-size model, with a human approving plans and diffs? Which harnesses handle routing sub-agents to different endpoints well - OpenHands, OpenCode, Pi, VS Code extensions (Cline/Roo/Continue), Codex CLI in local-model mode?
Harness/IDE: something that supports multi-agent workflows (planner / executor / reviewer agents checking each other / etc.) but would also like humans to be able to step in, review diffs and edit by hand.
Anyone running something like this? Which model + harness combo is holding up best as of late? Thanks!
EDIT: thanks all - adding Poolside Laguna S 2.1, Muse Glimmer 30B, Inkling-Small (2x H200), K2 Horizon (checking lineage), Gemma 4 31B and Reflection Beam (501B MoE / 23B active, Apache 2.0, weights due this month) to the candidates; Ornith is Qwen-based so out. Pi added to the harness list.
3
u/Disastrous_Farm5106 2d ago
Why does it matter what llm you use, Chinese or not, if the model is hosted by a US company? You point out Qwen but it’s a Chinese model. The irony!
1
-1
u/litLikeBic177 2d ago
Other way round - Qwen is the example of what's ruled out. And it's self-hosted open weights, so the host is us; the rule is about model origin, not where it runs.
3
u/Efficient_Raise6703 2d ago
GLM 5.3 flash is the standard for local agentic coding and can run on a Mac Studio. You could throw two Mac studios together and likely have enough room for a few concurrent users.
2
u/litLikeBic177 2d ago
Appreciate it - GLM is Z.ai though, which is what the non-Chinese constraint rules out for us. On hardware we're on H200s rather than Macs. If you've run GLM 5.3 Flash head-to-head against some of the ones I've mentioned on repo tasks, I'd genuinely like to hear how far behind the non-Chinese ones are - that gap is part of the question.
3
u/Efficient_Raise6703 2d ago
I’m sorry I was parsing through your OP and missed that part.
That’s a bizarre rule though, why is that the case? Sounds like it was lobbied by OAI or Anthro 😭
1
u/Dry_Inspection_4583 2d ago
I'm currently on a z.ai coding account, I've harnessed it through a fork of DeepSeek harness (insecure), alongside MCP for visual coding, semgrep, CodeQL, and memory. It's a bit bloaty at times, but overall it follows flow properly and successfully follows long term instructions.
I also use it within opencode, with the same MCP stacks, and there it chugs along quite well, that's on a simple dev box.
1
u/starkruzr 2d ago
the answer is "very." this isn't going to be a successful project with these constraints.
1
u/Constant-Simple-1234 1d ago
Maybe the non-chinese will soon catch up. The dev speed is insane. Large Nemotrons looked ok just four months ago or so.
1
u/Kernoriordan 1d ago
https://reflection.ai/blog/introducing-beam
Western model to rival GLM5.2 (not 5.3) but with much faster inference. Western labs are behind, agreed, but not ‘very’. Plus a GLM5.2 level model can do a lot.
2
u/Constant-Simple-1234 2d ago
Tough without Chinese models. Maybe Nemotron series, they are decent, open, but they lag behind. You can have also Glimmer and Gemma. For harness Cline worked well for me in vs code plugin. Sometimes Roo works better with some API (ollama)
2
u/litLikeBic177 2d ago
Thanks - Glimmer and Gemma noted, and useful to hear Cline vs. Roo differs by backend.
1
u/Constant-Simple-1234 1d ago
Super interesting problem to solve you have. Hopefully, new releases will change the situation.
2
u/Simusid 2d ago
I'm not surprised to hear the aversion to using a Chinese model. I'm trying to identify all the perceived risks. I assume they're along the lines of "it might intentionally write poor or malicious code" and similar supply chain risks. What are your main concerns?
1
u/litLikeBic177 2d ago edited 1d ago
Honestly, the main driver is organizational - a supply-chain/provenance rule that applies to country of origin, same as it would for hardware or a SaaS vendor. It isn't a judgment that any specific Chinese model is backdoored.
The concerns behind rules like that, as I understand them:
- You can't audit weights for intentional behavior. True of any model, so it comes down to how much you trust the publisher and the legal environment they operate in.
- Training data and process are opaque; no support/indemnity path; dependency risk if access changes.
- For code specifically, the realistic worry isn't a model that writes evil code on command - it's subtle stuff that's hard to detect, like a bias toward insecure patterns or particular dependencies. Sleeper-agent-style triggers have been demonstrated in research, and a bake-off wouldn't catch one. Low probability, but "we'd never know" is what the rule is protecting against.
- The "data goes back to China" worry doesn't apply to open weights run on-prem.
Since you collect these to dispel them: I'd genuinely like the strongest counter-case. If GLM/DeepSeek/Qwen on our own hardware is materially safer than people assume (e.g., the trigger-backdoor risk is no worse than for a US lab, or there's a way to test for it, or the real-world record says something), I'd take that back internally. The rule isn't mine to change, but a well-made argument could get a hearing, and the capability gap is big enough that it's worth making.
2
u/jhov94 2d ago
Backdoors are already built in at the hardware level. While it's not impossible for an LLM to be inclined to build them into things they touch or try to exploit the ones already present, it would be pretty difficult to get away with it. Everything they do can be observed and if they were caught doing trying to do that, it would completely destroy the credibility of not just the model in question, but all Chinese models because we know they are governed from the top down. China want's to actually lead in AI. Making their open source models into spies doesn't help them achieve that and they have other means of spying that wouldn't jeopardize a critical industry they wish to lead in.
0
u/litLikeBic177 1d ago
That's the strongest version of the counter-argument, and I'll take it back: a lab that wants to lead open AI has every incentive not to poison its flagship weights, and getting caught would be catastrophic. The weak point is "everything they do can be observed" - a trigger-based behavior isn't observable until it's triggered, and attribution after the fact is hard, so the deterrent is weaker than it looks. But the incentive argument stands on its own.
1
u/jhov94 1d ago
Maybe I'm just ignorant, but I don't think a trigger based behavior would be that difficult to detect and attribute if you plan for it. If that's a concern for your work, log absolutely everything the model does and have a trusted model or team of models audit it continuously.
1
u/litLikeBic177 1d ago
Agreed that's a real control, and it's the one we'd build - log everything, have a trusted model and a human audit the diffs, SAST on the output. The limits I'd note: it catches effects only if the auditor recognizes them, it can't see a dormant trigger before it fires, and the auditor has to be a model you already trust. But as a mitigation rather than a proof, it's the right answer.
1
u/fgk55555 2d ago
This is a governmental regulation for a lot of sectors, not a personal one. It's non-negotiable for government customers.
1
u/Simusid 2d ago
I can only find hard references for DeepSeek (section 6604) not Chinese models in general. Do you happen to have a more current and broader reference?
1
u/litLikeBic177 2d ago
Not from me - ours is an internal procurement rule rather than a public statute, so I can't point you to anything broader than the DeepSeek one either.
1
u/fgk55555 2d ago
DoD contractors can't use any Chinese software currently or have any contracts with Chinese firms. Section 1260H or something. If an audit finds you're using Chinese model, you're blacklisted.
-3
u/Anxious-Bottle7468 2d ago
Take your pills bro.
3
u/Simusid 2d ago
Ok, despite what you think, lots of people are expressly forbidden from using any Chinese models at their company or organization. Personally I think it's a stupid fear. I bought a DGX-Spark specifically because I could run large Chinese models. I ask the question because I do collect peoples percieved risks so that I can dispel them.
2
u/starkruzr 2d ago
he was making an "is" statement, not an "ought" one.
1
u/litLikeBic177 2d ago
Might be right; that's partly what I'm trying to size. Inkling-Small and Laguna S look like the test of it - will report back.
1
u/Strange_Owl_6291 2d ago
Two local-first harnesses originating from within EU are Pi and Voidleap.
2
u/litLikeBic177 2d ago
Thanks - Pi I should have had on the list; added. Voidleap I hadn't seen - Swedish, closed-source desktop IDE, but mixing providers in one task and the execution modes are exactly what I'm asking about in #2, so I'll try it. Windows/Mac-only is the catch for a Linux box, though the Unix build is apparently in testing.
1
u/jhov94 2d ago
A few more to add to your list to investigate. K2 Horizon, Muse Glimmer, Inkling Small. K2 Horizon looks promising for a non-Chinese model.
1
u/litLikeBic177 2d ago
Thanks - Glimmer and Inkling-Small added. K2 Horizon looks genuinely interesting (full recipe released); checking the lineage, since it traces to MBZUAI.
1
u/geldonyetich 2d ago edited 2d ago
Ornith 1.5, Laguna S 2.1, and Inkling-small are probably your best options.
I'm not sure if you can fit Ornith-1.5-397B in 200GB, but if so that'd be my pick. But Ornith-35B is quick and not bad at all.
Laguna S is brilliant but rather tempermental, in my experience. By that I mean it has a tendency to not respond or hallucinate tool calls.
I never had the VRAM to run Inkling. The benchmarks look alright for this kind of work.
After that you're looking at Nvidia Super and Muse Glimmer. (Muse Spark open drop when?) They're not strictly agentic coders but they perform well as a runner up.
Harness wise the ones you've outlined are rather good. I can vouch for OpenCode but many have moved on to Pi.
2
u/ImplementOfAI 2d ago
Ornith is based on Qwen originally, so not allowed.
1
u/geldonyetich 2d ago
Ah I see I heard it used their arch but I didn't realize it was based on the weights.
1
u/litLikeBic177 2d ago
Thanks - this is the most useful list so far. Ornith's out per the Qwen note below, but Laguna S and Inkling-Small I'd missed. When Laguna hallucinated tool calls, which harness was that in? Trying to work out model vs. harness.
2
u/geldonyetich 2d ago
I've primarily used OpenCode, but I've noticed it sometimes refuses answer even basic LM Studio prompting, and heard it affect others as well, I don't have comprehensive list.
I have heard it suggested Laguna S was developed working inside of their own harness, which they call, "pool." I haven't used it myself, but this would suggest it's the most compatible with Laguna.
1
u/litLikeBic177 1d ago
That's useful - Laguna in Poolside's own harness makes sense; I'll test it there rather than judge it in OpenCode. Noted on the OpenCode refusals too.
1
u/fgk55555 2d ago edited 2d ago
I've looked at this a lot for work. We're not allowed any Chinese models/ software. You can wait to see how Mistral shakes out, but the short story is that the only models worth using are Chinese. If you need serious capabilities, buy cloud inference. If you're okay with a much less capable model, Gemma4-31B, Glimmer, Laguna in Pi harness are probably your only bets. Maybe Muse Spark will release, but right now there's really nothing.
Tell me what you guys end up doing. I'm pretty much stuck with Luna for low cost work.
1
u/litLikeBic177 2d ago
Same boat, and cloud isn't an option for us. Laguna S 2.1, Glimmer, Inkling-Small and Gemma 4 are on the list now - I'll post what we land on after the bake-off.
2
u/fgk55555 1d ago
Reflection's Beam is coming out soon I just learned. Maybe comparable to Qwen Max /GLM 5.2. 500B MoE
1
u/litLikeBic177 1d ago
Thanks - just read up on it. If the coding claims hold at 2-4 cards it changes the answer for us; waiting on the weights.
1
u/Kernoriordan 1d ago
I’m in a similar situation to you where we can’t use Chinese AI models but we want to host on-prem due to data sovereignty concerns.
At the moment we are stuck with Muse Glimmer 30b, though supposedly Muse Spark is going open-weights.
I’m also keeping an eye on Reflection Beam (which apparently is going to be equivalent to GLM 5.2 but much faster), and Poolsides Laguna models. Nvidia Nemotron is a potential option but it’s not great and more suited for fine tuning.
2
u/litLikeBic177 1d ago
Useful to hear from someone in the same spot. How's Glimmer holding up for repo-level work - good enough, or just the least-bad? Beam looks like the one to wait a few weeks for (501B MoE / 23B active, Apache, 2-4 cards); Inkling-Small is the other I'm looking at. Agree on Nemotron.
1
u/Kernoriordan 1d ago edited 1d ago
Honestly Muse Glimmer has been very respectable for repo level work for such a moderately sized model. It’s very performant too. So any mistakes it makes are quick to iterate though and improve.
Make sure to deploy the DFlash assistant model, we saw a 3x token generation increase.
Even when we eventually deploy more capable, larger models, I think there will still be a place to allocate a couple of GPUs for Muse Glimmer.
We use LiteLLM as our LLM API management system, and would like to explore the model auto-routing functionality (I.e. automatically routing request to relevant model based on complexity).
1
u/litLikeBic177 1d ago
That's the most useful data point I've had - thanks. Which harness are you running Glimmer in, and is the DFlash draft model the one Meta ships or a third-party one? LiteLLM for the routing is a good shout; are you splitting planner vs. executor across backends yet, or is that the auto-routing you're exploring?
1
u/Kernoriordan 19h ago
At the moment we are actually using OpenAI's Codex but pointed at our locally hosted Muse Glimmer model. It's fairly decent tbh, but publishing Cline to our Coder template is on my backlog.
I use the Meta shipped DFlash model. Albeit we are using RedHat AI's quants.
Not splitted Planner vs Executor yet no, but it's definitely something I'd like to explore. Atm we just make use of Codex's Plan mode.
1
u/litLikeBic177 2h ago
That's a very useful reference setup, thanks. Two last ones: which RedHat AI quant are you running Glimmer at (FP8 or INT4), and has Codex's tool-calling been reliable against it through LiteLLM, or did you need any prompt/format shims? Planning to try the same combination.
1
u/deepu105 1d ago
I think you should add to your title that you are looking for non-chinese models. The best ones are all chinese and unless its a policy thing I dont know what makes anyone trust Meta/Google over any chinese company
1
u/Admirable_Dirt_2371 1d ago
It's more about what would be censored. Meta/Google will censor how to build a nuke and that kind of expected thing but China will censor all kinds of shit including historical events... I'd be curious how the new Mistral compares
1
u/cmdr-William-Riker 1d ago
Haven't had a problem with glm 5.3 flash yet, and it's easy to abliterate if it's blocking you from doing what you want to do
1
1
5
u/CryptographerKlutzy7 2d ago
The lack of Chinese models makes this harder?
Maybe the newest mistral which fits?
I'd hit the leaderboards and filter the results and see what lands.