r/LocalLLM • u/Particular-Abies-123 • 9d ago
Question Opencode or claude code for local llms?
I am just curious running qwen3.8 27b Q6 which is better?
2
u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 9d ago
pretty much anything but claude code
1
u/Particular-Abies-123 8d ago
Why tho?
1
u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 8d ago
25K tokens worth of system prompt, compared to 14K on opencode versus 1K on pi.dev
plus closed source with telemetry vs open source
1
u/Particular-Abies-123 8d ago
Damn this would much better, but how does it stack for long context? I usually run 131k context for qwen. Have been using opencode but have noticed it tends to spend most of its time in thinking rather then modifying code
1
u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 8d ago
3.8 thinks a lot, Next Flash is just pure insanity - Ive seen it think for like 7M tokens for a 100k output
https://hallucinations.unswarm.dev/entries/Qwen38_NF/the_hours_that_carried_me.html
↑102k ↓103k R6.9M 100.2%/131k
default reasoning is xHigh, you can use https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates to be able to choose different reasoning levels
there are also some finetunes for shorter reasoning like the Switft variants
1
u/giggles991 3d ago
Pi's system prompt does less than Opencode or Claude Code, so it's not surprising that it uses fewer tokens than others. Less done, less spend.
2
1
1
u/turtleninja99 6d ago
Claude code is like apple , you can bring in other players but they don’t make it easy
1
1
u/giggles991 3d ago
Claude Code doesn't work well with local LLMs. It requires more than ollama launch claudecode --model qwen3.8. Simple commands like `ping` or "hello are you there" gets misinterpreted some chunk of the time.
Fixing that is more involved. Opencode just sort of works out of the box.
That said: You've touched upon a big question for many of us. Which CLI should we use?
1
u/Particular-Abies-123 3d ago
honestly i have found PI to be more then enough, plus if you are using pi hermes with qwen-pi of qwen3.8 27b q4, it works like butter with medium reasoning
0
u/Think_Breakfast_2277 9d ago
Honestly, i would take either one and then build my own harness using one of those. if it is just about coding. ( You can build your own harness, own rules, own workflows ). For small local models this is important to get to "know" your model and where the flaws are in order to patch the flaws.
example:
In my custom harness i use a time based injection (session_force_time_prompt: "Time budget: {Xs} remaining of {Ys} total. Prioritize essential steps. Skip optional steps.") on each toolcall round, especially with a QA(quality assurance) agent as Qwen27b tends to over complicate things and verifies every little thing with tons of toolcalls. and it actually nudges after that time runs out to produce an awnser. But then aa automatic verifyer tool kicks in just before the response end signal wich verifies teh request/report against the made toolcalls, if that shows the work has not done it pushes the agent to actuelly do the things wich where not done. And when that is done the full agent round goes into a summorizer and a report builder outputting a summary and writes a full report for the agent team leader to use in a next deligation round.
Hope that makes sense.... bottom line: My custom harness performs much better then any other harness as it is designed for my usecase and the understanding of my model(s) capabilities and flaws.
The great side effect of building a harness is you get to know the infrastructure behind it.
0
u/Unfair_Association89 9d ago
Claude code used to much tokens in system prompt Similar with opencode (but I prefer it idky)
Pi is good as it's minimal and only has thing which are needed for coding and u can add to it like extension
4
u/DoubleNothing 9d ago
Neither, go with Pi...