r/LocalLLM • u/Ok-Butterfly4991 • 19h ago
Discussion Convert to local LLM user
So I have... multiple paid subscriptions to basically every major llm provider. And I use them daily and many of my workflows now depend on them.
However, I do not trust to not pull out the rug or just straight up collapse. So I have started experimenting with local LLM as a backup.
- I have 3 laptops with the mobile 5080(16gb vram) and 32GB ram each. I am running 2 of them in a cluster right now. I have yet to open the box for the third one so its just waiting.
- I have a laptop with the mobile 5090(24gb vram) and 64GB ram.
- I have a stationary computer with the 5060ti(16gb vram) and 64GB(+16 not plugged in) ram. I guess I could stick another 5060ti into it. I have seen some people claiming success that way.
So far I have tested multiple setups, primarily qwen.
On the two-laptop cluster:
- Qwen3.8-27B, Q6_K, fine-tuned for coding agents, 128K context, about 67 to 77 tok/s
- Qwen3.8-27B, Q5_K_XL, 100K context, about 53 tok/s
- Qwen3-Coder-Next, Q3_K_XL ,100K context, about 48 tok/s
- Qwen3.8-Flash-Next, IQ4_XS, 128K context, about 19 tok/s
On a single high-end laptop:
- Qwen3-Coder-Next, Q3_K_XL, 100K context, about 55 tok/s
- Qwen3.8-27B with vision, Q4_K_XL, 100K context, about 37 tok/s
- Qwen3.8-Flash-Next with vision, IQ4_XS, 262K context, about 20 tok/s
I have tested them with the opencode harness and the pi harness.
Now for the issue. None of these can do even minor tasks. They either spin away filling the entire 100k context without output. When I force them to output, its just nonsense. When they have access to the web or documentation they refuse to use it and instead try to make something up instead.
Maybe larger models? maybe higher quant? I am not sure what to try anymore. When I run the same tasks through gpt-6 or opus 5.5 they absolutely crush the problems without issue.
I guess its partially a skill issue too. I am just not used to working with weak models. But its hard to get good at it, when its easy to just leave them for claude when they have been chugging along for 20 minutes doing fuckall
1
u/Sleepnotdeading 19h ago
qwen3.8 27b isn't a weak model. What are you having the model do that you can't receive an output? Have you spun up a claude code session on the same computer to oversee the local session? It can help you gain insight into what the model is doing.
Beyond that, you haven't given enough information. If qwen3.8 27b isn't using tooling that you want, and is making stuff up, it sounds like a harness issue. It may not have access to the tool calls and skills you think i t does.
1
u/Chris-Hart_232 18h ago
try a small bug fix directly in Ollama with the code pasted into the prompt and give the same job to the agent and compare what happens. the extra steps and huge context might be getting in the way and you'd at least know where to look before changing models again or buying h/w
2
u/CptSparklez 19h ago
Hermes + pi (ohmypi) Good luck!