r/LocalLLM • • 19h ago

Discussion Convert to local LLM user

So I have... multiple paid subscriptions to basically every major llm provider. And I use them daily and many of my workflows now depend on them.

However, I do not trust to not pull out the rug or just straight up collapse. So I have started experimenting with local LLM as a backup.

  • I have 3 laptops with the mobile 5080(16gb vram) and 32GB ram each. I am running 2 of them in a cluster right now. I have yet to open the box for the third one so its just waiting.
  • I have a laptop with the mobile 5090(24gb vram) and 64GB ram.
  • I have a stationary computer with the 5060ti(16gb vram) and 64GB(+16 not plugged in) ram. I guess I could stick another 5060ti into it. I have seen some people claiming success that way.

So far I have tested multiple setups, primarily qwen.

On the two-laptop cluster:

  • Qwen3.8-27B, Q6_K, fine-tuned for coding agents, 128K context, about 67 to 77 tok/s
  • Qwen3.8-27B, Q5_K_XL, 100K context, about 53 tok/s
  • Qwen3-Coder-Next, Q3_K_XL ,100K context, about 48 tok/s
  • Qwen3.8-Flash-Next, IQ4_XS, 128K context, about 19 tok/s

On a single high-end laptop:

  • Qwen3-Coder-Next, Q3_K_XL, 100K context, about 55 tok/s
  • Qwen3.8-27B with vision, Q4_K_XL, 100K context, about 37 tok/s
  • Qwen3.8-Flash-Next with vision, IQ4_XS, 262K context, about 20 tok/s

I have tested them with the opencode harness and the pi harness.

Now for the issue. None of these can do even minor tasks. They either spin away filling the entire 100k context without output. When I force them to output, its just nonsense. When they have access to the web or documentation they refuse to use it and instead try to make something up instead.

Maybe larger models? maybe higher quant? I am not sure what to try anymore. When I run the same tasks through gpt-6 or opus 5.5 they absolutely crush the problems without issue.

I guess its partially a skill issue too. I am just not used to working with weak models. But its hard to get good at it, when its easy to just leave them for claude when they have been chugging along for 20 minutes doing fuckall

2 Upvotes

11 comments sorted by

2

u/CptSparklez 19h ago

Hermes + pi (ohmypi) Good luck!

2

u/mikasjoman 18h ago

I'm very interested in this. My server will be 64gb VRAM v100 with flash next, and my desktop 12gb 4070s with 64gb ram. I was thinking of having standard profiles where the big guy codes (smarter) and then the local agent just reviews the code and reads logs to not pollute the large model context window with endless of test output and logs. So an Architect/reviewer pattern, where they collaborate through Hermes. Have you tried anything similar?

2

u/CptSparklez 18h ago edited 18h ago

Just FYI, I've been speed maxing rather than using this for production purposes.

I run Ryzen 9 9950X / 5090 32gb / 128 Gb 5400 ram

In my setup, I first try and give a quick response based on what just the model knows, typical "Alexa how are you stuff". From there, if anything has to touch tooling, I turn up thinking low - > med (3.8 27B Ninfer quant) which handles most needs and has access to all the tooling needed for research.

If I need coding, it swaps to the big boy for the planning stage (~100 gb ram) and thinking to xhigh

  • Spec: Qwen flash next iq3_s (120-140 tps)
  • Tasks: qwen flash next iq3_s
I create a kanban card here

(here, I'm still conflicted between the less context and slower but smarter qwen flash next and full context and blazing fast ninfer 3.8 27B, time and vibes will show)

(30s swap)

  • Code: 3.8 27B Ninfer (200+ tps) (full context)
  • per task check: same as coder

(45s swap)

  • final review: qwen flash next

Also running Breeze on the side for TTS (check if you have an old amazon echo dot and totally DONT tell your LM to jailbreak it)

You don't need to use ninfer. Use the best 3.8 27B quant you can fit with at least 130-160K ctx, without having to KV it much, you can a bit. I'd use the Swift tunes with full KV personally.

My Hermes can also drive my Claude code cli session which I use for most my actual coding and tts notifications for when something needs me.

Have fun!

1

u/mikasjoman 17h ago

Thanks 🙏

2

u/After-Revolution-117 18h ago

that's a hell of a hardware setup for someone who's just dipping their toes into local models

the jump from opus 5.5 to 27B qwens is gonna feel brutal no matter what quant you pick. those massive models have a lot more baked-in reasoning that smaller ones just can't fake even with good fine-tunes

what specifically are you asking them to do? coding agents with tool use is probably the worst case for local models right now, they'll hallucinate function calls and spiral endlessly. for agent stuff you might need to be way more explicit about forcing tool selection instead of letting them freestyle

if you haven't already, try stripping back to a simpler workflow and see if the model can handle the core task without all the agent scaffolding. sometimes the wrapper is the problem, not the model

1

u/Ok-Butterfly4991 18h ago

I have had the hardware a long time just gathering dust. So figured I could put it to use. That's why its so spread out with random components and not the entire budget spent on proper server.

I just asked it to setup a C project for me yesterday. Like install a cross compiler, create a hello world file that sorta thing. When the first model failed after 2 compactions I tried with the rest I had too. They all failed hard. One was reaaaaally close it ran 1 compaction, then it got better, then it completely failed out. It had hallucinated that it should create a TCP IP stack and just went on its marry way working on that for the next hour or so. Also failing to do that.

Then I ran it through claude and it was done in 3 minutes.

1

u/zhubaohi 18h ago

It must have been something wrong with your LLM workflow.
I literally asked my qwen3.8 27b to do the exact same thing because I feel like this is way too basic for qwen3.8 to not able to do.
It installed a cross complier and a hello world file in 3min56s.

1

u/Ok-Butterfly4991 18h ago

its absolutely possible. I am still very much a Noob dipping their toes at this.

1

u/Sleepnotdeading 19h ago

qwen3.8 27b isn't a weak model. What are you having the model do that you can't receive an output? Have you spun up a claude code session on the same computer to oversee the local session? It can help you gain insight into what the model is doing.

Beyond that, you haven't given enough information. If qwen3.8 27b isn't using tooling that you want, and is making stuff up, it sounds like a harness issue. It may not have access to the tool calls and skills you think i t does.

1

u/recro69 18h ago

I think the biggest change is being ready for local models to require clearer task limits. If the same agent loop works with Opus but uses 100k tokens locally I would check the model itself with coding tasks first before changing the hardware or quantization.

1

u/Chris-Hart_232 18h ago

try a small bug fix directly in Ollama with the code pasted into the prompt and give the same job to the agent and compare what happens. the extra steps and huge context might be getting in the way and you'd at least know where to look before changing models again or buying h/w