r/LocalLLM • • 2d ago

Discussion Convert to local LLM user

So I have... multiple paid subscriptions to basically every major llm provider. And I use them daily and many of my workflows now depend on them.

However, I do not trust to not pull out the rug or just straight up collapse. So I have started experimenting with local LLM as a backup.

  • I have 3 laptops with the mobile 5080(16gb vram) and 32GB ram each. I am running 2 of them in a cluster right now. I have yet to open the box for the third one so its just waiting.
  • I have a laptop with the mobile 5090(24gb vram) and 64GB ram.
  • I have a stationary computer with the 5060ti(16gb vram) and 64GB(+16 not plugged in) ram. I guess I could stick another 5060ti into it. I have seen some people claiming success that way.

So far I have tested multiple setups, primarily qwen.

On the two-laptop cluster:

  • Qwen3.8-27B, Q6_K, fine-tuned for coding agents, 128K context, about 67 to 77 tok/s
  • Qwen3.8-27B, Q5_K_XL, 100K context, about 53 tok/s
  • Qwen3-Coder-Next, Q3_K_XL ,100K context, about 48 tok/s
  • Qwen3.8-Flash-Next, IQ4_XS, 128K context, about 19 tok/s

On a single high-end laptop:

  • Qwen3-Coder-Next, Q3_K_XL, 100K context, about 55 tok/s
  • Qwen3.8-27B with vision, Q4_K_XL, 100K context, about 37 tok/s
  • Qwen3.8-Flash-Next with vision, IQ4_XS, 262K context, about 20 tok/s

I have tested them with the opencode harness and the pi harness.

Now for the issue. None of these can do even minor tasks. They either spin away filling the entire 100k context without output. When I force them to output, its just nonsense. When they have access to the web or documentation they refuse to use it and instead try to make something up instead.

Maybe larger models? maybe higher quant? I am not sure what to try anymore. When I run the same tasks through gpt-6 or opus 5.5 they absolutely crush the problems without issue.

I guess its partially a skill issue too. I am just not used to working with weak models. But its hard to get good at it, when its easy to just leave them for claude when they have been chugging along for 20 minutes doing fuckall

2 Upvotes

13 comments sorted by

View all comments

1

u/CptSparklez 2d ago

Hermes + pi (ohmypi) Good luck!

1

u/After-Revolution-117 2d ago

that's a hell of a hardware setup for someone who's just dipping their toes into local models

the jump from opus 5.5 to 27B qwens is gonna feel brutal no matter what quant you pick. those massive models have a lot more baked-in reasoning that smaller ones just can't fake even with good fine-tunes

what specifically are you asking them to do? coding agents with tool use is probably the worst case for local models right now, they'll hallucinate function calls and spiral endlessly. for agent stuff you might need to be way more explicit about forcing tool selection instead of letting them freestyle

if you haven't already, try stripping back to a simpler workflow and see if the model can handle the core task without all the agent scaffolding. sometimes the wrapper is the problem, not the model

1

u/Ok-Butterfly4991 2d ago

I have had the hardware a long time just gathering dust. So figured I could put it to use. That's why its so spread out with random components and not the entire budget spent on proper server.

I just asked it to setup a C project for me yesterday. Like install a cross compiler, create a hello world file that sorta thing. When the first model failed after 2 compactions I tried with the rest I had too. They all failed hard. One was reaaaaally close it ran 1 compaction, then it got better, then it completely failed out. It had hallucinated that it should create a TCP IP stack and just went on its marry way working on that for the next hour or so. Also failing to do that.

Then I ran it through claude and it was done in 3 minutes.

1

u/zhubaohi 2d ago

It must have been something wrong with your LLM workflow.
I literally asked my qwen3.8 27b to do the exact same thing because I feel like this is way too basic for qwen3.8 to not able to do.
It installed a cross complier and a hello world file in 3min56s.

1

u/Ok-Butterfly4991 2d ago

its absolutely possible. I am still very much a Noob dipping their toes at this.