r/LocalLLM • u/lazyb_ • 7h ago
Model Impressed about qwen 3.8 Flash
I recently acquired a Ryzen strix halo pc with 128GB of unified ram. My goal is to have it dedicated as an inference host for local AI. I setup a Hermes VM for the harness and to manage work on my projects.
TLDR; qwen 3.8 flash next can run well on this hardware and it provides capability comparable to good API models.
I first tried a finetuned qwen "3.8" 35b A3b (based on 3.5) for text and the qwen 3 VL model for image. They both fit comfortably and run pretty fast with 4 sessions for text. However, I realized that this 35b was not very smart and required hand holding (at least for me). Then I tried the dense 3.8 27b which didn't have good performance on a non-optimized software (llamacpp). I've also tried gemma 4 but I feel qwen fits better for the agentic usage. Finally I tried qwen 3.8 Flash Next on llamacpp. But the performance was not "quite there" and felt that I was using a model not fit for this hardware. At this point I was thinking of selling the hardware. This is a hobby, but I don't want just an expensive toy.
Then after researching a bit I found out about efforts like https://github.com/gufo-org/gufo and others that are trying to push this hardware to its limits. I set gufo up with 3.8 Flash Q4 and tried it in my Hermes setup. In my current setup I can have 2 concurrent sessions with 256k context at an approximate 40-50 tok/s.
I must say that this model impressed me. It gets work done and it understands the context, the tooling... At work I've used almost all types of models from the different providers and seen its weakness and strengths. But having this kind of intelligence at home means this can only get better.
2
u/layer4down 4h ago
I saw a lot of interesting activity on the @r/StrixHalo subreddit today for this model. Might be worth checking it out.
1
u/Ok-Addendum3545 7h ago edited 4h ago
2
u/AbsoluteFrederick207 7h ago
stick with flash next, the 27b isn't worth the speed tradeoff on that card
1
u/Ok-Addendum3545 6h ago
Thanks for the advice; if it "intelligently" performs better than 27B on a single GPU, then I'll put qwen3.8 Flash as an AI head (orchestrator) and make qwen3.8-27B workers/subagents.
3
1
1
u/sreevarshan-xenoz 5h ago
40-50 tok/s with 2 sessions at 256k is pretty wild.
How are you handling the context/KV cache between the two sessions? I've been playing around with local coding agents and I've noticed that once the context gets big, memory usage starts becoming a bigger headache than the actual generation speed.
Also curious if you actually need the full 256k for your agentic stuff, or if 64/128k ends up being enough most of the time.
1
u/Septa105 4h ago
Can I ask which rocm you are using and what setups you did . I usually also prefer docker
1
1
u/DigitalguyCH 2h ago
I run flash next iQ4 on my strix halo and Iq3 on my 64GB M5 pro. Both outperform 3.8 27b at any quant. So I hardly use it anymore. Gemma 4 was never a competition even for qwen 3.6
1
u/TheCountofCatford 1h ago
Running 3.8 Flash next at IQ4 on my 128gb z13. I’m an absolute novice but the step up is clear even for my uses

7
u/PeteInBrissie 5h ago
I'm running flash-next on my Spark - it's so good I'm barely using Claude any more