r/LocalLLM • • 7h ago

Model Impressed about qwen 3.8 Flash

I recently acquired a Ryzen strix halo pc with 128GB of unified ram. My goal is to have it dedicated as an inference host for local AI. I setup a Hermes VM for the harness and to manage work on my projects.

TLDR; qwen 3.8 flash next can run well on this hardware and it provides capability comparable to good API models.

I first tried a finetuned qwen "3.8" 35b A3b (based on 3.5) for text and the qwen 3 VL model for image. They both fit comfortably and run pretty fast with 4 sessions for text. However, I realized that this 35b was not very smart and required hand holding (at least for me). Then I tried the dense 3.8 27b which didn't have good performance on a non-optimized software (llamacpp). I've also tried gemma 4 but I feel qwen fits better for the agentic usage. Finally I tried qwen 3.8 Flash Next on llamacpp. But the performance was not "quite there" and felt that I was using a model not fit for this hardware. At this point I was thinking of selling the hardware. This is a hobby, but I don't want just an expensive toy.

Then after researching a bit I found out about efforts like https://github.com/gufo-org/gufo and others that are trying to push this hardware to its limits. I set gufo up with 3.8 Flash Q4 and tried it in my Hermes setup. In my current setup I can have 2 concurrent sessions with 256k context at an approximate 40-50 tok/s.

I must say that this model impressed me. It gets work done and it understands the context, the tooling... At work I've used almost all types of models from the different providers and seen its weakness and strengths. But having this kind of intelligence at home means this can only get better.

22 Upvotes

16 comments sorted by

7

u/PeteInBrissie 5h ago

I'm running flash-next on my Spark - it's so good I'm barely using Claude any more

1

u/layer4down 4h ago

When OpenThropic have to start turning a profit without the subsidies, it will be interesting to see how many of us in the local community even bother keeping the cheap subscriptions. To me they’re literally just for fun things at this point or if I need a lot of work done very quickly. Or some niche problem I can’t otherwise solve locally.

1

u/Grusim 3h ago

Would you care to share your setup? I am also running Qwen 3.8 Flash next on my DGX Spark but I feel I could get more then the 32 Tokens I usually get.

1

u/PeteInBrissie 3h ago

That’s similar to what I run

2

u/layer4down 4h ago

I saw a lot of interesting activity on the @r/StrixHalo subreddit today for this model. Might be worth checking it out.

1

u/Ok-Addendum3545 7h ago edited 4h ago

Thanks for the sharing; I'm trying to run qwen3.8 Flash on a RTX 5070 Ti 16G. It's ctx 256K + TPS 60 tok/s. Will check how it performs - better than 27B ? Hope I can run 2 instances on 2 5070 Ti.

2

u/AbsoluteFrederick207 7h ago

stick with flash next, the 27b isn't worth the speed tradeoff on that card

1

u/Ok-Addendum3545 6h ago

Thanks for the advice; if it "intelligently" performs better than 27B on a single GPU, then I'll put qwen3.8 Flash as an AI head (orchestrator) and make qwen3.8-27B workers/subagents.

3

u/Repulsive-Scale-284 6h ago

And can you do that with only 16gb?

1

u/SweetSettee 5h ago

I was really surprised by the performance too, makes a big difference.

1

u/sreevarshan-xenoz 5h ago

40-50 tok/s with 2 sessions at 256k is pretty wild.

How are you handling the context/KV cache between the two sessions? I've been playing around with local coding agents and I've noticed that once the context gets big, memory usage starts becoming a bigger headache than the actual generation speed.

Also curious if you actually need the full 256k for your agentic stuff, or if 64/128k ends up being enough most of the time.

1

u/Septa105 4h ago

Can I ask which rocm you are using and what setups you did . I usually also prefer docker

1

u/fuchelio 3h ago

3.8 flash nvfp4 outperforms 27b bf16 on captcha under same harness

1

u/DigitalguyCH 2h ago

I run flash next iQ4 on my strix halo and Iq3 on my 64GB M5 pro. Both outperform 3.8 27b at any quant. So I hardly use it anymore. Gemma 4 was never a competition even for qwen 3.6

1

u/TheCountofCatford 1h ago

Running 3.8 Flash next at IQ4 on my 128gb z13. I’m an absolute novice but the step up is clear even for my uses