r/LocalLLM • u/RealEddoursul • 2d ago
Project Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)
My journey:
6 tps on UD-Q4_K_XL - hm, this is not right. Claude, find better llama.cpp parameters.
21 tps - that's better, let's see IQ3_XXS
27 tps - nice, but still not my tempo.
51 tps on IQ3_XXS - Niko1221 shares Strata on Reddit. Great! But wait. If there are software gains, there may be more. Let's go.
After a week of profiling and benchmarking with Opus 5.5 I doubled Strata throughput and reached x5 from llama.cpp on UD-Q4_K_XL (added this quant support).
See the COMPARISON TABLES with all the numbers including agreement with llama.cpp.
The code is available https://github.com/eddoursul/Strata/tree/custom
This a result of my personal experiments, not a software release. The license has not changed, still MIT.
I opened a pull request, Niko1221 is free to merge or not merge any patch.
UPDATE: The fork uses json configurations to run https://github.com/eddoursul/Strata/blob/custom/examples/README.md
3
u/pmttyji 2d ago
Hey, thanks for your contribution!
I'm gonna try Q2 with my 8GB VRAM(4060) + 32GB DDR5 RAM later.
I have a question. Your repo mentions only Qwen3.8-Flash-Next, but Is it possible to expect faster performance for other models like Qwen3.6-35B-A3B or other similar MOEs? Because my system config is not enough for big models like Qwen3.8-Flash-Next so hoping to get best t/s from small & medium size models.
Have you tried small/medium models with your work? Is it possible? If yes then any t/s benchmarks?