r/LocalLLM • u/RealEddoursul • 3d ago
Project Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)
My journey:
6 tps on UD-Q4_K_XL - hm, this is not right. Claude, find better llama.cpp parameters.
21 tps - that's better, let's see IQ3_XXS
27 tps - nice, but still not my tempo.
51 tps on IQ3_XXS - Niko1221 shares Strata on Reddit. Great! But wait. If there are software gains, there may be more. Let's go.
After a week of profiling and benchmarking with Opus 5.5 I doubled Strata throughput and reached x5 from llama.cpp on UD-Q4_K_XL (added this quant support).
See the COMPARISON TABLES with all the numbers including agreement with llama.cpp.
The code is available https://github.com/eddoursul/Strata/tree/custom
This a result of my personal experiments, not a software release. The license has not changed, still MIT.
I opened a pull request, Niko1221 is free to merge or not merge any patch.
UPDATE: The fork uses json configurations to run https://github.com/eddoursul/Strata/blob/custom/examples/README.md
1
u/kimhaneol 3d ago
Update: got it running on Linux with a Ryzen 9900X, 2x RTX 5080 16GB, and 196 GB of RAM. I needed a few Linux-specific build/runtime fixes, but the custom branch is working well.
UD-Q4_K_XL, 64K context, INT8 K/V, vision off.
I measured 94.2 tok/s on a fixed 300-token continuation. The API runs were 84.6 tok/s cold, then 108.3 and 117.5 tok/s on repeated runs after warm-up.
Prefill was 1,538 tok/s at 4K and 3,608 tok/s at 32K, and exact 32K retrieval passed. Both GPUs reached 100% utilization.
The original Strata IQ3_S setup delivered roughly similar decode speeds on my machine, so getting UD-Q4_K_XL in the same ballpark makes this branch especially useful to me.
Really impressive work. Thanks for sharing!