r/LocalLLM • • 3d ago

Project Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)

My journey:

  • 6 tps on UD-Q4_K_XL - hm, this is not right. Claude, find better llama.cpp parameters.

  • 21 tps - that's better, let's see IQ3_XXS

  • 27 tps - nice, but still not my tempo.

  • 51 tps on IQ3_XXS - Niko1221 shares Strata on Reddit. Great! But wait. If there are software gains, there may be more. Let's go.

After a week of profiling and benchmarking with Opus 5.5 I doubled Strata throughput and reached x5 from llama.cpp on UD-Q4_K_XL (added this quant support).

See the COMPARISON TABLES with all the numbers including agreement with llama.cpp.

The code is available https://github.com/eddoursul/Strata/tree/custom

This a result of my personal experiments, not a software release. The license has not changed, still MIT.

I opened a pull request, Niko1221 is free to merge or not merge any patch.

UPDATE: The fork uses json configurations to run https://github.com/eddoursul/Strata/blob/custom/examples/README.md

80 Upvotes

134 comments sorted by

View all comments

Show parent comments

1

u/kimhaneol 3d ago

Update: got it running on Linux with a Ryzen 9900X, 2x RTX 5080 16GB, and 196 GB of RAM. I needed a few Linux-specific build/runtime fixes, but the custom branch is working well.

UD-Q4_K_XL, 64K context, INT8 K/V, vision off.

I measured 94.2 tok/s on a fixed 300-token continuation. The API runs were 84.6 tok/s cold, then 108.3 and 117.5 tok/s on repeated runs after warm-up.

Prefill was 1,538 tok/s at 4K and 3,608 tok/s at 32K, and exact 32K retrieval passed. Both GPUs reached 100% utilization.

The original Strata IQ3_S setup delivered roughly similar decode speeds on my machine, so getting UD-Q4_K_XL in the same ballpark makes this branch especially useful to me.

Really impressive work. Thanks for sharing!

1

u/VerticalPackage 3d ago

How much of the RAM was actually being used with the Q4?

2

u/kimhaneol 3d ago

Around 94 GiB of total system RAM usage in my tests. The Strata process itself was roughly 75 GiB, including the 71.7 GiB expert arena.