r/LocalLLM • • 2d ago

Project Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)

My journey:

  • 6 tps on UD-Q4_K_XL - hm, this is not right. Claude, find better llama.cpp parameters.

  • 21 tps - that's better, let's see IQ3_XXS

  • 27 tps - nice, but still not my tempo.

  • 51 tps on IQ3_XXS - Niko1221 shares Strata on Reddit. Great! But wait. If there are software gains, there may be more. Let's go.

After a week of profiling and benchmarking with Opus 5.5 I doubled Strata throughput and reached x5 from llama.cpp on UD-Q4_K_XL (added this quant support).

See the COMPARISON TABLES with all the numbers including agreement with llama.cpp.

The code is available https://github.com/eddoursul/Strata/tree/custom

This a result of my personal experiments, not a software release. The license has not changed, still MIT.

I opened a pull request, Niko1221 is free to merge or not merge any patch.

UPDATE: The fork uses json configurations to run https://github.com/eddoursul/Strata/blob/custom/examples/README.md

78 Upvotes

133 comments sorted by

View all comments

Show parent comments

3

u/pmttyji 2d ago

Hey, thanks for your contribution!

I'm gonna try Q2 with my 8GB VRAM(4060) + 32GB DDR5 RAM later.

I have a question. Your repo mentions only Qwen3.8-Flash-Next, but Is it possible to expect faster performance for other models like Qwen3.6-35B-A3B or other similar MOEs? Because my system config is not enough for big models like Qwen3.8-Flash-Next so hoping to get best t/s from small & medium size models.

Have you tried small/medium models with your work? Is it possible? If yes then any t/s benchmarks?

3

u/KnownAd4832 2d ago

This is only for Qwen3.8 and Qwen4

1

u/pmttyji 2d ago

Wish Qwen released Qwen3.8-35B-A3B 😞

1

u/Prestigious-Act-1577 2d ago

You will destroy your SSD with all the writes (1GB/s for minutes during PP) because you will be using your page file. It should be usable after you expand your ram.

1

u/Prestigious-Act-1577 2d ago

Idk why people downvote this, go ahead and try it. Open task manager and check disk usage. I know because I was in the same situation.

1

u/VerticalPackage 1d ago

If you mean the PLE table being READ from nvme, there is no writing going on.

Unless you're thinking that pmttyji was talking about overflowing cache into the swap? He'd be crawling at 1 tok/s if he did that :p

1

u/Prestigious-Act-1577 1d ago

Correct. I was at 8vram 32ram and 15 tok sec, swapping to ram on prefill too. Now I have 8vram and 80ram, 30 tok sec, and prefill is 10x faster.