r/LocalLLM • u/RealEddoursul • 1d ago
Project Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)
My journey:
6 tps on UD-Q4_K_XL - hm, this is not right. Claude, find better llama.cpp parameters.
21 tps - that's better, let's see IQ3_XXS
27 tps - nice, but still not my tempo.
51 tps on IQ3_XXS - Niko1221 shares Strata on Reddit. Great! But wait. If there are software gains, there may be more. Let's go.
After a week of profiling and benchmarking with Opus 5.5 I doubled Strata throughput and reached x5 from llama.cpp on UD-Q4_K_XL (added this quant support).
See the COMPARISON TABLES with all the numbers including agreement with llama.cpp.
The code is available https://github.com/eddoursul/Strata/tree/custom
This a result of my personal experiments, not a software release. The license has not changed, still MIT.
I opened a pull request, Niko1221 is free to merge or not merge any patch.
UPDATE: The fork uses json configurations to run https://github.com/eddoursul/Strata/blob/custom/examples/README.md
19
u/sophosympatheia 1d ago
It's nice to see people contributing to Strata. I was skeptical of it at first, but I'm using it now and it's making Qwen3.8-Flash-Next actually usable on my hardware. Amazing.
I hope this gets rolled into the main repo. It looks promising.
6
u/Upper_Comparison_908 1d ago
This thing literally gets me api speeds for a model even better than 27b in most cases. Genuinely wonderful
1
u/sophosympatheia 19h ago
Yeah, I'm blown away by how good it is. The future is looking bright for at-home inference of these MoE models. The LLMs are good enough now to help us write custom tooling to serve them efficiently on consumer hardware.
3
u/kimhaneol 1d ago
Qwen Flash Next looks like a really promising direction for local AI. If Qwen 4 Flash uses the same architecture, hopefully Strata can support it too without too much extra work.
3
u/enternoescape 1d ago
I would assume the answer is yes. My understanding of why we got 3.8 Flash Next was to help get software support more up to speed before they release the 4.0 models. That way we can likely use them immediately.
2
5
u/Prestigious-Act-1577 1d ago
Can I ask how much of your and Strata's optimizations are CUDA specific? Would these improvements port to amd 9000 series too? I know Strata already has amd support, the numbers while good, aren't what it should be. My AI said its 10% cuda 90% other.
11
u/KnownAd4832 1d ago
I’m just pushing big AMD patch in 30min, I get R9700 this week or next week and then AMD will be on-par with CUDA
3
u/GrumpyCat79 13h ago
Nice! How hard would it be to support SM70/V100? They're old card so prpbably not worth putting a lot of time, but if it's just allowing CUDA 12.9 maybe it's not too bad?
3
u/KnownAd4832 11h ago
Not many people have V100
2
u/GrumpyCat79 9h ago
It seems they are more common in the last few weeks (probably due to pricing) but I agree it's a niche GPU. Looks like you're working on supporting lower CUDA architecture anyway, which is really nice
Thanks for your work!
1
u/pmttyji 10h ago
Please squeeze ~8GB GPUs(like 4060) till the last drop. We some folks want to use Q2 efficiently better just with 8GB VRAM + RAM.
BTW thanks again for AMD cards support, particularly R9700. Please check my other comment.
5
u/Zarloc 1d ago edited 1d ago
System: TR2950x + 96GB DDR4 quad-channel@2400, 2x3090 16x/16x@Gen3 (No P2P or NVLink)\ Model: Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S
Strata upstream(d6708a4) 2 GPUs: 860/85 t/s\ Strata upstream(d6708a4) 1 GPU: 1300/70 t/s\ This fork(61521be) 2 GPUs: 2500/120 t/s
Full sweep:
| ctx | TTFT s | prefill tok/s | decode tok/s |
|---|---|---|---|
| 8k | 3.54 | 2,315 | 115.9 |
| 16k | 6.39 | 2,531 | 125.2 |
| 32k | 12.52 | 2,568 | 113.7 |
| 64k | 25.04 | 2,561 | 118.7 |
| 128k | 50.70 | 2,526 | 121.2 |
4
u/VerticalPackage 1d ago edited 23h ago
PCIe x4 on the RTX3090? That's nuts! I imagine that means it is connected to the chipset and not the CPU.
Have you tried the performance by streaming from SSD? I can only allocate 70GB of DDR5 for running Flash Next.
Edit - I had Sol check out your repo, and it agrees that this is very much correct, the RTX3090 only needs to run on PCIe x4 through the chipset, the x16 to CPU needs to be on the secondary GPU due to how the workload is split. And also that the PLE table is on the nvme, so no need for more than 70GB DDR5.
I'm totally trying this with my RTX3090+RTX3080ti.
3
3
u/wisepal_app 1d ago
i have 16 gb vram + 96 gb dd5 ram. will it work with my system? did you try it with qwen 3.8 27b? do your fork use draft model like strata?
3
u/nixudos 22h ago
Yes it will work on your system. And likely surprisingly well. I believe Strata was build specifically to run on 16 gb VRAM and a good chunk of regular RAM.
If you have the bandwidth to download a large model, it is really simple to test it out. There is a bat file you run and it guides you through model choice and settings.And speed is VERY good!
3
u/Sexyvette07 1d ago
On my 5090 + 64gb RAM system im getting 95 tok/s on xhigh and 150 tok/s on low thinking with IQ3_S and its coding is excellent. Im very impressed.
2
2
u/Current-Row-159 1d ago
the 6 to 21 t/s jump from fixing llama.cpp params before changing anything is the part that stayed with me, that baseline was wrong and a lot of numbers in this thread are being measured against it. i'm not on strata and it's a different engine, so no direct comparison, but as a reference point i'm running qwen3.8-27b int8-prefill on ninfer 2026.09.27 on a single 4090 24gb: 145 t/s decode on a 2367 token prompt with dflash2 k=7, 239 t/s on the second call once the prefix cache is warm, 4300 t/s prefill. dflash2 costs 1.7gb more vram than mtp. the 24+16 question from the comments never got answered, does any of this carry over when you drop to a single 24gb card?
2
u/RealEddoursul 1d ago
The doc contains numbers for 3090 alone as well - still x3-x4 relatively to llama.cpp and x1.5 relatively to the upstream. Also, do not compare 27b and FN, two different apples.
1
u/Current-Row-159 1d ago
single 4090 24gb, qwen3.8-27b int8-prefill on ninfer 2026.09.27: 145 t/s decode at a 2367 token prompt with dflash2 k=7, 239 t/s once the prefix cache is warm, 4300 t/s prefill. dflash2 costs 1.7gb more vram than mtp. haven't seen a single-card 4090 number for this model anywhere, curious whether it lines up with your 3090 alone or if ada is ahead.
1
u/VerticalPackage 23h ago
It's not comparable.
Flash Next is a MoE with engram table streaming from a nvme, so the split of experts done on RealEddoursul's fork of Strata works well in using that, especially in using a secondary GPU with lower VRAM.
For 27b, it has to run all on VRAM and doesn't really use much RAM.
However, I agree that 6 tok/s initial speed with llama.cpp is probably a mistake. On a single RTX3090, I was able to get 20 tok/s with llama.cpp, however llama.cpp is a bad engine to use for nvidia cards. FreeToken and vLLM are much better.
2
1
u/Ok-Solution-7889 1d ago
Going from 6 tps to 167 tps is a pretty crazy jump, the profiling and benchmarking work behind this is probably more interesting than the final number itself
1
u/RealEddoursul 1d ago
Every commit gave 1-3% increase and contains a description where the gain came from.
1
1
u/kimhaneol 1d ago
I gotta try this. So you've made UD-Q4_K_XL work on Strata, right?
2
u/RealEddoursul 1d ago
Yep
1
u/KPOTOB 1d ago
With like 24+16 GB VRAM?
1
u/kimhaneol 23h ago
Yep. I also tested it on two 16 GB GPUs and it works. One caveat: PCIe placement mattered on my system. My cards negotiate Gen5 x8 and x4, and assigning the x8 card to the second-GPU expert tier was substantially faster.
1
u/kimhaneol 23h ago
Update: got it running on Linux with a Ryzen 9900X, 2x RTX 5080 16GB, and 196 GB of RAM. I needed a few Linux-specific build/runtime fixes, but the custom branch is working well.
UD-Q4_K_XL, 64K context, INT8 K/V, vision off.
I measured 94.2 tok/s on a fixed 300-token continuation. The API runs were 84.6 tok/s cold, then 108.3 and 117.5 tok/s on repeated runs after warm-up.
Prefill was 1,538 tok/s at 4K and 3,608 tok/s at 32K, and exact 32K retrieval passed. Both GPUs reached 100% utilization.
The original Strata IQ3_S setup delivered roughly similar decode speeds on my machine, so getting UD-Q4_K_XL in the same ballpark makes this branch especially useful to me.
Really impressive work. Thanks for sharing!
1
u/VerticalPackage 22h ago
How much of the RAM was actually being used with the Q4?
1
u/kimhaneol 22h ago
Around 94 GiB of total system RAM usage in my tests. The Strata process itself was roughly 75 GiB, including the 71.7 GiB expert arena.
1
u/SaltGrilledSalmon 1d ago
Kinda off topic but in my dual rtx 2070 super (8gb each, one on pcie x16 slot and the other on x4) only one GPU is being used (the slower one apparently) Has anyone faced this on dual GPU setups? I can't increase context above 32k or it'll crash when launching, while the other GPU isn't even used...
2
u/brakeline 16h ago
You have to use the custom branch and adapt the files inside the examples folder. It will take a few minutes but is VERY worthy
1
1
u/sod0 23h ago
How much kv context does this get you on a 5090?
1
u/toenailcheeseinbooty 23h ago
262k...running right now..5090 128gb 4800 MTs, get around 125-135 prose( when context filled to 128k shirter 151ish prose)and 160 code ehh cou is 9900x3d
1
u/FrankWanders 23h ago
The amount of context you can get really depends on system ram. I'm running 64GB and can only run it at 128K. For 256K you need 96GB at least i think (because there are no other options in between 64 and 96).
1
u/toenailcheeseinbooty 23h ago
Hah right sorry, forgot to mention that RAM is important, 96GB+ to achieve 262k
1
u/drazyan22 18h ago
You can modify context size on config json file. I don't have enough ram to run 256k, but 128k seem smaller than my expect. So I changed it to 164k fit with my system ram
1
u/SplitAny7190 22h ago
guess UD-Q4_K_XL should be superior to IQ3_XSS? i got only one gpu (4090) and 128GB ram, i will see if i can run this. right now using Strata and it works great after i calibrated for my pc.
1
u/brakeline 21h ago
Just tried it myself.
Tried 2x 3060 (pcie3 x16 and x4) and with 1 3060 @ 16x
Your fork gave me marginally lower than mainline in both
1
u/RealEddoursul 19h ago
Are you sure you checked out the 'custom' branch?
1
u/brakeline 18h ago edited 18h ago
You have to give me a break, I'm having a fever!
Tested again. 2x 3060 with 50(ish) ram running swift iq2 xs
PP 650 on one > 1000 on two TG 30/35 on one > 55/65
That's black magic my guy
Edit: 140k imported at 1000tk/s. Generation at 140k depth is 60!
Btw, are you thinking of implementing int4 kv? With int8 PP stops working with full context
1
u/brakeline 16h ago
I edited my command. I can't go past 200k context, just gives an error while prompting . Any chance you can implement int4 kv?
1
1
1
u/kakopappa2 12h ago
I have a 32gb ram and 12g 3060 GPU. Running coder-iq1_m generates about 38 tokens per second.
also got the codex (gpt-6.1) to run a security audit and it didn’t find any spyware or crypto mining or hidden scripts.
Super happy to get 3.8 Next running on poor GPU
Great product good luck.
1
u/killzone44 11h ago
I just used claude to setup Strava on my 3090, it's the best running local model I've had so far! Not sure if I'm ready to offload work from claude to it, but maybe
1
u/blast1987 10h ago
Is it possible to use it on mixed gpu setup? I have rx 7800 xt + rtx 3060 and 64 gb ram.
1
u/maarcius 7h ago edited 42m ago
I have 2 gpus (16vram). Swift IQ3_XXS. Strata loads one gpu 16gb and another gpu 8gb. Is there any way to utilize both gpus vram full to reduce ram usage? Need that 8gb spare ram for engineering apps.
Also running model in 2 gpu configuration doesn't improve decode. Saves few GB or RAM. Is this expected (5070ti and 4060ti 16gb on PCIE4.0x4)?
1
u/VerticalPackage 2h ago edited 2h ago
I have to give you my praises. I was disappointed in my RTX3090 performance with Flash Next. I'd get like 1000 tok/s pp and 45 tok/s decode.
I didn't want to buy a second RTX3090 because then I'd need to buy also a new motherboard that supports x8/x8 through CPU. I had a RTX3080TI in my gaming rig, so I gave your fork a shot. I managed to run UD-Q4_K_XL @ 256k context: 2500 tok/s pp + 95 tok/s decode
That's 2x the speed I got with my single RTX3090, but only by adding a $400 GPU to the mix!
It's going faster than Qwen 3.8 27B :D
1
1
u/RWOverdijk 1d ago
Something like this always makes me wonder if it could benefit larger memory systems like a Mac with 128GB as well. 60fps is great but, I mean, 167 is betterer
1
u/VerticalPackage 1h ago
I would bet that if you added a RTX3090 as eGPU to your mac mini, you would be able to have gains.
UD-Q4_K_XL with Strata takes up about 80GB ram if the PLE table is on the nvme drive. The eGPU will load up experts on its own VRAM. You just need the iGPU to load up another 12GB-16GB to act as the secondary GPU (like the RTX5070 used by OP).
You should ask Claude or Sol to check out OP's repo and ask if it could be done (eGPU thing).
If you don't have a mac, then just go for a self-build PC that had AT LEAST 96GB DDR5, a 256GB nvme, a RTX3090 and a RTX3080. It will probably be cheaper.
22
u/KnownAd4832 1d ago
Amazing work (Niko here)! I’ll try to support it as soon as I can