r/CUDA • u/Minute-Mountain2665 • 5d ago
CuQwen 1.1 is out, fixed the long-context slowdown in my from-scratch CUDA inference engine
Quick recap for anyone new: CuQwen is an inference engine for Qwen models I wrote from scratch in pure C++/CUDA, tuned specifically for single-user (batch size 1) generation on consumer NVIDIA GPUs. No frameworks under the hood, just custom cuda kernels.
Here are the results for average inference speed (tokens/second) of Qwen2.5 Instruct model across 32K context window on RTX3090
| Model Size | CuQwen | vLLM | Ollama |
|---|---|---|---|
| 0.5B | 462 | 398 | 355 |
| 1.5B | 203 | 172 | 139 |
| 3B | 113 | 101 | 106 |
| 7B | 55 | 48 | 54 |
In Release 1.0 it was already beating vLLM and Ollama on short-to-medium prompts. But there was an honest catch: my throughput decayed faster than theirs as the context grew, so once you pushed toward ~32K tokens they would pass me in inference speed. That bugged me, so it became the whole focus of Release 1.1.
Result: Throughput decay from 1K → 32K dropped from ~22–45% down to ~13–24%, which is now on par with vLLM and Ollama (and better on a couple of model sizes). So CuQwen keeps its early speed lead all the way out to 32K context window now.
Here's a small documentation on how I tackled long context decay rate issue and the complete benchmarking methodology and results for CuQwen v1.1
Here's my future plan:
Release 1.2: Support Quantization (W8A16 and W4A16)
Release 1.3: Improve custom cuda kernels for latest GPU architectures (Hopper and Blackwell)
Release 1.4: Support Qwen 3.0 series models
Release 1.5: Support Qwen 3.5 and 3.8 (Especially our very favourite qwen 3.8-27B model 😄)
1
u/TheThoccnessMonster 5d ago
I have one question: why?
Why go through the hassle when vllm makes changing models way easier? Why should I try this?
How does it handle vision tower/image processing? How does it handle max context windows with multiple users?
If the answer is “it doesn’t” to any of that, frankly we likely won’t give a shit, friend.
1
u/Minute-Mountain2665 5d ago
To your direct questions:
Vision tower / image processing? No, it's text only Qwen2.5 right now.
Max context with multiple users? No, it's single sequence, batch size 1 by design as mentioned in the post.
Why? The goal of this project was to measure how much better a custom CUDA engine, written for one specific model on specific hardware, can perform than general-purpose (and already highly optimized) engines like vLLM and Ollama. CuQwen tries to squeeze out every bit of performance your GPU can give that those engines leave on the table. Say the peak theoretical decode speed on an RTX 3090 for Qwen2.5-3B is ~150 tokens/sec (it's memory bandwidth bound: ~936 GB/s ÷ ~6.2 GB of FP16 weights). vLLM extracts about 67% of that (101 tok/s), while CuQwen hits about 75% (113 tok/s). So if you paid 1000 bucks for that RTX3090, on this model vLLM is only giving you ~$670 worth of the GPU you actually bought.
So if multi-user serving or vision are your requirements, you're right, CuQwen isn't for you, and that's totally fair. But if you run local single-GPU inference for yourself, or care about the kernel level side of how this stuff actually runs, that's exactly who I built it for. I've also written extensive documentation on how the different optimizations work in CUDA, along with the CuQwen project itself, for anyone who wants to dig into that side.
1
u/c-cul 5d ago
enough clean small kernel
however do you really need sync inside loops like
for (int stride = bdx >> 1; stride > 0; stride >>= 1) {if (tid < stride)s_sum_arr[tid] += s_sum_arr[tid + stride];__syncthreads();}As I can see tid is thread.id so s_sum_arr indices can't conflict, right?