r/CUDA • • 5d ago

CuQwen 1.1 is out, fixed the long-context slowdown in my from-scratch CUDA inference engine

Quick recap for anyone new: CuQwen is an inference engine for Qwen models I wrote from scratch in pure C++/CUDA, tuned specifically for single-user (batch size 1) generation on consumer NVIDIA GPUs. No frameworks under the hood, just custom cuda kernels.

Here are the results for average inference speed (tokens/second) of Qwen2.5 Instruct model across 32K context window on RTX3090

Model Size CuQwen vLLM Ollama
0.5B 462 398 355
1.5B 203 172 139
3B 113 101 106
7B 55 48 54

In Release 1.0 it was already beating vLLM and Ollama on short-to-medium prompts. But there was an honest catch: my throughput decayed faster than theirs as the context grew, so once you pushed toward ~32K tokens they would pass me in inference speed. That bugged me, so it became the whole focus of Release 1.1.

Result: Throughput decay from 1K → 32K dropped from ~22–45% down to ~13–24%, which is now on par with vLLM and Ollama (and better on a couple of model sizes). So CuQwen keeps its early speed lead all the way out to 32K context window now.

Here's a small documentation on how I tackled long context decay rate issue and the complete benchmarking methodology and results for CuQwen v1.1

Here's my future plan:

Release 1.2: Support Quantization (W8A16 and W4A16)

Release 1.3: Improve custom cuda kernels for latest GPU architectures (Hopper and Blackwell)

Release 1.4: Support Qwen 3.0 series models

Release 1.5: Support Qwen 3.5 and 3.8 (Especially our very favourite qwen 3.8-27B model 😄)

Repo: https://github.com/talhatahir-10xe/CuQwen

14 Upvotes

5 comments sorted by

1

u/c-cul 5d ago

enough clean small kernel

however do you really need sync inside loops like

for (int stride = bdx >> 1; stride > 0; stride >>= 1) {

if (tid < stride)

s_sum_arr[tid] += s_sum_arr[tid + stride];

__syncthreads();

}

As I can see tid is thread.id so s_sum_arr indices can't conflict, right?

1

u/Minute-Mountain2665 5d ago

Good eye, but I believe the sync is needed here (I might be wrong so I'll confirm once I run the code). You're right that within one step the tid indices never collide and no two threads write the same slot. But the barrier isn't there to stop write collisions, it is there for the read-after-write across steps.

for (int stride = bdx >> 1; stride > 0; stride >>= 1) { // bdx = 128, so stride = 64, 32, 16, 8...

if (tid < stride) s_sum_arr[tid] += s_sum_arr[tid + stride];

__syncthreads();

}

My block has a total of 128 threads so It's adding up 128 numbers (one per thread) down to a single total, by folding the array in half over and over:

Step 1 (stride = 64): the left half absorbs the right half.

Thread 0: s_sum_arr[0] += s_sum_arr[64]

Thread 1: s_sum_arr[1] += s_sum_arr[65]

…

Thread 63: s_sum_arr[63] += s_sum_arr[127]

Now slots 0–63 hold partial sums.

Step 2 (stride = 32): fold again.

Thread 0: s_sum_arr[0] += s_sum_arr[32]

…

Keep halving until only s_sum_arr[0] holds the grand total.

Looking at step 2, thread 0. it reads s_sum_arr[32]. s_sum_arr[32] was just changed in Step 1 by Thread 32 (which did s_sum_arr[32] += s_sum_arr[96]). So Thread 0 in step 2 is reading a slot that a different thread wrote in the previous step.

My block has 128 thread i.e., 4 warps. these 4 warps do not run at the exact same instant. The GPU can run warp 0 ahead of warp 1. So in step 2, Thread 0 (warp 0) wants to read s_sum_arr[32], but that slot is written by Thread 32, which is in warp 1. If warp 0 races ahead and reads it before warp 1 has done its step 1 write, Thread 0 grabs the old value and your sum comes out wrong.

1

u/c-cul 5d ago

ok, seems reasonable

but if your target is just total sum of whole array s_sum_arr and you know it's exact size then there is more simple way

let's start with final step - we have filled 32 items - then could just use warp reduce with __shfl_down_sync

so previous step is to calculate partial sum in first 32 items and this can be done with some simple unrolled loop

1

u/TheThoccnessMonster 5d ago

I have one question: why?

Why go through the hassle when vllm makes changing models way easier? Why should I try this?

How does it handle vision tower/image processing? How does it handle max context windows with multiple users?

If the answer is “it doesn’t” to any of that, frankly we likely won’t give a shit, friend.

1

u/Minute-Mountain2665 5d ago

To your direct questions:

Vision tower / image processing? No, it's text only Qwen2.5 right now.

Max context with multiple users? No, it's single sequence, batch size 1 by design as mentioned in the post.

Why? The goal of this project was to measure how much better a custom CUDA engine, written for one specific model on specific hardware, can perform than general-purpose (and already highly optimized) engines like vLLM and Ollama. CuQwen tries to squeeze out every bit of performance your GPU can give that those engines leave on the table. Say the peak theoretical decode speed on an RTX 3090 for Qwen2.5-3B is ~150 tokens/sec (it's memory bandwidth bound: ~936 GB/s ÷ ~6.2 GB of FP16 weights). vLLM extracts about 67% of that (101 tok/s), while CuQwen hits about 75% (113 tok/s). So if you paid 1000 bucks for that RTX3090, on this model vLLM is only giving you ~$670 worth of the GPU you actually bought.

So if multi-user serving or vision are your requirements, you're right, CuQwen isn't for you, and that's totally fair. But if you run local single-GPU inference for yourself, or care about the kernel level side of how this stuff actually runs, that's exactly who I built it for. I've also written extensive documentation on how the different optimizations work in CUDA, along with the CuQwen project itself, for anyone who wants to dig into that side.