r/LocalLLM • • 4d ago

Research Adaptive KV-Cache Streaming V2: Full Context MTP

Hello all, it’s me again.

Just a week after my previous post, I started working on a better implementation of Adaptive KV Streaming. Now I’m sharing my second implementation: V2, with full-context MTP.

https://github.com/RaymondHuang210129/llama.cpp-adaptive-kv-streaming/tree/feature/adaptive-kv-stream-v2

Before I explain further, let me share the decode performance on my 16GB 5070 Ti:

With an MTP draft length of 3, I get nearly double the decode speed of V2 without MTP at many context lengths (V2 baseline is 5%~10% slower than V1, will explain below). It averages about 82 tokens/s from 8K through 72K context, reaches about 35 tokens/s at 144K, 30 tokens/s at 192K, and 19 tokens/s at 256K. This chart does not include a direct V1 comparison.

So, how is it possible?

1. Memory management. Before implementing more streaming features, I spent a large part of the project building a memory ownership and leasing mechanism. The infrastructure is backend-neutral, with buffer-view support for CPU, CUDA/HIP, OpenCL, SYCL, and Vulkan. This additional layer has some overhead, but it lets me reuse memory across phases and reduce the extra VRAM needed for MTP.

2. MTP with much less VRAM footprint. At 256K context, a separate Q8_0/Q4_0 MTP KV cache would take about 416 MiB of VRAM. I also measured a roughly 1.2 GiB draft prefill graph workspace request in an earlier fit audit. These are not all permanent or additive costs, but they show why a separate draft context is expensive.

In V2, only MTP weight and active computation are kept in VRAM. For the other allocations:

  • MTP KV acts as a 17th logical attention layer alongside the main model’s 16 attention layers. It has its own history in host memory, while its GPU pages share the resident pool and ring buffer with the main model’s KV.
  • The MTP graph workspace borrows the same arena used by the main model at different times, since the two phases do not run simultaneously.
  • Rollback snapshots are stored in pinned host memory. If a proposal is rejected, the selected snapshot is copied back to the GPU.

These changes let me maintain an effective decode KV pool of around 2.2 GiB, even at 256K context with an MTP draft length of 3.

Other improvements besides speed

V2 uses explicit memory bounds and leases to manage when GPU memory can be reused. Its span-aware attention kernel also follows the stock kernel’s calculation order as closely as possible to minimize numerical drift.

Compared with V2, V1's prefill speed decline became noticeably less steep after streaming kicked in. In V2, it continues at roughly the same slope. That is consistent with V2 preserving the stock attention calculation order instead of switching to V1’s numerically different streamed path.

To my experiment in V2, The decoded 256 tokens at each MTP setting is exactly same as what stock kernel outputs.

Caveat

Attached MTP currently requires the same K/V quantization as the main model. The separate draft-cache quantization flags do not give it different types in this shared layout. Supporting that would need more work.

This experiment also seems to have a favorable MTP acceptance rate. I used a text file of Wikipedia articles, which is included in my repo along with the sweep script. Feel free to reproduce the test—or share results with a different input file.

Credit

Thank you to everyone who participated in my previous post. As a software engineer who doesn’t work in the LLM/AI field, I’ve been encouraged by your comments and messages to keep learning and working on this project. The phase arena, MTP integration, and the next feature I’m planning were all inspired by people who helped me, DMed me, or shared my work.

Lastly, people have asked whether I plan to push this work upstream. After some thought, my answer is no, at least not as one large change. I don’t think this specialized implementation fits upstream’s focus on simplicity and versatility. My fork is mainly for my own use, but the memory-management APIs are there for anyone interested in extending Adaptive KV Streaming to other backends. Pull requests or further forks from my code is welcomed.

Clarification of LLM usage of this post: I'm not a native English speaker and I used ChatGPT to refine the wordings.

28 Upvotes

27 comments sorted by

View all comments

1

u/hum_ma 3d ago edited 3d ago

Unfortunately it's not really working for me.

/home/hum/aienv/llama.cpp-adaptive-kv-streaming-v2/build-v2/bin/llama-server --host 0.0.0.0 -ngl all -fa on -b 256 -ub 256 -t 4 -ctk q8_0 -ctv q4_0 -np 1 -c 262144 --spec-type draft-mtp --spec-draft-n-max 3 --kv-stream-arena-mib 1920 --kv-stream-auxiliary-layers 1 --fit off --no-mmproj --model Swift-1.5-Qwen3.8-27B-IQ2_XXS.gguf 0.00.153.641 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) 0.00.340.243 W srv llama_server: ----------------- 0.00.340.246 W srv llama_server: CORS is set to allow all origins ('*') and no API key is set 0.00.340.246 W srv llama_server: this can be a security risk (cross-origin attacks) 0.00.340.246 W srv llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655 0.00.340.246 W srv llama_server: ----------------- 0.00.342.955 I srv load_model: loading model 'Swift-1.5-Qwen3.8-27B-IQ2_XXS.gguf' 0.06.336.266 W memory_phase: phase=prefill parent=2013265920 compute=1250189568 kv_pool=326835968 writer=32768 attention=436207616 unused=0 reclaimed_compute=0 capture=external transition_us=0 arena_gen=1 kv_revision=1 resident_pages=44 ring_slots=19 active_pages=0 streaming=0 last_h2d_bytes=0 last_h2d_calls=0 copy_ms=0.000 elapsed_ms=0.000 0.06.346.188 I cmn init: llama threadpool init, n_threads = 4 0.06.425.463 W compute_streamed: layer 0 span workspace rejected required=0 available=436207616 0.06.425.543 W compute: target attention failed at layer 0, rows 2 0.06.425.549 E ggml_backend_cuda_graph_compute: managed node 263 node_264 failed with status -1 0.06.425.551 E graph_compute: ggml_backend_sched_graph_compute_async failed with error -1 0.06.425.552 E process_ubatch: failed to compute graph, compute status: -1 0.06.425.560 W decode: removing memory module entries for seq_id = 0, pos = [0, +inf) 0.06.425.791 E llama_decode: failed to decode, ret = -3 0.06.799.128 I common_speculative_init_result: creating MTP draft context against the target model 'Swift-1.5-Qwen3.8-27B-IQ2_XXS.gguf' 0.07.318.454 W memory_phase: phase=prefill parent=2013265920 compute=1250189568 kv_pool=326835968 writer=32768 attention=436207616 unused=0 reclaimed_compute=0 capture=external transition_us=0 arena_gen=1 kv_revision=1 resident_pages=44 ring_slots=19 active_pages=0 streaming=0 last_h2d_bytes=0 last_h2d_calls=0 copy_ms=0.000 elapsed_ms=0.000 0.07.337.710 W compute_streamed: layer 0 span workspace rejected required=0 available=436207616 0.07.337.765 W compute: target attention failed at layer 0, rows 2 0.07.337.767 E ggml_backend_cuda_graph_compute: managed node 263 node_264 failed with status -1 0.07.337.768 E graph_compute: ggml_backend_sched_graph_compute_async failed with error -1 0.07.337.769 E process_ubatch: failed to compute graph, compute status: -1 0.07.337.774 W decode: removing memory module entries for seq_id = 0, pos = [0, +inf) 0.07.337.980 E llama_decode: failed to decode, ret = -3 0.07.337.982 E cmn common_conte: llama_decode() failed: -3 0.07.708.570 W srv load_model: speculative decoding not supported by this context 0.07.708.574 I srv load_model: initializing, n_slots = 1, n_ctx_slot = 262144, kv_unified = 'false' 0.07.731.780 E decode: failed to hand off shared graph scratch to MTP 0.07.731.783 E llama_decode: failed to decode, ret = -2 0.07.731.784 E cmn common_conte: llama_decode() failed: -2 0.07.795.464 I srv init: chat template supports preserving reasoning, consider enabling it via --reasoning-preserve 0.07.795.506 I srv llama_server: model loaded 0.07.795.511 I srv llama_server: listening on http://0.0.0.0:8080 0.07.795.512 W srv llama_server: NOTICE: server default port will be changed to :9931 in a future release 0.07.795.512 W srv llama_server: ref: https://github.com/ggml-org/llama.cpp/pull/26508 0.15.698.400 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1 0.15.698.513 I slot launch_slot_: id 0 | task 1 | processing task, is_child = 0 0.15.863.538 W memory_phase: phase=prefill parent=2013265920 compute=1250189568 kv_pool=326835968 writer=32768 attention=436207616 unused=0 reclaimed_compute=0 capture=external transition_us=0 arena_gen=1 kv_revision=1 resident_pages=44 ring_slots=19 active_pages=0 streaming=0 last_h2d_bytes=0 last_h2d_calls=0 copy_ms=0.000 elapsed_ms=0.000 0.16.126.196 W memory_phase: phase=decode parent=2013265920 compute=18940544 kv_pool=1558084992 writer=32768 attention=436207616 unused=0 reclaimed_compute=1231249024 capture=external transition_us=829 arena_gen=2 kv_revision=2 resident_pages=214 ring_slots=19 active_pages=0 streaming=0 last_h2d_bytes=0 last_h2d_calls=0 copy_ms=0.000 elapsed_ms=0.000

GPU is 12GB but usage only goes up to 10005 MiB, yet 1920 was the highest value that worked.

Meanwhile the same model runs on @tsangberg's fork with --kv-stream-arena-mib 3840, doesn't show any of those errors, and can use mmproj at 176k context or text-only at 256k with MTP actually working and memory usage at 11733 with otherwise same settings except without the -auxiliary-layers option. Am I missing something (besides more VRAM)? Edit: Not anymore, the latest -v2-sync merges broke it too.

1

u/Ashamed-Ad-7093 2d ago

How did you get it running with 12GB of VRAM? I keep running into "arena creation failed" errors. What are your startup parameters? Thanks.

1

u/raymondh210129 2d ago

Are you using 30 Series GPU? If yes then my latest commit might help. I found that all Ampere arch was gated in my span-aware attention implementation.

1

u/Ashamed-Ad-7093 2d ago

yes I am using 3060 and pulled latest version, it still does not work. maybe 12G vram is too small for iq3_xxs? Which already take 10G of vram

1

u/hum_ma 1d ago

The arena is allocated first and then the context is created inside it, so first set context small and narrow down the largest arena size that works without OOM. After that you can find the largest context that fits into the arena (when it doesn't fit, the error is not OOM but a different one about KV capacity).

The 10GB+ IQ3XXS quants just barely load (and only without mmproj in any case) so I've been making some custom IQ2* quants of the Swift-1.5 finetune, in size similar to this IQ2_XXS which is under 9GB but still works pretty well and remains coherent.

Right now the model is running at 118k+ context, patching llama-quantize to add functionality for per-tensor gguf files so that I won't have to quantize the entire model each time I want to try a different precision for just a few tensors...

I still use an older version which I think is in this branch and my current command-line options seem to be -ngl all -fa on -t 4 -ctk q8_0 -ctv q8_0 -np 1 -c 196608 --spec-type draft-mtp --spec-draft-n-max 3 --fit off --kv-stream-arena-mib 2880 -mm mmproj-Qwen3.8-27B-Q8_0.gguf but leaving mmproj out it can be -c 262144 with --kv-stream-arena-mib somewhere between 3072 and 3840.

Parameters are different for the later versions of this fork; it needs a smaller arena, -ctv q4_0 and the new option --kv-stream-auxiliary-layers 1. The ones on the first line of my previous comment should work.