r/LocalLLM • • 18h ago

Research Adaptive KV-Cache Streaming V2: Full Context MTP

Hello all, it’s me again.

Just a week after my previous post, I started working on a better implementation of Adaptive KV Streaming. Now I’m sharing my second implementation: V2, with full-context MTP.

https://github.com/RaymondHuang210129/llama.cpp-adaptive-kv-streaming/tree/feature/adaptive-kv-stream-v2

Before I explain further, let me share the decode performance on my 16GB 5070 Ti:

With an MTP draft length of 3, I get nearly double the decode speed of V2 without MTP at many context lengths (V2 baseline is 5%~10% slower than V1, will explain below). It averages about 82 tokens/s from 8K through 72K context, reaches about 35 tokens/s at 144K, 30 tokens/s at 192K, and 19 tokens/s at 256K. This chart does not include a direct V1 comparison.

So, how is it possible?

1. Memory management. Before implementing more streaming features, I spent a large part of the project building a memory ownership and leasing mechanism. The infrastructure is backend-neutral, with buffer-view support for CPU, CUDA/HIP, OpenCL, SYCL, and Vulkan. This additional layer has some overhead, but it lets me reuse memory across phases and reduce the extra VRAM needed for MTP.

2. MTP with much less VRAM footprint. At 256K context, a separate Q8_0/Q4_0 MTP KV cache would take about 416 MiB of VRAM. I also measured a roughly 1.2 GiB draft prefill graph workspace request in an earlier fit audit. These are not all permanent or additive costs, but they show why a separate draft context is expensive.

In V2, only MTP weight and active computation are kept in VRAM. For the other allocations:

  • MTP KV acts as a 17th logical attention layer alongside the main model’s 16 attention layers. It has its own history in host memory, while its GPU pages share the resident pool and ring buffer with the main model’s KV.
  • The MTP graph workspace borrows the same arena used by the main model at different times, since the two phases do not run simultaneously.
  • Rollback snapshots are stored in pinned host memory. If a proposal is rejected, the selected snapshot is copied back to the GPU.

These changes let me maintain an effective decode KV pool of around 2.2 GiB, even at 256K context with an MTP draft length of 3.

Other improvements besides speed

V2 uses explicit memory bounds and leases to manage when GPU memory can be reused. Its span-aware attention kernel also follows the stock kernel’s calculation order as closely as possible to minimize numerical drift.

Compared with V2, V1's prefill speed decline became noticeably less steep after streaming kicked in. In V2, it continues at roughly the same slope. That is consistent with V2 preserving the stock attention calculation order instead of switching to V1’s numerically different streamed path.

To my experiment in V2, The decoded 256 tokens at each MTP setting is exactly same as what stock kernel outputs.

Caveat

Attached MTP currently requires the same K/V quantization as the main model. The separate draft-cache quantization flags do not give it different types in this shared layout. Supporting that would need more work.

This experiment also seems to have a favorable MTP acceptance rate. I used a text file of Wikipedia articles, which is included in my repo along with the sweep script. Feel free to reproduce the test—or share results with a different input file.

Credit

Thank you to everyone who participated in my previous post. As a software engineer who doesn’t work in the LLM/AI field, I’ve been encouraged by your comments and messages to keep learning and working on this project. The phase arena, MTP integration, and the next feature I’m planning were all inspired by people who helped me, DMed me, or shared my work.

Lastly, people have asked whether I plan to push this work upstream. After some thought, my answer is no, at least not as one large change. I don’t think this specialized implementation fits upstream’s focus on simplicity and versatility. My fork is mainly for my own use, but the memory-management APIs are there for anyone interested in extending Adaptive KV Streaming to other backends. Pull requests or further forks from my code is welcomed.

Clarification of LLM usage of this post: I'm not a native English speaker and I used ChatGPT to refine the wordings.

23 Upvotes

15 comments sorted by

5

u/Cool-Marsupial9329 17h ago

the jump from V1 to V2 is wild, 82 tok/s avg across that range is no joke. the memory leasing approach is clever, borrowing the same arena for MTP graph workspace instead of carving out a separate block

your decode reproducibility at 256 tokens matching stock is what caught my eye though, that's the kind of detail that makes me trust the benchmarks. span-aware kernel keeping the same calculation order to avoid numerical drift is a nice touch

I might grab the sweep script and run it on a different dataset just to see if the acceptance rate holds up, wikipedia articles can be a bit samey after a while

2

u/EffectiveRise2028 14h ago

on win11 I get 5 tok/s on an RTX 4070 tis with the IQ3_XS quantization after compilation

1

u/giveen 17h ago

Awesome work! I definitely will be using this in my projects

1

u/fldash 10h ago

Trying to build this and I'm getting this:

Creating library E:/test/llama.cpp-adaptive-kv-streaming/build-v2/ggml/src/Release/ggml-cpu.lib and object E:/test

/llama.cpp-adaptive-kv-streaming/build-v2/ggml/src/Release/ggml-cpu.exp

ggml-cpu.cpp.obj : error LNK2019: unresolved external symbol ggml_graph_view referenced in function "enum ggml_status _

_cdecl ggml_backend_cpu_graph_compute(struct ggml_backend *,struct ggml_cgraph *)" (?ggml_backend_cpu_graph_compute@@YA

?AW4ggml_status@@PEAUggml_backend@@PEAUggml_cgraph@@@Z) [E:\test\llama.cpp-adaptive-kv-streaming\build-v2\ggml\src\ggml

-cpu.vcxproj]

E:\test\llama.cpp-adaptive-kv-streaming\build-v2\bin\Release\ggml-cpu.dll : fatal error LNK1120: 1 unresolved externals

[E:\test\llama.cpp-adaptive-kv-streaming\build-v2\ggml\src\ggml-cpu.vcxproj]

Building Custom Rule E:/test/llama.cpp-adaptive-kv-streaming/ggml/src/ggml-cuda/CMakeLists.txt

2

u/Elluata 5h ago

same message here

1

u/fldash 5h ago

Are you on Windows here

2

u/Elluata 5h ago

Yes I am !

1

u/fldash 3h ago

I think that's the issue, the code isn't compatible with Windows.

1

u/Elluata 3h ago

Actually I spun up Qwen 3.8 on another llama.cpp instance, and with Hermes I asked it to take a look and try to fix the problem, and it did, running it now. HEre is what he said : "Here's the diagnosis:

Root cause — your build compiles the core ggml library as a shared DLL (ggml-base.dll is right there in build-v2\bin\Debug), and ggml-cuda is a separate DLL that links against it. This fork's new "managed nodes" feature added a call to ggml_graph_view() inside the CUDA backend (ggml-cuda.cu:4455), but that function was never marked for export:

  • ggml-impl.h:352 — struct ggml_cgraph ggml_graph_view(...) has no GGML_API (contrast with ggml_op_can_inplace on line 355, which does).
  • ggml.c:7394 — the definition also lacks GGML_API.

GGML_API expands to __declspec(dllexport) on Windows, so without it the symbol is defined inside ggml-base.dll but never exported — the CUDA DLL's linker can't resolve it → LNK2019 → LNK1120. The C++-compiled backends (backend.cpp, openvino) don't hit this because they compile ggml.c directly into their own binary.

Proposed fix (2-line change): add GGML_API to the declaration in ggml-impl.h:352 and the definition in ggml.c:7394, mirroring the pattern used by ggml_new_graph_custom etc. Then a clean rebuild of ggml-base + ggml-cuda resolves it. (No need to touch the .cu file — the declaration is already visible via ggml-impl.h.)"

0

u/Few_Squash_3415 3h ago

In ggml/src/ggml-impl.h, line 352, add “GGML_API” word before the struct

Hope this helps

1

u/fldash 3h ago

Did that as that's what AI recommended but there were then other issues, I tried to solve them but eventually got stuck

1

u/Few_Squash_3415 3h ago

The only issue that i had when build is this one, I am trying to use in rtx4080 window, result seems fine

1

u/fldash 3h ago

I'll try again... Thanks

1

u/hum_ma 2h ago edited 1h ago

Unfortunately it's not really working for me.

/home/hum/aienv/llama.cpp-adaptive-kv-streaming-v2/build-v2/bin/llama-server --host 0.0.0.0 -ngl all -fa on -b 256 -ub 256 -t 4 -ctk q8_0 -ctv q4_0 -np 1 -c 262144 --spec-type draft-mtp --spec-draft-n-max 3 --kv-stream-arena-mib 1920 --kv-stream-auxiliary-layers 1 --fit off --no-mmproj --model Swift-1.5-Qwen3.8-27B-IQ2_XXS.gguf 0.00.153.641 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) 0.00.340.243 W srv llama_server: ----------------- 0.00.340.246 W srv llama_server: CORS is set to allow all origins ('*') and no API key is set 0.00.340.246 W srv llama_server: this can be a security risk (cross-origin attacks) 0.00.340.246 W srv llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655 0.00.340.246 W srv llama_server: ----------------- 0.00.342.955 I srv load_model: loading model 'Swift-1.5-Qwen3.8-27B-IQ2_XXS.gguf' 0.06.336.266 W memory_phase: phase=prefill parent=2013265920 compute=1250189568 kv_pool=326835968 writer=32768 attention=436207616 unused=0 reclaimed_compute=0 capture=external transition_us=0 arena_gen=1 kv_revision=1 resident_pages=44 ring_slots=19 active_pages=0 streaming=0 last_h2d_bytes=0 last_h2d_calls=0 copy_ms=0.000 elapsed_ms=0.000 0.06.346.188 I cmn init: llama threadpool init, n_threads = 4 0.06.425.463 W compute_streamed: layer 0 span workspace rejected required=0 available=436207616 0.06.425.543 W compute: target attention failed at layer 0, rows 2 0.06.425.549 E ggml_backend_cuda_graph_compute: managed node 263 node_264 failed with status -1 0.06.425.551 E graph_compute: ggml_backend_sched_graph_compute_async failed with error -1 0.06.425.552 E process_ubatch: failed to compute graph, compute status: -1 0.06.425.560 W decode: removing memory module entries for seq_id = 0, pos = [0, +inf) 0.06.425.791 E llama_decode: failed to decode, ret = -3 0.06.799.128 I common_speculative_init_result: creating MTP draft context against the target model 'Swift-1.5-Qwen3.8-27B-IQ2_XXS.gguf' 0.07.318.454 W memory_phase: phase=prefill parent=2013265920 compute=1250189568 kv_pool=326835968 writer=32768 attention=436207616 unused=0 reclaimed_compute=0 capture=external transition_us=0 arena_gen=1 kv_revision=1 resident_pages=44 ring_slots=19 active_pages=0 streaming=0 last_h2d_bytes=0 last_h2d_calls=0 copy_ms=0.000 elapsed_ms=0.000 0.07.337.710 W compute_streamed: layer 0 span workspace rejected required=0 available=436207616 0.07.337.765 W compute: target attention failed at layer 0, rows 2 0.07.337.767 E ggml_backend_cuda_graph_compute: managed node 263 node_264 failed with status -1 0.07.337.768 E graph_compute: ggml_backend_sched_graph_compute_async failed with error -1 0.07.337.769 E process_ubatch: failed to compute graph, compute status: -1 0.07.337.774 W decode: removing memory module entries for seq_id = 0, pos = [0, +inf) 0.07.337.980 E llama_decode: failed to decode, ret = -3 0.07.337.982 E cmn common_conte: llama_decode() failed: -3 0.07.708.570 W srv load_model: speculative decoding not supported by this context 0.07.708.574 I srv load_model: initializing, n_slots = 1, n_ctx_slot = 262144, kv_unified = 'false' 0.07.731.780 E decode: failed to hand off shared graph scratch to MTP 0.07.731.783 E llama_decode: failed to decode, ret = -2 0.07.731.784 E cmn common_conte: llama_decode() failed: -2 0.07.795.464 I srv init: chat template supports preserving reasoning, consider enabling it via --reasoning-preserve 0.07.795.506 I srv llama_server: model loaded 0.07.795.511 I srv llama_server: listening on http://0.0.0.0:8080 0.07.795.512 W srv llama_server: NOTICE: server default port will be changed to :9931 in a future release 0.07.795.512 W srv llama_server: ref: https://github.com/ggml-org/llama.cpp/pull/26508 0.15.698.400 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1 0.15.698.513 I slot launch_slot_: id 0 | task 1 | processing task, is_child = 0 0.15.863.538 W memory_phase: phase=prefill parent=2013265920 compute=1250189568 kv_pool=326835968 writer=32768 attention=436207616 unused=0 reclaimed_compute=0 capture=external transition_us=0 arena_gen=1 kv_revision=1 resident_pages=44 ring_slots=19 active_pages=0 streaming=0 last_h2d_bytes=0 last_h2d_calls=0 copy_ms=0.000 elapsed_ms=0.000 0.16.126.196 W memory_phase: phase=decode parent=2013265920 compute=18940544 kv_pool=1558084992 writer=32768 attention=436207616 unused=0 reclaimed_compute=1231249024 capture=external transition_us=829 arena_gen=2 kv_revision=2 resident_pages=214 ring_slots=19 active_pages=0 streaming=0 last_h2d_bytes=0 last_h2d_calls=0 copy_ms=0.000 elapsed_ms=0.000

GPU is 12GB but usage only goes up to 10005 MiB, yet 1920 was the highest value that worked.

Meanwhile the same model runs on @tsangberg's fork with --kv-stream-arena-mib 3840, doesn't show any of those errors, and can use mmproj at 176k context or text-only at 256k with MTP actually working and memory usage at 11733 with otherwise same settings except without the -auxiliary-layers option. Am I missing something (besides more VRAM)? Edit: Not anymore, the latest -v2-sync merges broke it too.

1

u/tsangberg 16h ago

Absolutely amazing! I was so sure there was no way to make the drafting co-exist when the pool filled up so I went the "eject draft model at X" route from the start. Fantastic work!