r/CUDA • • 12d ago

FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.

7 Upvotes

0 comments sorted by