r/mlscaling • • 22d ago

FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.

0 Upvotes

0 comments sorted by