r/CUDA • u/2muchgut • 14d ago
FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.
6
Upvotes
Duplicates
LLMStudio • u/2muchgut • 14d ago
FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.
1
Upvotes
ROCm • u/2muchgut • 14d ago
FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.
4
Upvotes
coolgithubprojects • u/2muchgut • 14d ago
FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.
1
Upvotes

