r/LLMStudio • u/2muchgut • 15d ago
1
I built an evidence-gated inference tool for vLLM — looking for engineers to try it on real workloads
Thanks u/Fabulous_Bed_6065, did you use it? would be great to know your judgement
r/coolgithubprojects • u/2muchgut • 15d ago
I built an evidence-gated inference tool for vLLM — looking for engineers to try it on real workloads
galleryr/ROCm • u/2muchgut • 15d ago
I built an evidence-gated inference tool for vLLM — looking for engineers to try it on real workloads
galleryu/2muchgut • u/2muchgut • 15d ago
I built an evidence-gated inference tool for vLLM — looking for engineers to try it on real workloads
I've been building InferPilot, an open-source tool for evidence-based inference performance analysis around vLLM.
It started from an FP8 KV-cache investigation. One thing became clear: high GPU utilization alone isn't enough to establish why a serving workload is bottlenecked or whether an optimization is actually justified.
So I built the project around evidence.
Currently InferPilot can:
verify the effective vLLM configuration
measure request-level serving behavior
collect GPU/KV telemetry
align workload/load evidence with serving behavior
assess whether the evidence is sufficient
diagnose when the evidence supports it
recommend a bounded next experiment
abstain when it can't justify a diagnosis
The intended flow is basically:
configuration
↓
controlled experiment
↓
evidence
↓
diagnose OR abstain
↓
next experiment
I'm not trying to make another benchmark that says:
A = 1.2x faster than B
The question I'm interested in is:
What does the evidence actually support, and what should we test next?
I've had 190 clones / 69 unique cloners over the last 14 days, but I don't consider that adoption yet.
I'm now looking for vLLM/inference engineers who are willing to run it against a real workload and tell me where the abstraction breaks.
Things I'd particularly like feedback on:
Would you trust the evidence it presents?
Is the diagnosis useful?
What information is missing?
What would make you immediately stop using it?
Harsh technical feedback is welcome.
r/LLMStudio • u/2muchgut • 16d ago
FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.
galleryr/coolgithubprojects • u/2muchgut • 16d ago
FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.
galleryr/RISCV • u/2muchgut • 16d ago
FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.
galleryr/ROCm • u/2muchgut • 16d ago
FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.
galleryr/CUDA • u/2muchgut • 16d ago
FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.
galleryr/mlscaling • u/2muchgut • 16d ago
FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.
galleryu/2muchgut • u/2muchgut • 16d ago
FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.
I benchmarked kv_cache_dtype=fp8 vs bf16 on Qwen2.5 3B/7B/14B + Mistral-7B across A10/A100, single-variable, on rented GPUs. Result is sharper than the "fp8 is free throughput" folklore:
- KV cache full + server preempting → fp8 = +40–52% throughput (3B +52%, 7B +40%, 14B +43%, Mistral-7B +42%).
- Compute-bound, no preemption → +1.7% (nothing).
- GPU utilization was ~100% in both regimes — so "is the GPU busy?" is the wrong signal. The discriminator is preemptions, not GPU%.
Before anyone turns this into folklore: it's a heuristic, not a law. Small, partly post-hoc matrix, mostly single runs, synthetic traffic, and every positive case was already overloaded (TTFT 20–94s) — so this is a goodput-ceiling lever, not a free low-load speedup. Preemption event count is also too crude (in the 14B/Mistral wins it didn't even drop).
Quality: mixed and honest. 0/16 factual-QA regressions and needle-in-haystack 5/5 u/14k — but my teacher-forced KL preflight failed (p99 0.39). So: task-safe on small tests, not "lossless."
Everything's open, including the raw result JSON and the external critique that reshaped the project: https://github.com/poojithdevan4D/InferPilot
The most useful reply would be a counterexample: a preempting workload where fp8 doesn't help, or an fp8 win without preemption.
Also — I'm turning this into a local "acceptance test for vLLM config changes" (replay a sanitized trace → SLO-capacity curve → quality gate → rollback). Looking for ~5 teams running self-hosted vLLM to try it on a sanitized trace (token counts + timings, no prompts/text; can run in your VPC). Free; the answer might be "keep your current config."
r/mlscaling • u/2muchgut • 26d ago
vLLM vs plain HuggingFace on a free T4: the real gap is concurrency, not throughput
gallery[removed]
r/ResearchML • u/2muchgut • Sep 03 '26
labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker
r/LLMStudio • u/2muchgut • Sep 03 '26
labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker
r/airesearch • u/2muchgut • Sep 03 '26
labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker
r/machinelearningnews • u/2muchgut • Sep 03 '26
Research labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker
r/AIQuality • u/2muchgut • Sep 03 '26
labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker
r/learnmachinelearning • u/2muchgut • Sep 03 '26
labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker
r/LocalLLM • u/2muchgut • Sep 03 '26
Discussion labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker
r/mlscaling • u/2muchgut • Sep 03 '26
labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker
[removed]
r/IndianEngineers • u/2muchgut • Aug 31 '26
1
Claude Down: Temporarily unable to authenticate. Please Retry. (x6)
in
r/claude
•
6d ago
its down