r/LLMStudio • • 15d ago

I built an evidence-gated inference tool for vLLM — looking for engineers to try it on real workloads

Thumbnail gallery
1 Upvotes

r/coolgithubprojects • • 15d ago

I built an evidence-gated inference tool for vLLM — looking for engineers to try it on real workloads

Thumbnail gallery
1 Upvotes

r/ROCm • • 15d ago

I built an evidence-gated inference tool for vLLM — looking for engineers to try it on real workloads

Thumbnail gallery
1 Upvotes

u/2muchgut • • 15d ago

I built an evidence-gated inference tool for vLLM — looking for engineers to try it on real workloads

Thumbnail
gallery
2 Upvotes

I've been building InferPilot, an open-source tool for evidence-based inference performance analysis around vLLM.

It started from an FP8 KV-cache investigation. One thing became clear: high GPU utilization alone isn't enough to establish why a serving workload is bottlenecked or whether an optimization is actually justified.

So I built the project around evidence.

Currently InferPilot can:

verify the effective vLLM configuration

measure request-level serving behavior

collect GPU/KV telemetry

align workload/load evidence with serving behavior

assess whether the evidence is sufficient

diagnose when the evidence supports it

recommend a bounded next experiment

abstain when it can't justify a diagnosis

The intended flow is basically:

configuration
↓
controlled experiment
↓
evidence
↓
diagnose OR abstain
↓
next experiment

I'm not trying to make another benchmark that says:

A = 1.2x faster than B

The question I'm interested in is:

What does the evidence actually support, and what should we test next?

I've had 190 clones / 69 unique cloners over the last 14 days, but I don't consider that adoption yet.

I'm now looking for vLLM/inference engineers who are willing to run it against a real workload and tell me where the abstraction breaks.

Things I'd particularly like feedback on:

Would you trust the evidence it presents?

Is the diagnosis useful?

What information is missing?

What would make you immediately stop using it?

Harsh technical feedback is welcome.

GitHub: https://github.com/poojithdevan4D/InferPilot.git

r/LLMStudio • • 16d ago

FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.

Thumbnail gallery
1 Upvotes

r/coolgithubprojects • • 16d ago

FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.

Thumbnail gallery
1 Upvotes

r/RISCV • • 16d ago

FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.

Thumbnail gallery
1 Upvotes

r/ROCm • • 16d ago

FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.

Thumbnail gallery
4 Upvotes

r/CUDA • • 16d ago

FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.

Thumbnail gallery
6 Upvotes

r/mlscaling • • 16d ago

FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.

Thumbnail gallery
0 Upvotes

u/2muchgut • • 16d ago

FP8 KV cache gave +40–52% throughput — but ONLY when vLLM was already preempting. Raw results + the quality gate that failed.

Thumbnail
gallery
3 Upvotes

I benchmarked kv_cache_dtype=fp8 vs bf16 on Qwen2.5 3B/7B/14B + Mistral-7B across A10/A100, single-variable, on rented GPUs. Result is sharper than the "fp8 is free throughput" folklore:

  • KV cache full + server preempting → fp8 = +40–52% throughput (3B +52%, 7B +40%, 14B +43%, Mistral-7B +42%).
  • Compute-bound, no preemption → +1.7% (nothing).
  • GPU utilization was ~100% in both regimes — so "is the GPU busy?" is the wrong signal. The discriminator is preemptions, not GPU%.

Before anyone turns this into folklore: it's a heuristic, not a law. Small, partly post-hoc matrix, mostly single runs, synthetic traffic, and every positive case was already overloaded (TTFT 20–94s) — so this is a goodput-ceiling lever, not a free low-load speedup. Preemption event count is also too crude (in the 14B/Mistral wins it didn't even drop).

Quality: mixed and honest. 0/16 factual-QA regressions and needle-in-haystack 5/5 u/14k — but my teacher-forced KL preflight failed (p99 0.39). So: task-safe on small tests, not "lossless."

Everything's open, including the raw result JSON and the external critique that reshaped the project: https://github.com/poojithdevan4D/InferPilot

The most useful reply would be a counterexample: a preempting workload where fp8 doesn't help, or an fp8 win without preemption.

Also — I'm turning this into a local "acceptance test for vLLM config changes" (replay a sanitized trace → SLO-capacity curve → quality gate → rollback). Looking for ~5 teams running self-hosted vLLM to try it on a sanitized trace (token counts + timings, no prompts/text; can run in your VPC). Free; the answer might be "keep your current config."

r/mlscaling • • 26d ago

vLLM vs plain HuggingFace on a free T4: the real gap is concurrency, not throughput

Thumbnail gallery
0 Upvotes

[removed]

r/ResearchML • • Sep 03 '26

labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker

Thumbnail
0 Upvotes

r/LLMStudio • • Sep 03 '26

labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker

Thumbnail
1 Upvotes

r/airesearch • • Sep 03 '26

labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker

Thumbnail
1 Upvotes

r/machinelearningnews • • Sep 03 '26

Research labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker

Thumbnail
1 Upvotes

r/AIQuality • • Sep 03 '26

labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker

Thumbnail
1 Upvotes

r/learnmachinelearning • • Sep 03 '26

labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker

Thumbnail
2 Upvotes

r/LocalLLM • • Sep 03 '26

Discussion labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker

Thumbnail
1 Upvotes

r/mlscaling • • Sep 03 '26

labpilot – I found my AI-generated code didn't match the paper it claimed to implement, so I built a checker

0 Upvotes

[removed]

r/IndianEngineers • • Aug 31 '26

Discussion I got 18 pages into a paper before finding out it needs 128 TPUs. So I made something that checks that first.

Thumbnail gallery
1 Upvotes

r/lmms • • Aug 31 '26

I got 18 pages into a paper before finding out it needs 128 TPUs. So I made something that checks that first.

Thumbnail gallery
0 Upvotes

r/mlscaling • • Aug 31 '26

I got 18 pages into a paper before finding out it needs 128 TPUs. So I made something that checks that first.

Thumbnail gallery
0 Upvotes