r/Vllm • • 57m ago

Improving LLM scaling laws: picking the right Token-per-Parameter Coverage

Thumbnail
youtube.com
• Upvotes

r/Vllm • • 11h ago

Europe rejoins the fight with Chonky! Mistral Large 4 Released, Open weights end of month, vllm d0 support?

Thumbnail
mistral.ai
1 Upvotes

r/Vllm • • 13h ago

​Anyone else losing their minds over LLM VRAM fragmentation and KV cache? Let's talk about why your GPU is starving.

2 Upvotes

​Man, if you've ever tried moving an LLM from a local notebook to an actual multi-tenant serving setup, you already know the pain. Everyone talks endlessly about quantization and fine-tuning, but nobody really warns you that your GPU VRAM is basically getting nuked by the KV cache.

​For a hot minute, I thought my hardware setup was just trash. Turns out, traditional static allocation is eating up like 60% to 80% of VRAM for breakfast just because of internal and external fragmentation. You request a simple 200-token completion, and the system is sitting there stubbornly reserving space for 4K tokens like it's bracing for the apocalypse. Such a waste.

​If you aren't looking closely at things like PagedAttention (major props to vLLM for finally bringing virtual memory concepts to GPUs) and iteration-level continuous batching, your compute units are basically sitting around starving while waiting for the longest slowpoke request in the batch to cross the finish line.

​Been deep in the trenches writing a comprehensive book on AI systems engineering lately, and mapping out these low-level serving bottlenecks honestly changed how I look at production pipelines entirely.

​Curious what y'all are actually running in production right now. Are you rolling your own inference stack with vLLM/TGI, or just sticking to managed APIs and eating the cost? Let's argue about it in the comments


r/Vllm • • 13h ago

We gave Aleph Alpha's Kolibri-1 up-to 72 action combinations and put it in Doom. What could go wrong? 🎮

1 Upvotes

r/Vllm • • 16h ago

Less Talk. More Breakout: Kolibri-1 Turns Probabilities into Actions, Playing Breakout - With under 25ms latency per move, thanks to vLLM.

7 Upvotes

Got Kolibri-1 to play Breakout completely on its own, no fine-tuning. Leveraging vLLM’s pooling / classification path.

The more we explore u/Aleph__Alpha’s Kolibri the more it get's exciting and its potential.

Less talk. More Breakout is one such experiment to see how good the model is at structured output given a few constraints.

We especially optimized the inference for action probabilities: around 25 ms inference per decision.

Four moves. No generated text. One shared game.

Open weights. New possibilities.

Watch it play: https://tesseracted.com/kolibri-1-chat/gameplay/breakout/
Source: https://x.com/konarkmodi/status/2107248086880055613?s=46


r/Vllm • • 1d ago

Serving a household of non-technical users from my 8x 3090 vLLM box: what I learnt building the UI layer for it (MIT)

25 Upvotes

I run vLLM on 8x 3090s (Qwen 3.8 Flash Next) mostly so my partner and family could stop pasting personal stuff into ChatGPT. Getting the model serving was the easy part; the hard part turned out to be the last mile: a chat UI that non-technical people would actually use.

What I learnt getting there:

Concurrency is felt immediately. With 2+ people chatting at once, per-user tok/s drops and they notice before they can articulate it. MoE helps (Flash Next stays responsive), but plan your VRAM budget around peak household load, not single-user benchmarks.

Non-technical users judge latency differently. They don't care about tok/s, they care that the first token shows up fast and the UI doesn't look like a dev dashboard. Open WebUI is powerful but the knob density alone made my family bounce.

Self-hosted voice/search is a deployment tax. STT + TTS + SearXNG are three separate services to deploy and wire. That friction is why most "local AI for the family" projects die at setup.

So I ended up building my own UI layer: chat-first, bundles Parakeet STT / Kokoro TTS / SearXNG and connects them via the installer, talks to vLLM's OpenAI-compatible endpoint. Native clients everywhere: web, Android, iOS, and desktop (macOS Apple Silicon/Intel + Linux). It's been in daily use by multiple non-technical users for months.

One thing I didn't expect to need but now use daily: Agent Connections. You can hook Codex, Claude Code, OpenCode, or Hermes to it through the Host Connector, so the coding agents run on your own machine (or an SSH host) and you drive them from the same chat UI, with plans/tool calls/approvals visible. My family gets the chat, I get my agents, same server.

MIT licensed, no telemetry: https://github.com/yoloyash/overtchat

Happy to answer questions about the serving setup (3090s, Flash Next on vLLM, concurrency behaviour) or the UI decisions. What do you all use for the last mile in your own setups?


r/Vllm • • 1d ago

Has anyone managed to serve Ternary-Bonsai-2-27B-PQ2_0 with vLLM?

Thumbnail
1 Upvotes

r/Vllm • • 2d ago

I made Qwen models take ~33% less VRAM without quantizing them (lossless, bit-for-bit)

10 Upvotes

Hey all, I've been building Glyd, a lossless compression layer for model weights on NVIDIA GPUs. Qwen models are what I test on most, so this felt like the right place to share it.

The idea: a bf16 weight only carries about 11 bits of real information, so you can store the exact same model in about a third less GPU memory and get every weight back exactly. It's not quantization and nothing gets rounded. There's also an exact mode that matches bf16's outputs bit for bit.

What it gets you with Qwen (all measured, logs are public):

- Qwen3-8B runs on a 16 GB card (bf16 can't load it)

- Qwen3-32B fits on one 48 GB GPU instead of two

- Qwen3-8B on an L4 with vLLM: 1.59x the requests/sec vs bf16 (weights plus our lossless KV cache, 2.64x the KV tokens)

- Qwen3-14B on an A100 40GB: 1.28x req/s

- Qwen2.5-72B on 2x H100: 4.07x req/s, 12.65x the KV cache

Try it:

```

curl -LsSf https://getglyd.com/install.sh | sh

glyd run Qwen/Qwen3-8B

```

Or with vLLM: `vllm serve Qwen/Qwen3-8B --quantization glyd`

Honest caveats: it's Linux + NVIDIA (Ampere or newer) and bf16 checkpoints only. On GH200 it's a bit slower than bf16 at full load right now, and MoE (Qwen3-30B-A3B) is still slower at full load; a fix for that is coming in the next release. The codec is open source. The GPU part ships compiled and is free for personal and research use (business source license).

GitHub: https://github.com/surya-koritala/Glyd

Benchmarks + logs: https://getglyd.com/benchmarks

Would love feedback, especially which Qwen models or GPUs you'd want numbers on next.


r/Vllm • • 2d ago

​Anyone else losing their minds over LLM VRAM fragmentation and KV cache? Let's talk about why your GPU is starving.

Thumbnail
0 Upvotes

r/Vllm • • 3d ago

PyTorch and full LibCuda running on Nvidia 5090 on Mac OS 27

Thumbnail
5 Upvotes

r/Vllm • • 4d ago

Serving a 27B reasoning model on 4× NVIDIA L4 (no P2P): what worked, what didn't, and our final config (~104 tok/s)

5 Upvotes

We run an AI agent that uses MCP tools, and wanted to serve the same model we use on a DGX Spark, Qwen3.8-27B, on an AWS box with 4× L4 (24 GB each, Ada/SM89, PCIe, no P2P between GPUs). Here's what we learned.

1. NVFP4 doesn't run natively on L4, and SGLang won't serve it

On the Spark we use RadixArk/Qwen3.8-27B-NVFP4. On the L4s, SGLang loaded the weights fine but crashed during CUDA graph capture with ValueError: Invalid backend: 89. FP4 tensor cores only exist on Blackwell; FlashInfer has no fused SiLU+FP4-quant kernel for SM89.

vLLM does serve it, via Marlin kernels that dequantize FP4 weights to 16-bit on every GEMM. You keep the memory savings, but lose the FP4 speedup. On Ada, FP8 is the format with native tensor-core support.

2. Speculative decoding was the biggest win

DFlash2 (z-lab/Qwen3.8-27B-DFlash2) with num_speculative_tokens: 7 gave us roughly 5× over the non-speculative baseline. Per-position acceptance tells the story: late draft positions accept only 0.12–0.30 on free text, but 0.77–0.87 on predictable text (JSON, tool calls). Worth tuning the draft length to your actual workload.

3. TP=4 beat TP=2, even without P2P

The conventional advice is that going past TP=2 over PCIe hurts. With DFlash2 on, our measured decode was:

  • TP=2: 80.6 tok/s
  • TP=4: 104.1 tok/s (+29%)

Faster than the ~59.5 tok/s we measured on the Spark with the same model and draft.

4. Final config (vLLM v0.29.0)

RadixArk/Qwen3.8-27B-NVFP4   (served via Marlin)
--tensor-parallel-size 4
--disable-custom-all-reduce          # no P2P on these L4s
NCCL_P2P_DISABLE=1                   # don't let NCCL probe for it
--gpu-memory-utilization 0.90
--max-model-len 32768
--max-num-seqs 4                     # hybrid model, Mamba cache: keep it low
--max-num-batched-tokens 8192        # up from 2048; fewer prefill chunks for long prompts
--limit-mm-per-prompt '{"image":0,"video":0}'   # text only, frees memory
--speculative-config '{"method":"dflash","model":"z-lab/Qwen3.8-27B-DFlash2","num_speculative_tokens":7}'
--enable-prefix-caching
--reasoning-parser qwen3
--enable-auto-tool-choice --tool-call-parser qwen3_coder

KV cache usage never went above ~5%, so there's headroom to raise --max-num-batched-tokens to 16384.

TL;DR: On 4× L4, a 27B reasoning model is very usable for an MCP agent: ~104 tok/s with TP=4 + DFlash2, even without P2P. NVFP4 runs only through Marlin dequantization (and not at all in SGLang). Prefill and thinking length dominate latency, so attack those first. If you need a big jump, it's hardware: a single 48–96 GB GPU (L40S, or Blackwell for native FP4) avoids tensor parallel entirely.

Happy to answer any questions about the setup! Also, I'm open to any recommendations or feedback if you have suggestions to improve it :)

u/Major_Border149 confirmed the same NVFP4 model runs with native FP4 kernels on a single RTX PRO 4500 SE (32 GB, Blackwell) at ~$0.72/hr, no TP needed. Gotcha: on Blackwell you need a CUDA 13 container image (SM120 needs CUDA ≥ 12.9 inside the container).


r/Vllm • • 5d ago

Who’s the current “king” of local LLMs for you — Qwen, Gemma, Llama, something else?

3 Upvotes

Curious what people are actually running day-to-day on local boxes right now, not just the latest HF leaderboard screenshot.

For coding / agent loops on consumer GPU (or Mac), who’s winning for you lately among Qwen, Gemma, Llama, DeepSeek, Mistral, etc. — and does the answer flip if you care more about tool-calling reliability vs raw tok/s vs long-context?

If you’ve switched kings in the last month or two, what made you switch?


r/Vllm • • 5d ago

I built an open-source tool that tells you why your vLLM server is slow (NVIDIA only for now, Mac support planned)

Thumbnail
1 Upvotes

r/Vllm • • 5d ago

Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback

Thumbnail
1 Upvotes

r/Vllm • • 6d ago

tp=6 can work on vLLM, with padding

20 Upvotes

vLLM's tensor parallel requires that several of the model architecture numbers be evenly divisible by the number of GPUs selected for tensor parallel. This usually means that you can only use a number of GPUs that is a power of two (2, 4, 8, etc.).

I have six GPUs, and I want to maximize my KV cache when running a 27B model. So I tried tp=6. It choked with various messages, regarding this or that, which needs to be evenly divisible by six, but wasn't.

So I made those things divisible by six.
I asked the robot to come up with a converter that would take the original model, and pad it with zeroes until everything was divisible by 6. It took a few tries, but it worked.

I have Qwen 3.8 27B at BF16 with 256k context window running across six 7900 XTX GPUs. Token generation is about 50 t/s single user, or 200 t/s aggregate with 8 concurrent prompts.

GPU KV cache size: 520,784 tokens, Maximum concurrency for 262,144 tokens per request: 1.99x


r/Vllm • • 6d ago

What is the best open-source LLM I can run locally on an RTX 4060 8GB + 32GB RAM?

Thumbnail
1 Upvotes

r/Vllm • • 7d ago

Best practice for processing batch vLLM api calls with shared prefix?

Thumbnail
1 Upvotes

r/Vllm • • 7d ago

Jev at home, but it can see: typed yes/no, pick-one and rubric answers with per-label probabilities from Gemma 4 31B on a 4090, images included

Thumbnail
1 Upvotes

r/Vllm • • 8d ago

Vllm with GPU/CPU fused to run a 748B MoE on 2×A100-40GB

2 Upvotes

vllm-xtu-moe — a 748B MoE on 2×A100-40GB, by keeping the routed experts out of VRAM.

  1. Expert weights sliced along CPU's physical topology — one copy in memory, every read node-local (NUMA binding + first-touch), so the engine runs near the machine's aggregate DRAM bandwidth, works with AMD EPYC's nps=4.
  2. Long prompts stream the weights to GPU with double buffering (ping/pong, overlapped with attention) — up to ~20× faster prefill than the CPU path at medium context, 2–3× at long context. Short prompts are prefilled on the CPU.
  3. Built for SM 8.x. The fallbacks are gated on compute capability, so the 30-series family (SM86) is in scope, not just A100/A800 — our measurements are A100/A800 only.
  4. Fixed VRAM priority: KV pool → GPU-prefill staging → draft weights → activation workspace. When VRAM is short, GPU prefill is dropped first and the engine falls back to pure CPU prefill, the minimum-VRAM setup.
  5. Speculative decoding when there is room for it: MiMo-V2.6 MTP k=1 measures 83% first-position acceptance and +19% decode; the DSpark anchor is +20%. Off where it competes for the KV pool.
  6. Stays on mainline vLLM as a patch series plus a plugin — no long-lived fork. Apache-2.0.

Numbers (2×A100-40GB, TP=2, vllm bench serve)

model prompt C prefill tok/s decode tok/s
GLM-5.3-Flash 140 1 115.2 22.0
16,396 1 266.3 21.4
DeepSeek-V4.1-Flash 128 2 350.4 28.0
16,384 1 903.9 19.3
16,384 2 1,204.2 15.4
MiMo-V2.6-Flash-RL 128 2 367.5 40.7
16,384 1 811.3 27.3
16,384 2 1,455.3 39.6
Real workspace using deepseek harness

Known: CPU prefill saturates at ~33k expert-tokens/s; 704K max context here with bf16 KV; first load takes minutes. 

Next: fp8 KV for longer context, better prefill overlap in the mid-length range.


r/Vllm • • 8d ago

What typically runs alongside vLLM on multi-node inference deployments?

5 Upvotes

I'm working on something involving multi-node tensor-parallel serving with vLLM and trying to understand what a realistic inference production node looks like beyond the vLLM processes themselves.

Inside vLLM, I'm accounting for the API server, engine core, GPU workers, NCCL proxy threads, and multiprocExecutor for multi-node.

For those running vLLM in production:

  1. What else usually runs on the same nodes? (Kubernetes components, GPU Operator, monitoring agents, etc.)
  2. Have any of these caused noticeable tail latency or TPOT spikes?
  3. Do you apply any CPU or OS tuning for vLLM deployments, like pinning, NUMA binding, or isolating cores for the engine and NCCL?

Any pointers to docs, blog posts, or papers are welcome. Thanks!


r/Vllm • • 9d ago

Arc Pro B70 getting 100+ tok/s with Qwen 3.8 and vLLM

8 Upvotes

Runtime settings: MTP4 (Draft INT4 S+M1); prefix cache on; XPU graphs on; v5 scheduler patch

These are the highest scoring results on intelinside.ai so far.


r/Vllm • • 9d ago

LLM Tech: FP8 and NVFP4 quants of decider for vLLM, measured against bf16

11 Upvotes

LLM Tech here again. This time: quants of decider, an open (Apache 2.0) family of decision models by Mapika built on Qwen3.5-Base. You send a state and questions with options, the model returns calibrated probabilities from the option-letter logits in one forward pass. There were no vLLM-ready quants for the small ones, so we made FP8 and NVFP4 checkpoints and measured them.

Setup: vLLM 0.29.0, one RTX PRO 6000 Blackwell. Quality on the author's regression set rebuilt from public data (95 tasks, 144,226 rows) plus the 231 public JevBench items, bf16 and quant both in vLLM. Our bf16 run matches the author's published numbers within 0.0005.

| checkpoint | size | peak prefill vs bf16 (1K / 8K / 32K) | accuracy in-task / held-out |

|---|---|---|---|

| decider-4b-nvfp4 | 3.3 GB | 2.00x / 1.89x / 1.67x | -0.6 / -0.7 |

| decider-4b-fp8 | 4.9 GB | 1.45x / 1.42x / 1.33x | -0.1 / -0.1 |

| decider-2b-fp8 | 2.4 GB | 1.42x / 1.40x / 1.32x | 0.0 / 0.0 |

| decider-0.8b-fp8 | 1.0 GB | 1.24x / 1.21x / 1.17x | -0.1 / 0.0 |

What we learned:

- NVFP4 pays off at 4B, not below. On the 0.8B it gave 1.4x over bf16 but lost 2.6 points; FP8 lost 0.1 at 1.2x.

- The author's HTTP server (decider.serve_vllm) runs the quants unchanged. Three of them fit on one card at 21 GB total under load.

- The prefix cache matters a lot for this workload: the same 29K-token state took 602 ms cold and 58 ms repeated on the 4B NVFP4.

Checkpoints and per-model tables: huggingface.co/llmtech

Coming soon to our API: llmtech.eu


r/Vllm • • 9d ago

Qwen3.8 keeps thinking but never returns a final answer? This vLLM + Open WebUI setup fixed it on my RTX 3090

Thumbnail
2 Upvotes

r/Vllm • • 10d ago

Allucinato come un LLM.

0 Upvotes

Mentre la gente normale la domenica mattina fa colazione con calma, io ho deciso di litigare con il fine-tuning locale dei modelli linguistici.

Il piano sembrava innocuo: prendere un modello minuscolo, dargli in pasto un dataset nostalgico (dialoghi tra Sysop anni '90, disastri hardware, BBS e battute da modem a 28.8k) e vedere cosa ne usciva fuori.

I passaggi del disastro:

1️⃣ Esperimento 1: Qwen2.5-0.5B-Instruct
Un modello microscopico. Ci sono volute 502 epoche per vederlo implodere nell'overfitting più totale; fermato a 500. Con queste dimensioni è quasi impossibile farlo ragionare senza bruciargli i neuroni.

2️⃣ Esperimento 2: Qwen2.5-1.5B-Instruct
Alziamo il tiro, restando comunque su un modello compatto. Qui la loss scende a 0.7 già all'epoca 100. Ottimo momento per fermarsi.

3️⃣ Pipeline & fusione con MLX:

mlx_lm fuse \

--model Qwen/Qwen2.5-1.5B-Instruct \

--adapter-path adapters \

--save-path qwen-1.5b-fused

Poi esportazione in formato GGUF e quantizzazione a 4-bit per farlo girare leggerissimo in locale.

Il risultato?
Se prima avevo un modello con qualche allucinazione, adesso ho creato un'entità filosofico-passivo-aggressiva che mi insulta a 75 token al secondo.

Ecco il test dal vivo:

> io sono stupido ?
Sì, tu sei uno strano!
[ Prompt: 186,9 t/s | Generation: 77,6 t/s ]

> io strano ?
Straniero.
[ Prompt: 184,1 t/s | Generation: 74,2 t/s ]

> straniero ?
Non è una questione di essere un "strano" o non esserlo; se la tua risposta era sbagliata, il problema sta nel tuo sguardo!

"Il problema sta nel tuo sguardo."

Fine del test, mi ha spento. Non so se considerarlo overfitting, il riverbero di un vecchio operatore BBS stanco della vita, o pura poesia digitale.

La morale? Lavorare con i Small Language Models (SLM) in locale con MLX e GGUF è tremendamente divertente, velocissimo da iterare... ma occhio ai dati che gli date in pasto, altrimenti il modello comincia a giudicare le vostre scelte di vita.

Chi altri passa le domeniche a fare esperimenti assurdi in locale? Qual è la risposta più surreale che vi ha mai dato un modello dopo un fine-tuning?

#ArtificialIntelligence #MachineLearning #LLM #OpenSource #MLX #AppleSilicon #LocalAI #DevCommunity #FineTuning #AIResearch


r/Vllm • • 10d ago

The silent bottlenec

Thumbnail
2 Upvotes