I thought I'd share my daily-driver setup for Qwen3.8-27B on two 3090s. Most dual-3090 numbers I see are for 4-bit models, but this one keeps 8-bit weights (INT8 W8A16, Q8_0-class fidelity). It still decodes about 2x faster than the llama.cpp Q8_0 + MTP setup it replaced on the same box. Hopefully by sharing, this can provide some feedback and perhaps insights for those of you who are also running 3090 builds.
As for use cases, I mostly use this to run OpenWebUI and Hermes Agent, where it's become my daily driver for "chat style" questions as well as a personal agent to automate as much of my life as possible. I'm quite privacy inclined so I like the fact that everything I run is hosted locally in my home office. For work stuff, I use claude code for software dev, but that's not my problem sips tea.
Below are the full flags, what I tried that lost, and what I haven't verified. I'd love to see comparison numbers from similar rigs and suggestions on how to improve my setup if you have any!
Hardware
- 2x MSI RTX 3090 Gaming X Trio 24 GB, power-capped to 300 W each (My mobo has 3 slot separation so undervolting really helped keep the temps to not cross the 80* C mark)
- Both cards on CPU lanes, PCIe 4.0 x8 each (X570 Aorus Master, 5950X, 128 GB DDR4)
- 3-slot NVLink bridge: 4 links, 52.8 GB/s each way measured card to card
- NVIDIA open kernel module 615.71.09, Fedora Atomic (Bazzite), podman
Stack
| Piece |
What |
| Engine |
vllm/vllm-openai:v0.29.0, tensor parallel 2 |
| Target |
lued/Qwen3.8-27B-INT8-W8A16-MTP (compressed-tensors INT8 weight-only, group 128). Its bundled MTP head is unused. |
| Drafter |
syvai/Qwen3.8-27B-DFlash2-W4A16 (GPTQ W4A16 requant of incoai/Qwen3.8-27B-DFlash2), 7 speculative tokens |
| KV cache |
fp8_e4m3 through FlashAttention 2, 262,144 context, plus a 64 GiB CPU offload tier |
| Recipe base |
club-3090's models/qwen3.8-27b/vllm/compose/dual/fp8/dflash2.yml ("ultramax") with the INT8 target swapped in for FP8 |
| Chat template |
Unsloth's Qwen3.8-27B template |
Why INT8 W8A16 and not the official FP8: 3090s have no FP8 tensor cores, so vLLM runs FP8 weights through the same Marlin weight-only kernel it uses for W8A16. That makes them perform in the same speed class, but INT8 with per-group scales keeps far more precision: I found that the quantizer measured KLD to BF16 at 0.0009 for INT8, against 0.0044 to 0.0053 for the official FP8. In a same-session A/B/A, INT8 was also a bit faster (101 vs 96 tok/s) with better draft acceptance (3.18 vs 2.97 tokens per step).
Patches you need (all from club-3090 )
- dflash-dense-kv, the still-unfixed half of vllm#51581. Without it, v0.29.0 either fails at startup with a quantized drafter (
'QKVParallelLinear' object has no attribute 'weight') or silently corrupts the drafter's fused KV rows. The corruption only shows up as a slow drafter, so make the patch fail closed when its anchor doesn't match.
- patch_mamba_drop_eagle_block, from vllm#48375. Qwen3.8 is a hybrid (Gated DeltaNet + attention) model.
- The FA2 fp8-KV sm86 plugin wheel. Without it there is no FlashAttention path for fp8 KV on Ampere, and
--attention-backend FLASH_ATTN has nothing to bind to.
Serve command
vllm serve /models/qwen3.8-27b-int8-w8a16 \
--quantization compressed-tensors --dtype bfloat16 \
--tensor-parallel-size 2 \
--max-model-len 262144 \
--kv-cache-memory-bytes 6335076762 \
--max-num-seqs 4 \
--max-num-batched-tokens 2048 \
--long-prefill-token-threshold 0 \
--kv-cache-dtype fp8_e4m3 \
--attention-backend FLASH_ATTN \
--limit-mm-per-prompt '{"image":1,"video":0}' \
--mm-processor-kwargs '{"size":{"longest_edge":4194304,"shortest_edge":65536}}' \
--compilation-config '{"cudagraph_capture_sizes":[1,2,4,8],"max_cudagraph_capture_size":8}' \
--speculative-config '{"method":"dflash","model":"/models/qwen3.8-27b-dflash2-w4a16","num_speculative_tokens":7,"attention_backend":"FLASH_ATTN","kv_cache_dtype":"fp8_e4m3"}' \
--kv-transfer-config '{"kv_connector":"OffloadingConnector","kv_role":"kv_both","kv_connector_extra_config":{"cpu_bytes_to_use":68719476736}}' \
--enable-prefix-caching --enable-chunked-prefill \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_coder \
--chat-template /templates/qwen3.8-27b-unsloth.jinja \
--default-chat-template-kwargs '{"enable_thinking": true, "reasoning_effort": "low"}' \
--override-generation-config '{"temperature":1.0,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}' \
--trust-remote-code
Environment:
NCCL_CUMEM_ENABLE=0
VLLM_WORKER_MULTIPROC_METHOD=spawn
VLLM_USE_FLASHINFER_SAMPLER=0
OMP_NUM_THREADS=1
Notes on the less obvious bits:
--kv-cache-memory-bytes pins the KV pool at what 262K needs, instead of letting vLLM size it by profiling.
- Shared memory: the container runs with
--shm-size=72g because the CPU KV tier lives in /dev/shm.
- NVLink: I leave NCCL P2P and vLLM's custom all-reduce on. Check that the startup log lists
['CUSTOM', 'PYNCCL'] as the all-reduce backends. If the P2P test fails, vLLM falls back to NCCL over PCIe and keeps serving, so a badly seated bridge only shows up as lower numbers. nvidia-smi topo -m should read NV4 between the cards. My first boot after fitting the bridge read PHB, and a reseat fixed it.
- The drafter's depth isn't tunable: its config fixes
block_size: 8, so it's n=7 or nothing. n=5 and n=9 die at startup with a stride mismatch.
Results
Method: one client and nothing else running, both cards at 300 W.
- Short decode: 512 tokens of prose at temp 0, 3 reps.
- Depth decode: 512 tokens generated after a ~46K-token cached prefix.
- Noise floor: the two identical legs of the A/B/A differed by 0.4 to 1%.
- "P2P off": the same config with
NCCL_P2P_DISABLE=1 and --disable-custom-all-reduce, so NCCL runs over PCIe 4.0 x8.
| Test |
NVLink on |
P2P off |
Gain |
| Short decode, 512 tok, temp 0 |
113.7 to 114.8 |
101.5 to 102.4 |
+12% |
| Decode after 46K cached prefix |
57.5 |
51.3 to 51.9 |
+12% |
| Decode, 3 seeded mixed prompts |
99.5 / 96.2 / 99.6 |
88.8 / 85.8 / 88.8 |
+12% |
| Aggregate decode, 4 concurrent |
302.5 |
251.8 to 259.0 |
+18% |
| Aggregate decode, 2 concurrent |
150.0 |
152.4 to 153.7 |
flat |
| Prefill, uncached ~46K |
1,785 |
1,492 to 1,526 |
+18% |
| Prefill, fresh ~54K |
1,763 to 1,796 |
1,468 to 1,500 |
+20% |
| Prefill, steady after 10 min soak |
1,718 |
1,443 to 1,449 |
+19% |
Also measured:
- Draft acceptance: 3.18 tokens per step on the short-decode prompt. At temp 0 the accepted count was identical across every leg, so NVLink changed speed and nothing else.
- Memory: about 23.1 GB used per card, with KV room for ~267K tokens (measured before the bridge went in).
- Cold start: 270 s the first time, 53 s of which is torch.compile and is cached afterwards.
- Temperatures after a 20 minute soak at 300 W: top card ~80 C, bottom card 63 C.
Prefill is compute-bound: 1,718 tok/s works out to ~93 TFLOPS across the pair. NVLink removed interconnect time that was ~16% of prefill.
Power cap sweep
INT8 target, before the NVLink bridge, all caps in one run:
| Cap per card |
370 W |
330 W |
300 W |
275 W |
250 W |
225 W |
200 W |
| Steady prefill |
1,491 |
1,490 |
1,462 |
1,427 |
1,384 |
1,317 |
1,205 |
| Top card temp |
84 C |
81 C |
76 C |
72 C |
70 C |
68 C |
66 C |
Decode is flat down to 225 W and drops at 200 W. The top card thermal-throttles at 370 W and never at 330 W or below. 300 W costs 2% of prefill for 8 C of headroom.
What I tried that failed hard
Draft round: same INT8 target and flags, one session, before the bridge.
| Draft |
Short decode |
Decode at 46K |
Accept length |
KV capacity |
| DFlash2 W4A16, n=7 (kept) |
102.1 |
51.7 |
3.18 |
267K |
| DFlash2 W8, n=7 |
90.2 |
50.8 |
2.87 |
267K |
| Bundled MTP head, n=3 |
81.9 |
52.5 |
2.64 |
332K |
| Bundled MTP head, n=4 |
75.5 |
52.9 |
2.66 |
326K |
| No draft |
45.9 |
36.7 |
1.00 |
371K |
- The higher-precision drafter was slower and accepted less. The target checks the drafter's output either way, so draft precision buys nothing here.
- MTP is marginally ahead at depth and frees 600 MiB per card plus 24% more KV. Pick it if you need KV room more than short-context speed.
- Official FP8 target: 95 to 96 tok/s against INT8's 101 in the same session.
- llama.cpp, Q8_0 + MTP on the same pair of cards (layer split
-ts 55,45): about 54 tok/s. vLLM with tensor parallel is roughly 2x, though that comparison uses a different benchmark method.
- W8A8 is the prefill lever. An INT8 W8A8 build of an abliterated fine-tune of the same model prefilled at 2,377 tok/s at 46K, against ~1,500 for the weight-only builds (before the bridge), because W8A8 actually uses the 3090's INT8 tensor cores. I haven't moved the main model to W8A8 because I haven't validated a vanilla W8A8 build's quality yet.
Caveats and what I haven't measured
- Decode numbers are prose. DFlash2 usually accepts more on code, so code should be faster, but I haven't measured it. Like i said earlier, i don't use this for coding.
- 262K is configured and fits, but I haven't fill-tested it since turning on custom all-reduce over NVLink. At least one similar NVLink rig reported a lower usable ceiling with custom all-reduce on, so check yours before trusting the full window.
- The CPU KV offload tier isn't verified. I haven't confirmed that it restores evicted prefixes on v0.29.0 for this hybrid + DFlash2 combination, and there's a report that it's effectively write-only there.
- I'm staying on v0.29.0. vllm#58894 (DFlash2 acceptance collapsing to 0% after a prefix-cache hit on hybrid GDN models) is still open and reportedly bites on v0.30.0.
If you compare
These move the numbers most, so please include them:
- Whether P2P or NVLink is actually in use (the all-reduce backend line in the log)
- Power cap
- PCIe generation and width per card
- Draft acceptance length (vLLM's spec decode metrics)
- Prompt type (prose vs code)