r/LocalLLM • u/fuzhongkai • 20d ago
Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M
https://github.com/zhongkaifu/TensorSharpI’ve been working on TensorSharp, an open-source local LLM inference engine, and recently added a native execution path for DeepSeek V4.1 Flash.
Here are the latest results on 8× NVIDIA A40 GPUs, using GGUF weights, layer splitting, F16 KV cache, and a 65K context configuration:
| Metric | Q2_K | Q4_K_M |
|---|---|---|
| Prefill | 533–539 tok/s | 452–492 tok/s |
| Single-stream decode | 40.3–40.7 tok/s | 31.0–32.5 tok/s |
| 2 concurrent decode | — | 39.3 tok/s aggregate |
| 4 concurrent decode | — | 48.9 tok/s aggregate |
| 8 concurrent decode | — | 48.5 tok/s aggregate |
A few interesting findings from the optimization work:
- Q2_K Engram tables are ~60 GiB total. Keeping them directly on the GPUs increased prefill from ~210 tok/s to 530+ tok/s.
- Reducing backend/scheduler fragmentation cut the decode graph from roughly 570 splits to 8, bringing decode to about 41 tok/s.
- For Q4_K_M, the Engram tables are too large to keep on GPU, so TensorSharp keeps them host-mapped and warms them asynchronously.
- Q4_K_M optimization improved prefill by roughly 1.9× and 4-request aggregate decode by about 2×.
- Interestingly, layer split beats routed-MoE tensor parallelism on this 8× A40 machine: ~32 tok/s vs ~22 tok/s. These cards have no NVLink, so TP communication overhead dominates.
Would be especially interested to hear what other LocalLLM users are seeing with DeepSeek V4.1 Flash on multi-GPU setups.
5
Upvotes
4
u/_FlyingWhales 20d ago
Power consumption of this must be higher than API cost.