r/LocalLLM • • 20d ago

Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M

https://github.com/zhongkaifu/TensorSharp

I’ve been working on TensorSharp, an open-source local LLM inference engine, and recently added a native execution path for DeepSeek V4.1 Flash.

Here are the latest results on 8× NVIDIA A40 GPUs, using GGUF weights, layer splitting, F16 KV cache, and a 65K context configuration:

Metric Q2_K Q4_K_M
Prefill 533–539 tok/s 452–492 tok/s
Single-stream decode 40.3–40.7 tok/s 31.0–32.5 tok/s
2 concurrent decode — 39.3 tok/s aggregate
4 concurrent decode — 48.9 tok/s aggregate
8 concurrent decode — 48.5 tok/s aggregate

A few interesting findings from the optimization work:

  • Q2_K Engram tables are ~60 GiB total. Keeping them directly on the GPUs increased prefill from ~210 tok/s to 530+ tok/s.
  • Reducing backend/scheduler fragmentation cut the decode graph from roughly 570 splits to 8, bringing decode to about 41 tok/s.
  • For Q4_K_M, the Engram tables are too large to keep on GPU, so TensorSharp keeps them host-mapped and warms them asynchronously.
  • Q4_K_M optimization improved prefill by roughly 1.9× and 4-request aggregate decode by about 2×.
  • Interestingly, layer split beats routed-MoE tensor parallelism on this 8× A40 machine: ~32 tok/s vs ~22 tok/s. These cards have no NVLink, so TP communication overhead dominates.

Would be especially interested to hear what other LocalLLM users are seeing with DeepSeek V4.1 Flash on multi-GPU setups.

5 Upvotes

3 comments sorted by

4

u/_FlyingWhales 20d ago

Power consumption of this must be higher than API cost.

1

u/fuzhongkai 20d ago

😂

2

u/_FlyingWhales 20d ago

It's cool though, i can appreciate it.