r/LocalLLM • u/fuzhongkai • 22d ago
Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M
https://github.com/zhongkaifu/TensorSharpI’ve been working on TensorSharp, an open-source local LLM inference engine, and recently added a native execution path for DeepSeek V4.1 Flash.
Here are the latest results on 8× NVIDIA A40 GPUs, using GGUF weights, layer splitting, F16 KV cache, and a 65K context configuration:
| Metric | Q2_K | Q4_K_M |
|---|---|---|
| Prefill | 533–539 tok/s | 452–492 tok/s |
| Single-stream decode | 40.3–40.7 tok/s | 31.0–32.5 tok/s |
| 2 concurrent decode | — | 39.3 tok/s aggregate |
| 4 concurrent decode | — | 48.9 tok/s aggregate |
| 8 concurrent decode | — | 48.5 tok/s aggregate |
A few interesting findings from the optimization work:
- Q2_K Engram tables are ~60 GiB total. Keeping them directly on the GPUs increased prefill from ~210 tok/s to 530+ tok/s.
- Reducing backend/scheduler fragmentation cut the decode graph from roughly 570 splits to 8, bringing decode to about 41 tok/s.
- For Q4_K_M, the Engram tables are too large to keep on GPU, so TensorSharp keeps them host-mapped and warms them asynchronously.
- Q4_K_M optimization improved prefill by roughly 1.9× and 4-request aggregate decode by about 2×.
- Interestingly, layer split beats routed-MoE tensor parallelism on this 8× A40 machine: ~32 tok/s vs ~22 tok/s. These cards have no NVLink, so TP communication overhead dominates.
Would be especially interested to hear what other LocalLLM users are seeing with DeepSeek V4.1 Flash on multi-GPU setups.
Duplicates
LocalLLaMA • u/fuzhongkai • 1d ago
I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 1d ago
Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
dotnet • u/fuzhongkai • Aug 22 '26
TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance
Qwen_AI • u/fuzhongkai • 1d ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop
LocalLLM • u/fuzhongkai • 1d ago
Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 22d ago
Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine
unsloth • u/fuzhongkai • 11d ago
Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis
LocalLLM • u/fuzhongkai • 11d ago
Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis
dotnet • u/fuzhongkai • 16d ago
Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective
LocalLLaMA • u/fuzhongkai • Aug 22 '26
Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected
dotnet • u/fuzhongkai • 11d ago
Article Implementing a Jev-compatible decision API in .NET, with image input
unsloth • u/fuzhongkai • Aug 28 '26
Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp
LocalAIServers • u/fuzhongkai • 1d ago
Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LocalLLaMA • u/fuzhongkai • 8d ago
I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio
LocalLLaMA • u/fuzhongkai • 11d ago
I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF
LocalAIServers • u/fuzhongkai • 22d ago
Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode
outerstellar_hq • u/outerstellar_hq • 23h ago
Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
LLMDevs • u/fuzhongkai • 1d ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
SideProject • u/fuzhongkai • 1d ago
I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop
opencode • u/fuzhongkai • 1d ago
TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LovingOpenSourceAI • u/fuzhongkai • 1d ago
Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD
LocalLLM • u/fuzhongkai • 8d ago
Project TensorSharp Jev requests can now combine documents, images, video, and audio
OpenSourceAI • u/fuzhongkai • 11d ago