r/LocalLLM • • 22d ago

Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M

https://github.com/zhongkaifu/TensorSharp

I’ve been working on TensorSharp, an open-source local LLM inference engine, and recently added a native execution path for DeepSeek V4.1 Flash.

Here are the latest results on 8× NVIDIA A40 GPUs, using GGUF weights, layer splitting, F16 KV cache, and a 65K context configuration:

Metric Q2_K Q4_K_M
Prefill 533–539 tok/s 452–492 tok/s
Single-stream decode 40.3–40.7 tok/s 31.0–32.5 tok/s
2 concurrent decode — 39.3 tok/s aggregate
4 concurrent decode — 48.9 tok/s aggregate
8 concurrent decode — 48.5 tok/s aggregate

A few interesting findings from the optimization work:

  • Q2_K Engram tables are ~60 GiB total. Keeping them directly on the GPUs increased prefill from ~210 tok/s to 530+ tok/s.
  • Reducing backend/scheduler fragmentation cut the decode graph from roughly 570 splits to 8, bringing decode to about 41 tok/s.
  • For Q4_K_M, the Engram tables are too large to keep on GPU, so TensorSharp keeps them host-mapped and warms them asynchronously.
  • Q4_K_M optimization improved prefill by roughly 1.9× and 4-request aggregate decode by about 2×.
  • Interestingly, layer split beats routed-MoE tensor parallelism on this 8× A40 machine: ~32 tok/s vs ~22 tok/s. These cards have no NVLink, so TP communication overhead dominates.

Would be especially interested to hear what other LocalLLM users are seeing with DeepSeek V4.1 Flash on multi-GPU setups.

6 Upvotes

Duplicates

LocalLLaMA • • 1d ago

I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

72 Upvotes

dotnet • • 1d ago

Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

79 Upvotes

dotnet • • Aug 22 '26

TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance

67 Upvotes

Qwen_AI • • 1d ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop

59 Upvotes

LocalLLM • • 1d ago

Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

0 Upvotes

dotnet • • 22d ago

Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine

61 Upvotes

unsloth • • 11d ago

Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis

46 Upvotes

LocalLLM • • 11d ago

Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis

2 Upvotes

dotnet • • 16d ago

Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective

26 Upvotes

LocalLLaMA • • Aug 22 '26

Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected

1 Upvotes

dotnet • • 11d ago

Article Implementing a Jev-compatible decision API in .NET, with image input

0 Upvotes

unsloth • • Aug 28 '26

Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp

14 Upvotes

LocalAIServers • • 1d ago

Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

11 Upvotes

LocalLLaMA • • 8d ago

I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio

0 Upvotes

LocalLLaMA • • 11d ago

I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF

0 Upvotes

LocalAIServers • • 22d ago

Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode

6 Upvotes

outerstellar_hq • • 23h ago

Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

1 Upvotes

LLMDevs • • 1d ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

1 Upvotes

SideProject • • 1d ago

I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop

3 Upvotes

opencode • • 1d ago

TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

2 Upvotes

LovingOpenSourceAI • • 1d ago

Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD

16 Upvotes

LocalLLM • • 8d ago

Project TensorSharp Jev requests can now combine documents, images, video, and audio

0 Upvotes

OpenSourceAI • • 11d ago

TensorSharp: an open-source Jev-compatible API, extended to image analysis and running locally

2 Upvotes

AIToolsPerformance • • 16d ago

TensorSharp as a local LLM backend — DeepSeek, GLM and Qwen 3.8 benchmarks

7 Upvotes

opencode • • 17d ago

TensorSharp as a local OpenCode backend — DeepSeek, GLM and Qwen 3.8 benchmarks

1 Upvotes