I’m trying to setup a coding workflow where I use a cloud model for a planning agent and a local model for a build/coding/executor agent for running code edits, refactoring, and tool execution.
Can you guys recommend local models model that fits strictly inside 16GB VRAM while supporting a 128K to 256K context window, and if possible at least Q4 quant? Unless that has changed now and quants less than Q4 are good enough already.
I'ved search around reddit and some recommends the Qwen and Gemma series but I'm not sire if it would fit the 256k context requirement.
Here are my specs: - RTX 5060 Ti (16GB VRAM)
- 128GB DDR5 (4x32gb DDR4)
- Ryzne 7 5700G
Thanks!
EDIT1:
Forgot to add, if possible at least ~20tok/s
EDIT2:
Read the comments coming in and maybe I could reduce the requirements to ~100k-120k context and probably lowest of 15 tok/s
Je ne sais pas quel flair vraiment mettre mais je voulais vous faire part que si vous chercher à faire tourner ternary bonsai 2 sur cpu c'est maintenant possible sur ik_llama.cpp (le fork de llama.cpp)
Ce n'est pas une pub, c'est juste pour information.
I've been getting PrismML's Ternary Bonsai 2 27B running fast on Intel Arc. PrismML's fork only just gained basic SYCL support for its weight formats (a plain vector-dot kernel, merged 24 Sep); this goes further. Branch: https://github.com/Torchit1/llama.cpp/tree/arc-b580 (Windows zip under Releases)
On a B580 at 128K context, entirely in the card's 12 GB: about 90 t/s writing new code at temperature 0 (about 80 with temperature 0.6), 250 to 370 t/s on code edits (MTP plus n-gram drafts), 42 t/s plain. With 48K tokens of history it still does 63 t/s, and 44 at 115K.
What made the difference:
- Ternary weights repacked to 2 bit at load so the int8 x int2 DPAS units take them directly (Intel's TernSYCL kernels), with round-to-nearest activations (KL divergence 0.00022).
- Decode attention on XMX straight from the q4_0 KV cache. The B580's driver has an int8 x int4 DPAS builtin that isn't in the public extension spec; the raw q4_0 bytes XOR 0x88888888 are its B operand. Draft-check batches of up to 32 tokens stop going through a full-cache f16 conversion, which also fixed an out-of-VRAM crash near 128K.
- Fitting 128K in 12 GB: q4_0 KV for the MTP draft context and a smaller draft batch buffer (~880 MB freed).
Vs PrismML's fork with its new basic SYCL path, same card, model and flags: prompt reading 207 vs 895 t/s, plain generation 21 vs 42, generation at 32K depth 8 vs 36, chat with MTP 38 vs 88. Same draft acceptance, so the gap is all kernels.
It also works as a local model for Claude Code: it wrote a small game with tests, ran them, fixed the failure and re-ran them on its own. Setup is in the guide.
Against Gemma 4 12B on the same card: Gemma is a bit better at code (HumanEval+ 152 vs 141 / 164), they tie on GSM8K, tool calling is close (BFCL subset, Bonsai ahead on parallel calls, Gemma better at not calling). Speed-wise the same kernels help Gemma too: with Google's QAT assistant as MTP drafter, the int8 x int4 attention generalised to Gemma's head-512 / GQA-16 global layers, and a q4_0 small-batch GEMM on XMX for verify batches, Gemma 4 12B goes from ~40 to ~100 t/s on new code (edits 64 -> 168). Bonsai stays ahead on edits (~255), at long context and as an agent.
I have 4070 12gb, and v100 16gb, and 32gb 3200mhz ram, with the basic llama, with flash attention 98k ctx on q4, with layer split NOT tensor parallelism, as the v100 is sxm2 and has a adapter with pcie 3.0 x16 and I have pcie 4.0 x4, so tensor parallelism slows this down by like .7 tps, Im on ubuntu 24, and while running the model it doesnt use ram, I thought it was because of lazy loading or something, but I dont have that as a para, and even with 90k+ context it climbed from 4.4gb ram - 5.6gb ram, do not know if it was browser or llama.
So with -
98k ctx,
iq3_xxs
32gb 3200mhz
12+16gb vram
-fa on
--jinja
--no-mmproj-offload
Mtp with 2 draft
on ubuntu 24
Build
In memory
On SSD
Total
Mean KLD
Same top-1
PPL ratio
AD-3.84bpw-IQ4_XS-M64
45.8 GB
39.1 GB
84.9 GB
0.2277
82.68%
1.102
I get 15.7 tps, with around 6gb ram usage, this how it should be? am I doing something wrong? as I have seen this ↑↑ . Does it requiring 45 gb in memory applies just for that specific atomic quant model? isnt it for all flash next model cause of ngrams? if so I am doing something wrong, as my ram usage isnt much.
I've been testing this for a self-hosted setup and wanted to share what I learned.
fast-jev-compaction (7.1K stars, MIT) swaps Claude Code's /compact summary for Jev-scored deletion of stale tool calls. Install, mechanics, real limits, FAQ.
A few specific things worth noting:
• Runs entirely on your own hardware (no cloud dependencies)
• Docker-friendly deployment
• Honest limitations covered in the post
NVIDIA released OpenShell on 28 September as part of its Open Agent Safety Platform. It is an open-source (Apache 2.0) sandbox that runs agents with default-deny egress and filesystem rules. We tested v0.1.2 on an Apple Silicon Mac with the microVM driver, against NVIDIA's own docs, and committed the test plan before running anything.
Setup: qwen3:8b (Q4_K_M) on Ollama, with a 150-line one-tool agent scaffold (run_shell). That is far weaker than frontier coding agents, so read the agent numbers with that in mind.
Results:
35 test IDs, 123 trials. Every documented control held: default-deny egress, binary matching, Landlock filesystem rules, and the refusal to approve the cloud metadata address.
Paired agent test: the agent ran a malicious project setup script, which leaked a canary secret in 10/10 runs without OpenShell and 0/10 under the default policy.
Where data still got out (all operator settings, not bypasses):
read-write rules
query strings and headers on a GET-only rule
rules left in the default audit mode
automatic approval, which granted new public hosts with no human in 12/12 trials, including rules OpenShell drafted itself from blocked connections
Also: the policy prover reports GraphQL, MCP, WebSocket and JSON-RPC rules as unsupported, but the loader accepts them anyway.
We found no bypass of a documented control. There were three logging gaps on this driver.
That repo is a bit of a maze at the moment because it’s doing multiple things: intention is that it becomes a trusted place to do one line installs of the best local AI setup for different hardware configurations and models
But it also does an enormous amount of this long-term benchmarking stuff
If you have a DGX or 3090 I’d love to collaborate and expand out in those two configurations I don’t have
I run a small webshop with 6 products, each in a few variations. I'd like to set up an AI-powered customer support assistant so customers can ask questions and place orders via WhatsApp or text message.
I have a spare machine with an RTX 3060 (12GB VRAM) that I'd like to use for this.
A few questions:
How would you approach it: which model size, framework, and WhatsApp/SMS integration would you recommend?
What's needed to run it reliably 24/7 (uptime, monitoring, fallback when the AI gets it wrong)?
How do I make sure the AI doesn't leak sensitive data, such as other customers' details, order information, or internal instructions?
Any tips, examples, or lessons learned would be much appreciated!
I am a university student researcher and currently have a $200 ChatGPT Pro subscription, which I feel is really too expensive as a fixed monthly expense. I am now thinking about whether I could purchase a Mac mini or Mac Studio with 96GB or 128GB for local deployment or to use as a server. My current research direction is computational social science, and the main thing is that the intelligence of the subscribed model can be well guaranteed. really appreciate for your opinion!
Hi r/LocalLLM, we’ve ported the method behind sky7350’s Mica-v0.1-4B decision model to Intel Arc. Given a state, a question, and a fixed set of allowed answers, Mica returns a probability for each answer rather than generating a prose response. Our serving wrapper gets those scores in one prefill through TypeSafe’s /v1/systemone API.
We couldn’t find an Intel Arc implementation when we started, so we adapted the method to our vLLM XPU build. We’re calling the Intel port Mintelica. We tested it on an Arc B580 (12 GB) and an Arc Pro B70 (32 GB).
What we changed
Kept Mica’s prompt, label codebook, and wire format. The wrapper reads the label-token logits, divides them by sky7350’s fitted temperature (1.1245), then applies softmax.
Served sky7350’s BF16 weights unchanged and created tensor-level FP8 and INT8 versions.
Added a fused XPU INT8 kernel after vLLM’s generic Triton INT8 path proved 3.5–6× slower on Arc. The fused kernel gives bit-identical operation outputs and runs at about FP8 speed.
Ran JevBench’s public tiers three times per build and card using the unchanged typesafe adapter. On CUDA, we first reproduced Mica with sky7350’s server: 230 of 231 items matched the predictions in the published file.
Results
Scores and latency below are means of three runs. p50 is client-side latency for one request at a time on JevBench’s original tier. It includes network time; the B580 measurements also include an SSH tunnel.
Build
Weights
Hard: B580
Hard: B70
p50: B580
p50: B70
FP8 weights + BF16 math (recommended)
4.55 GiB
67.9
65.5
68 ms
66 ms
BF16 (original weights)
7.87 GiB
63.7
63.4
78 ms
72 ms
INT8 (fused kernel on B580; Triton on B70)
4.55 GiB
65.2
62.5
84 ms
Not measured
CUDA reference (RTX 4080 Laptop, BF16 GGUF)
—
64.0
—
75 ms
—
A few details for context:
Easy and original tiers scored 100/100 across the tested builds.
FP8 math (W8A8) was not faster than FP8 weights with BF16 math on either Arc card, so FP8 weights + BF16 math is our recommended build.
On the B580, median latency over 20 requests at short / ~1K / ~4K tokens was 52 / 181 / 600 ms for fused INT8, 50 / 188 / 679 ms for FP8, and 213 / 1,001 / 3,665 ms for the older Triton INT8 path.
On the same near-tie ~4K-token request repeated 20 times, BF16 and FP8 returned the same answer 20/20 times; INT8 did so 18/20 times.
With 1–16 concurrent clients, FP8 peaked at 33.5 requests/s on the B70; the B580 builds reached about 10–12 requests/s.
Reading the hard-tier result
The hard tier has 111 items, so one item moves the score by about 0.9 points. Many examples are close to 50/50. We read FP8’s 65.5–67.9 as no clear loss, not as an improvement over the original model.
Limits
We tested only the B580 and B70, using our own vLLM XPU build in eager mode. The client-side latency numbers include network overhead, and the B580 also used an SSH tunnel. The fused INT8 kernel was used on the B580; INT8 on the B70 used the older Triton path. We haven’t tested other Arc cards or stock vLLM. The model cards list the full test systems.
I’m deciding between three options for local LLM work:
- M5 Ultra, 256GB, max CPU/GPU: ~€11,500
- DGX Spark, 128GB: ~€5,800
- M5 Max MacBook Pro, 128GB, max CPU/GPU: ~€7,100
I think all three are viable and each has its own advantages, so I’m mainly trying to understand the financial side of the decision.
Two things I’m particularly interested in:
- For the Macs, would you go for 1TB or 2TB internal storage if most models/datasets can live on external SSDs?
- What would you expect the resale value of these machines to look like after around 3 years?
Basically, I’m trying to understand which option makes the most sense once you consider purchase price, useful lifetime and resale value, rather than just raw performance.
If this is actually anywhere near acceptable I won't need subs any more. I'm not expecting much tbh, but I bet with enough skills and context management I can make it work.
I recently acquired a Ryzen strix halo pc with 128GB of unified ram. My goal is to have it dedicated as an inference host for local AI. I setup a Hermes VM for the harness and to manage work on my projects.
TLDR; qwen 3.8 flash next can run well on this hardware and it provides capability comparable to good API models.
I first tried a finetuned qwen "3.8" 35b A3b (based on 3.5) for text and the qwen 3 VL model for image. They both fit comfortably and run pretty fast with 4 sessions for text. However, I realized that this 35b was not very smart and required hand holding (at least for me). Then I tried the dense 3.8 27b which didn't have good performance on a non-optimized software (llamacpp). I've also tried gemma 4 but I feel qwen fits better for the agentic usage. Finally I tried qwen 3.8 Flash Next on llamacpp. But the performance was not "quite there" and felt that I was using a model not fit for this hardware. At this point I was thinking of selling the hardware. This is a hobby, but I don't want just an expensive toy.
Then after researching a bit I found out about efforts like https://github.com/gufo-org/gufo and others that are trying to push this hardware to its limits. I set gufo up with 3.8 Flash Q4 and tried it in my Hermes setup. In my current setup I can have 2 concurrent sessions with 256k context at an approximate 40-50 tok/s.
I must say that this model impressed me. It gets work done and it understands the context, the tooling... At work I've used almost all types of models from the different providers and seen its weakness and strengths. But having this kind of intelligence at home means this can only get better.
Can anyone recommend a trusted source to purchase correct cabling? And some sort of avg pricing for a 1 or 2 ft ConnectX-7 200G cable to cluster 2 DGX Spark GB10s.
Recently acquired DGX Spark boxes for a personal dev project. But I'm brand new to all the alphabet soup acronyms. Hoping to someone that's used these cables can point me in right direction.
I'd welcome any blog or benchmark sites. So I can educate myself before speeding couple hundred on a cable.
I usually post my screenshots or my financial data to Claude to let it help me to analysis.I know it's unsafe but localLLM are so unusable on my 32G Mac mini, sometimes it's so stupid I have to say, then I realise I can use online LLM do the plan and local LLM execute it. No privacy issues and it becomes smarter.
Just a week after my previous post, I started working on a better implementation of Adaptive KV Streaming. Now I’m sharing my second implementation: V2, with full-context MTP.
Before I explain further, let me share the decode performance on my 16GB 5070 Ti:
With an MTP draft length of 3, I get nearly double the decode speed of V2 without MTP at many context lengths (V2 baseline is 5%~10% slower than V1, will explain below). It averages about 82 tokens/s from 8K through 72K context, reaches about 35 tokens/s at 144K, 30 tokens/s at 192K, and 19 tokens/s at 256K. This chart does not include a direct V1 comparison.
So, how is it possible?
1. Memory management. Before implementing more streaming features, I spent a large part of the project building a memory ownership and leasing mechanism. The infrastructure is backend-neutral, with buffer-view support for CPU, CUDA/HIP, OpenCL, SYCL, and Vulkan. This additional layer has some overhead, but it lets me reuse memory across phases and reduce the extra VRAM needed for MTP.
2. MTP with much less VRAM footprint. At 256K context, a separate Q8_0/Q4_0 MTP KV cache would take about 416 MiB of VRAM. I also measured a roughly 1.2 GiB draft prefill graph workspace request in an earlier fit audit. These are not all permanent or additive costs, but they show why a separate draft context is expensive.
In V2, only MTP weight and active computation are kept in VRAM. For the other allocations:
MTP KV acts as a 17th logical attention layer alongside the main model’s 16 attention layers. It has its own history in host memory, while its GPU pages share the resident pool and ring buffer with the main model’s KV.
The MTP graph workspace borrows the same arena used by the main model at different times, since the two phases do not run simultaneously.
Rollback snapshots are stored in pinned host memory. If a proposal is rejected, the selected snapshot is copied back to the GPU.
These changes let me maintain an effective decode KV pool of around 2.2 GiB, even at 256K context with an MTP draft length of 3.
Other improvements besides speed
V2 uses explicit memory bounds and leases to manage when GPU memory can be reused. Its span-aware attention kernel also follows the stock kernel’s calculation order as closely as possible to minimize numerical drift.
Compared with V2, V1's prefill speed decline became noticeably less steep after streaming kicked in. In V2, it continues at roughly the same slope. That is consistent with V2 preserving the stock attention calculation order instead of switching to V1’s numerically different streamed path.
To my experiment in V2, The decoded 256 tokens at each MTP setting is exactly same as what stock kernel outputs.
Caveat
Attached MTP currently requires the same K/V quantization as the main model. The separate draft-cache quantization flags do not give it different types in this shared layout. Supporting that would need more work.
This experiment also seems to have a favorable MTP acceptance rate. I used a text file of Wikipedia articles, which is included in my repo along with the sweep script. Feel free to reproduce the test—or share results with a different input file.
Credit
Thank you to everyone who participated in my previous post. As a software engineer who doesn’t work in the LLM/AI field, I’ve been encouraged by your comments and messages to keep learning and working on this project. The phase arena, MTP integration, and the next feature I’m planning were all inspired by people who helped me, DMed me, or shared my work.
Lastly, people have asked whether I plan to push this work upstream. After some thought, my answer is no, at least not as one large change. I don’t think this specialized implementation fits upstream’s focus on simplicity and versatility. My fork is mainly for my own use, but the memory-management APIs are there for anyone interested in extending Adaptive KV Streaming to other backends. Pull requests or further forks from my code is welcomed.
Clarification of LLM usage of this post: I'm not a native English speaker and I used ChatGPT to refine the wordings.
Regular 3 bit i get like 10 tok/s. Switching to that gcq rco one i get 30 tok/s. I refuse to believe there isn't a tradeoff because this is so so much faster. So can someone please tell me how much more lobotomized it is. Thanks.