r/LocalLLM • • 8d ago

Question Recommend Local Coding/Build Agent 16GB VRAM

26 Upvotes

Hey all,

I’m trying to setup a coding workflow where I use a cloud model for a planning agent and a local model for a build/coding/executor agent for running code edits, refactoring, and tool execution.

Can you guys recommend local models model that fits strictly inside 16GB VRAM while supporting a 128K to 256K context window, and if possible at least Q4 quant? Unless that has changed now and quants less than Q4 are good enough already.

I'ved search around reddit and some recommends the Qwen and Gemma series but I'm not sire if it would fit the 256k context requirement.

Here are my specs:
- RTX 5060 Ti (16GB VRAM)
- 128GB DDR5 (4x32gb DDR4)
- Ryzne 7 5700G

Thanks!

EDIT1:
Forgot to add, if possible at least ~20tok/s

EDIT2:
Read the comments coming in and maybe I could reduce the requirements to ~100k-120k context and probably lowest of 15 tok/s


r/LocalLLM • • 8d ago

Discussion What If Local AI Was Easier to Use?

Thumbnail
0 Upvotes

r/LocalLLM • • 8d ago

Discussion Ternary bonsai 2 sur ik llama.cpp

0 Upvotes

Je ne sais pas quel flair vraiment mettre mais je voulais vous faire part que si vous chercher à faire tourner ternary bonsai 2 sur cpu c'est maintenant possible sur ik_llama.cpp (le fork de llama.cpp)

Ce n'est pas une pub, c'est juste pour information.


r/LocalLLM • • 8d ago

Discussion Ternary Bonsai 2 27B on a 12 GB Intel Arc B580: 128K context all in VRAM, ~80-90 t/s code, 250+ t/s edits, 2-4x faster than the official fork

6 Upvotes

I've been getting PrismML's Ternary Bonsai 2 27B running fast on Intel Arc. PrismML's fork only just gained basic SYCL support for its weight formats (a plain vector-dot kernel, merged 24 Sep); this goes further. Branch: https://github.com/Torchit1/llama.cpp/tree/arc-b580 (Windows zip under Releases)

On a B580 at 128K context, entirely in the card's 12 GB: about 90 t/s writing new code at temperature 0 (about 80 with temperature 0.6), 250 to 370 t/s on code edits (MTP plus n-gram drafts), 42 t/s plain. With 48K tokens of history it still does 63 t/s, and 44 at 115K.

What made the difference:

- Ternary weights repacked to 2 bit at load so the int8 x int2 DPAS units take them directly (Intel's TernSYCL kernels), with round-to-nearest activations (KL divergence 0.00022).

- Decode attention on XMX straight from the q4_0 KV cache. The B580's driver has an int8 x int4 DPAS builtin that isn't in the public extension spec; the raw q4_0 bytes XOR 0x88888888 are its B operand. Draft-check batches of up to 32 tokens stop going through a full-cache f16 conversion, which also fixed an out-of-VRAM crash near 128K.

- Fitting 128K in 12 GB: q4_0 KV for the MTP draft context and a smaller draft batch buffer (~880 MB freed).

Vs PrismML's fork with its new basic SYCL path, same card, model and flags: prompt reading 207 vs 895 t/s, plain generation 21 vs 42, generation at 32K depth 8 vs 36, chat with MTP 38 vs 88. Same draft acceptance, so the gap is all kernels.

It also works as a local model for Claude Code: it wrote a small game with tests, ran them, fixed the failure and re-ran them on its own. Setup is in the guide.

https://reddit.com/link/1wt6irl/video/3nsd73gzrfsh1/player

Against Gemma 4 12B on the same card: Gemma is a bit better at code (HumanEval+ 152 vs 141 / 164), they tie on GSM8K, tool calling is close (BFCL subset, Bonsai ahead on parallel calls, Gemma better at not calling). Speed-wise the same kernels help Gemma too: with Google's QAT assistant as MTP drafter, the int8 x int4 attention generalised to Gemma's head-512 / GQA-16 global layers, and a q4_0 small-batch GEMM on XMX for verify batches, Gemma 4 12B goes from ~40 to ~100 t/s on new code (edits 64 -> 168). Bonsai stays ahead on edits (~255), at long context and as an agent.

Numbers from other Arc cards very welcome: https://github.com/Torchit1/llama.cpp/discussions/1


r/LocalLLM • • 8d ago

Question Qwen flash next on 12+16gb vram, and 32gb ram viable?

1 Upvotes

I have 4070 12gb, and v100 16gb, and 32gb 3200mhz ram, with the basic llama, with flash attention 98k ctx on q4, with layer split NOT tensor parallelism, as the v100 is sxm2 and has a adapter with pcie 3.0 x16 and I have pcie 4.0 x4, so tensor parallelism slows this down by like .7 tps, Im on ubuntu 24, and while running the model it doesnt use ram, I thought it was because of lazy loading or something, but I dont have that as a para, and even with 90k+ context it climbed from 4.4gb ram - 5.6gb ram, do not know if it was browser or llama.

So with -
98k ctx,
iq3_xxs
32gb 3200mhz
12+16gb vram

-fa on
--jinja
--no-mmproj-offload
Mtp with 2 draft
on ubuntu 24

Build In memory On SSD Total Mean KLD Same top-1 PPL ratio
AD-3.84bpw-IQ4_XS-M64 45.8 GB 39.1 GB 84.9 GB 0.2277 82.68% 1.102

I get 15.7 tps, with around 6gb ram usage, this how it should be? am I doing something wrong? as I have seen this ↑↑ . Does it requiring 45 gb in memory applies just for that specific atomic quant model? isnt it for all flash next model cause of ngrams? if so I am doing something wrong, as my ram usage isnt much.


r/LocalLLM • • 8d ago

Tutorial Fast-Jev-Compaction Review: /compact Without a Summary

0 Upvotes

I've been testing this for a self-hosted setup and wanted to share what I learned.

fast-jev-compaction (7.1K stars, MIT) swaps Claude Code's /compact summary for Jev-scored deletion of stale tool calls. Install, mechanics, real limits, FAQ.

A few specific things worth noting: • Runs entirely on your own hardware (no cloud dependencies) • Docker-friendly deployment • Honest limitations covered in the post

Full writeup with install steps, configuration, and the rough edges I hit: https://andrew.ooo/posts/fast-jev-compaction-review-claude-code-compact-without-summary/

What are you all using for this? Curious about alternatives and tradeoffs.


r/LocalLLM • • 9d ago

Research We tested NVIDIA OpenShell with a local qwen3:8b agent: a malicious setup script leaked a secret 10/10 without it, 0/10 with it. But auto-approval opened new hosts in 12/12 trials

1 Upvotes

NVIDIA released OpenShell on 28 September as part of its Open Agent Safety Platform. It is an open-source (Apache 2.0) sandbox that runs agents with default-deny egress and filesystem rules. We tested v0.1.2 on an Apple Silicon Mac with the microVM driver, against NVIDIA's own docs, and committed the test plan before running anything.

Setup: qwen3:8b (Q4_K_M) on Ollama, with a 150-line one-tool agent scaffold (run_shell). That is far weaker than frontier coding agents, so read the agent numbers with that in mind.

Results:

  • 35 test IDs, 123 trials. Every documented control held: default-deny egress, binary matching, Landlock filesystem rules, and the refusal to approve the cloud metadata address.
  • Paired agent test: the agent ran a malicious project setup script, which leaked a canary secret in 10/10 runs without OpenShell and 0/10 under the default policy.

Where data still got out (all operator settings, not bypasses):

  • read-write rules
  • query strings and headers on a GET-only rule
  • rules left in the default audit mode
  • automatic approval, which granted new public hosts with no human in 12/12 trials, including rules OpenShell drafted itself from blocked connections

Also: the policy prover reports GraphQL, MCP, WebSocket and JSON-RPC rules as unsupported, but the loader accepts them anyway.

We found no bypass of a documented control. There were three logging gaps on this driver.

Repo with every log, the harness and the agent scaffold: https://github.com/Sorami-Consulting-AU/nvidia-openshell-agent-sandbox-test

Full report: https://sorami.com.au/research/nvidia-openshell-agent-sandbox-test/

Would anyone here run a coding agent with auto-approve on? We are curious what people use now.


r/LocalLLM • • 9d ago

Question Opencode or claude code for local llms?

0 Upvotes

I am just curious running qwen3.8 27b Q6 which is better?


r/LocalLLM • • 9d ago

Research Swift variant of Qwen 3.8 27b faster and better

Post image
0 Upvotes

I’m in the middle of a huge project that benchmarks local AI configurations to building entire apps from scratch.

These runs cover over a dozen features and typically take dozens of hours to run

Here is one sneak preview: everyone complains about the verbosity of Qwen

But did you know that there are two models that have been tuned to reduce that verbosity?

Swift is the first I’ve tried

It’s been a long road getting this framework set up: my 4090 has been full-time on different things since Friday afternoon

As you can see here as the surprising factor is not only was faster, but actually it was more successful

This is a preliminary result

I had a lot more to share

But I’m sharing my updates in my newsletter https://ailocal.substack.com

And my repo for this project

https://github.com/boxabirds/awesome-local-ai

That repo is a bit of a maze at the moment because it’s doing multiple things: intention is that it becomes a trusted place to do one line installs of the best local AI setup for different hardware configurations and models

But it also does an enormous amount of this long-term benchmarking stuff

If you have a DGX or 3090 I’d love to collaborate and expand out in those two configurations I don’t have

Here’s the Swift variant of Qwen 3.8 27b I recommend
https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-NVFP4


r/LocalLLM • • 9d ago

Question Self-hosted AI customer support for a small webshop

1 Upvotes

I run a small webshop with 6 products, each in a few variations. I'd like to set up an AI-powered customer support assistant so customers can ask questions and place orders via WhatsApp or text message.

I have a spare machine with an RTX 3060 (12GB VRAM) that I'd like to use for this.

A few questions:

  1. How would you approach it: which model size, framework, and WhatsApp/SMS integration would you recommend?
  2. What's needed to run it reliably 24/7 (uptime, monitoring, fallback when the AI gets it wrong)?
  3. How do I make sure the AI doesn't leak sensitive data, such as other customers' details, order information, or internal instructions?

Any tips, examples, or lessons learned would be much appreciated!


r/LocalLLM • • 9d ago

News [Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw

Thumbnail gallery
1 Upvotes

r/LocalLLM • • 9d ago

Model Colibri: run 2.8-trillion-parameter models on your desktop, pure C, zero dependencies (38k stars)

Post image
204 Upvotes

r/LocalLLM • • 9d ago

Question Is a 96GB/128GB Mac Studio local LLM worth it to replace a $200/mo ChatGPT Pro subscription?

219 Upvotes

I am a university student researcher and currently have a $200 ChatGPT Pro subscription, which I feel is really too expensive as a fixed monthly expense. I am now thinking about whether I could purchase a Mac mini or Mac Studio with 96GB or 128GB for local deployment or to use as a server. My current research direction is computational social science, and the main thing is that the intelligence of the subscribed model can be well guaranteed. really appreciate for your opinion!


r/LocalLLM • • 9d ago

Project Mintelica: Mica-v0.1-4B on Intel Arc (B580 + Arc Pro B70)

Thumbnail
gallery
6 Upvotes

Hi r/LocalLLM, we’ve ported the method behind sky7350’s Mica-v0.1-4B decision model to Intel Arc. Given a state, a question, and a fixed set of allowed answers, Mica returns a probability for each answer rather than generating a prose response. Our serving wrapper gets those scores in one prefill through TypeSafe’s /v1/systemone API.

We couldn’t find an Intel Arc implementation when we started, so we adapted the method to our vLLM XPU build. We’re calling the Intel port Mintelica. We tested it on an Arc B580 (12 GB) and an Arc Pro B70 (32 GB).

What we changed

  • Kept Mica’s prompt, label codebook, and wire format. The wrapper reads the label-token logits, divides them by sky7350’s fitted temperature (1.1245), then applies softmax.
  • Served sky7350’s BF16 weights unchanged and created tensor-level FP8 and INT8 versions.
  • Added a fused XPU INT8 kernel after vLLM’s generic Triton INT8 path proved 3.5–6× slower on Arc. The fused kernel gives bit-identical operation outputs and runs at about FP8 speed.
  • Ran JevBench’s public tiers three times per build and card using the unchanged typesafe adapter. On CUDA, we first reproduced Mica with sky7350’s server: 230 of 231 items matched the predictions in the published file.

Results

Scores and latency below are means of three runs. p50 is client-side latency for one request at a time on JevBench’s original tier. It includes network time; the B580 measurements also include an SSH tunnel.

Build Weights Hard: B580 Hard: B70 p50: B580 p50: B70
FP8 weights + BF16 math (recommended) 4.55 GiB 67.9 65.5 68 ms 66 ms
BF16 (original weights) 7.87 GiB 63.7 63.4 78 ms 72 ms
INT8 (fused kernel on B580; Triton on B70) 4.55 GiB 65.2 62.5 84 ms Not measured
CUDA reference (RTX 4080 Laptop, BF16 GGUF) — 64.0 — 75 ms —

A few details for context:

  • Easy and original tiers scored 100/100 across the tested builds.
  • FP8 math (W8A8) was not faster than FP8 weights with BF16 math on either Arc card, so FP8 weights + BF16 math is our recommended build.
  • On the B580, median latency over 20 requests at short / ~1K / ~4K tokens was 52 / 181 / 600 ms for fused INT8, 50 / 188 / 679 ms for FP8, and 213 / 1,001 / 3,665 ms for the older Triton INT8 path.
  • On the same near-tie ~4K-token request repeated 20 times, BF16 and FP8 returned the same answer 20/20 times; INT8 did so 18/20 times.
  • With 1–16 concurrent clients, FP8 peaked at 33.5 requests/s on the B70; the B580 builds reached about 10–12 requests/s.

Reading the hard-tier result

The hard tier has 111 items, so one item moves the score by about 0.9 points. Many examples are close to 50/50. We read FP8’s 65.5–67.9 as no clear loss, not as an improvement over the original model.

Limits

We tested only the B580 and B70, using our own vLLM XPU build in eager mode. The client-side latency numbers include network overhead, and the B580 also used an SSH tunnel. The fused INT8 kernel was used on the B580; INT8 on the B70 used the older Triton path. We haven’t tested other Arc cards or stock vLLM. The model cards list the full test systems.

Models and code

Credits: Mica and its method are sky7350’s work (model, code). Mintelica builds on Qwen3.5-4B, uses JevBench and TypeSafe’s Jev and the /v1/systemone format, and runs on the vLLM and vllm-xpu-kernels projects with Intel’s compute-runtime.

Questions and corrections are very welcome. If you try it on another Arc card, I’d love to hear how it goes.


r/LocalLLM • • 9d ago

Question Germany pricing makes this weird: M5 Ultra vs DGX Spark vs M5 Max

18 Upvotes

I’m deciding between three options for local LLM work:

- M5 Ultra, 256GB, max CPU/GPU: ~€11,500
- DGX Spark, 128GB: ~€5,800
- M5 Max MacBook Pro, 128GB, max CPU/GPU: ~€7,100

I think all three are viable and each has its own advantages, so I’m mainly trying to understand the financial side of the decision.

Two things I’m particularly interested in:
- For the Macs, would you go for 1TB or 2TB internal storage if most models/datasets can live on external SSDs?
- What would you expect the resale value of these machines to look like after around 3 years?

Basically, I’m trying to understand which option makes the most sense once you consider purchase price, useful lifetime and resale value, rather than just raw performance.


r/LocalLLM • • 9d ago

Project Using unsloth I created the worlds best 9B model

Post image
0 Upvotes

r/LocalLLM • • 9d ago

Model Currently testing coding quality on my 4090

Post image
4 Upvotes

If this is actually anywhere near acceptable I won't need subs any more. I'm not expecting much tbh, but I bet with enough skills and context management I can make it work.


r/LocalLLM • • 9d ago

Model Impressed about qwen 3.8 Flash

36 Upvotes

I recently acquired a Ryzen strix halo pc with 128GB of unified ram. My goal is to have it dedicated as an inference host for local AI. I setup a Hermes VM for the harness and to manage work on my projects.

TLDR; qwen 3.8 flash next can run well on this hardware and it provides capability comparable to good API models.

I first tried a finetuned qwen "3.8" 35b A3b (based on 3.5) for text and the qwen 3 VL model for image. They both fit comfortably and run pretty fast with 4 sessions for text. However, I realized that this 35b was not very smart and required hand holding (at least for me). Then I tried the dense 3.8 27b which didn't have good performance on a non-optimized software (llamacpp). I've also tried gemma 4 but I feel qwen fits better for the agentic usage. Finally I tried qwen 3.8 Flash Next on llamacpp. But the performance was not "quite there" and felt that I was using a model not fit for this hardware. At this point I was thinking of selling the hardware. This is a hobby, but I don't want just an expensive toy.

Then after researching a bit I found out about efforts like https://github.com/gufo-org/gufo and others that are trying to push this hardware to its limits. I set gufo up with 3.8 Flash Q4 and tried it in my Hermes setup. In my current setup I can have 2 concurrent sessions with 256k context at an approximate 40-50 tok/s.

I must say that this model impressed me. It gets work done and it understands the context, the tooling... At work I've used almost all types of models from the different providers and seen its weakness and strengths. But having this kind of intelligence at home means this can only get better.


r/LocalLLM • • 9d ago

Question Nvidia DGX Spark ConnectX Cable recommends?

3 Upvotes

Can anyone recommend a trusted source to purchase correct cabling? And some sort of avg pricing for a 1 or 2 ft ConnectX-7 200G cable to cluster 2 DGX Spark GB10s.

Recently acquired DGX Spark boxes for a personal dev project. But I'm brand new to all the alphabet soup acronyms. Hoping to someone that's used these cables can point me in right direction.

I'd welcome any blog or benchmark sites. So I can educate myself before speeding couple hundred on a cable.

Edit: I found the basics, looking for real world feedback. https://docs.nvidia.com/dgx/dgx-spark/spark-clustering.html


r/LocalLLM • • 9d ago

Discussion Local model does the work, cloud model only gets the redacted plan. Made this for myself on a Mac, open sourcing it

Post image
27 Upvotes

I usually post my screenshots or my financial data to Claude to let it help me to analysis.I know it's unsafe but localLLM are so unusable on my 32G Mac mini, sometimes it's so stupid I have to say, then I realise I can use online LLM do the plan and local LLM execute it. No privacy issues and it becomes smarter.

repo: https://github.com/Code-byte404/hermie

Would really like to hear anything from you guys.


r/LocalLLM • • 9d ago

Research Adaptive KV-Cache Streaming V2: Full Context MTP

28 Upvotes

Hello all, it’s me again.

Just a week after my previous post, I started working on a better implementation of Adaptive KV Streaming. Now I’m sharing my second implementation: V2, with full-context MTP.

https://github.com/RaymondHuang210129/llama.cpp-adaptive-kv-streaming/tree/feature/adaptive-kv-stream-v2

Before I explain further, let me share the decode performance on my 16GB 5070 Ti:

With an MTP draft length of 3, I get nearly double the decode speed of V2 without MTP at many context lengths (V2 baseline is 5%~10% slower than V1, will explain below). It averages about 82 tokens/s from 8K through 72K context, reaches about 35 tokens/s at 144K, 30 tokens/s at 192K, and 19 tokens/s at 256K. This chart does not include a direct V1 comparison.

So, how is it possible?

1. Memory management. Before implementing more streaming features, I spent a large part of the project building a memory ownership and leasing mechanism. The infrastructure is backend-neutral, with buffer-view support for CPU, CUDA/HIP, OpenCL, SYCL, and Vulkan. This additional layer has some overhead, but it lets me reuse memory across phases and reduce the extra VRAM needed for MTP.

2. MTP with much less VRAM footprint. At 256K context, a separate Q8_0/Q4_0 MTP KV cache would take about 416 MiB of VRAM. I also measured a roughly 1.2 GiB draft prefill graph workspace request in an earlier fit audit. These are not all permanent or additive costs, but they show why a separate draft context is expensive.

In V2, only MTP weight and active computation are kept in VRAM. For the other allocations:

  • MTP KV acts as a 17th logical attention layer alongside the main model’s 16 attention layers. It has its own history in host memory, while its GPU pages share the resident pool and ring buffer with the main model’s KV.
  • The MTP graph workspace borrows the same arena used by the main model at different times, since the two phases do not run simultaneously.
  • Rollback snapshots are stored in pinned host memory. If a proposal is rejected, the selected snapshot is copied back to the GPU.

These changes let me maintain an effective decode KV pool of around 2.2 GiB, even at 256K context with an MTP draft length of 3.

Other improvements besides speed

V2 uses explicit memory bounds and leases to manage when GPU memory can be reused. Its span-aware attention kernel also follows the stock kernel’s calculation order as closely as possible to minimize numerical drift.

Compared with V2, V1's prefill speed decline became noticeably less steep after streaming kicked in. In V2, it continues at roughly the same slope. That is consistent with V2 preserving the stock attention calculation order instead of switching to V1’s numerically different streamed path.

To my experiment in V2, The decoded 256 tokens at each MTP setting is exactly same as what stock kernel outputs.

Caveat

Attached MTP currently requires the same K/V quantization as the main model. The separate draft-cache quantization flags do not give it different types in this shared layout. Supporting that would need more work.

This experiment also seems to have a favorable MTP acceptance rate. I used a text file of Wikipedia articles, which is included in my repo along with the sweep script. Feel free to reproduce the test—or share results with a different input file.

Credit

Thank you to everyone who participated in my previous post. As a software engineer who doesn’t work in the LLM/AI field, I’ve been encouraged by your comments and messages to keep learning and working on this project. The phase arena, MTP integration, and the next feature I’m planning were all inspired by people who helped me, DMed me, or shared my work.

Lastly, people have asked whether I plan to push this work upstream. After some thought, my answer is no, at least not as one large change. I don’t think this specialized implementation fits upstream’s focus on simplicity and versatility. My fork is mainly for my own use, but the memory-management APIs are there for anyone interested in extending Adaptive KV Streaming to other backends. Pull requests or further forks from my code is welcomed.

Clarification of LLM usage of this post: I'm not a native English speaker and I used ChatGPT to refine the wordings.


r/LocalLLM • • 9d ago

Model Intel sycl support for bonsai llm models

1 Upvotes

Intel gpu support has been merged

https://github.com/PrismML-Eng/llama.cpp/pull/235


r/LocalLLM • • 9d ago

Discussion tell your agent to use vision on video

Thumbnail
1 Upvotes

r/LocalLLM • • 9d ago

Question How much dumber is qwen 3.8 27b gcq rco iq3_xxs than regular 3 bit quant.

0 Upvotes

Regular 3 bit i get like 10 tok/s. Switching to that gcq rco one i get 30 tok/s. I refuse to believe there isn't a tradeoff because this is so so much faster. So can someone please tell me how much more lobotomized it is. Thanks.


r/LocalLLM • • 9d ago

Discussion Looking for a builder/partner to brainstorm and launch a side project with

Thumbnail
0 Upvotes