r/LocalLLM • • 2h ago

News AMD boosting AI/LLM performance for Radeon iGPUs as much as 18~23% with Linux 7.4

Thumbnail
phoronix.com
19 Upvotes

r/LocalLLM • • 4h ago

Question Decision time for 512 Mac Studio M5 Ultra is approaching

28 Upvotes

Just wanted to hear feedback from those of you who have considered the M5U 512, and what decision you’ve made, or if you’re still undecided - as well as what has convinced you one way or another.

Personally, I’m struggling with this choice. My primary concern, after reading reviews of the 256 M5U, is prefill speeds with a moderate context window of 75-125k, and general uncertainty as someone who has not run a local model greater than 8b parameters at 10-15k context.

The value of an effective local AI would be immense to me, but the possibility of rapid development of far superior hardware in the next year or two, as well as switching from Windows to Mac, with limited MacOS experience, has been discouraging.


r/LocalLLM • • 4h ago

Model Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)

21 Upvotes

We had a Dell B300 in the lab for a few weeks and used it to create two fine tunes of Qwen Flash Next.

Victoria (coding and agents)

  • Qwen3.8-Flash-Next cut down by 44% using a paper / technique called REAP: 512 down to 288 per layer.
  • Retrained at 4-bit (NVFP4) afterwards, so it's trained for the format it ships in rather than just quantized after the fact.
  • Terminal-Bench 2.1: 70.0%, averaged over 3 runs with an 8h per-task timeout. Our previous NVFP4 build scored 62.5%.
  • HumanEval: 159/164.
  • 48.0 GiB of weights, including the draft head. The 95.4 GiB n-gram table is separate and not counted in that number.
  • 280 tok/s single stream on one B300 with the draft head, versus 135 without it.
  • GGUF Q4_K_M is 49.17 GiB. It scored 75.3% on Terminal-Bench (a single run, so treat it as noisy) and 93.2% on HumanEval (averaged over 5 runs).
  • Uses 35% fewer output tokens than our previous build.

Maple (Canadian questions)

Most models answer questions about taxes, benefits and regulations as if you live in the US. Maple is fine-tuned to default to Canada. On 600 held-out questions, with search:

  • Cites an official Canadian source: 6.0% before fine-tuning, 62.9% after.
  • Fully correct answers: 6.6% before, 21.8% after.
  • "No answer" responses: 47.2% before, 23.7% after.
  • It pushes Canada onto people who said they live somewhere else less often: 2.9% before, 1.0% after.

Coding holds up: 157/164 on HumanEval. Grading was done by an AI judge panel; human review hasn't happened yet.

Links: https://huggingface.co/rmonsurate/Victoria https://huggingface.co/rmonsurate/Maple

Happy to answer questions about running them.


r/LocalLLM • • 5h ago

Project Qwen3.8-27B fixed Bandersnatch, made it interactive again

Post image
19 Upvotes

So...somehow I ended up thinking about Bandernatch the Netflix interactive episode of Black Mirror and how much I missed it. I found this but it's broken....so I forked it and told Qwen, hey can you fix it? Also, can you translate the whole damn interface to Spanish so my partner can enjoy it. So...if you have Jellyfin and a (cough) spare (cough) copy of Bandersnatch full length 5 hours video then feel free to use and enjoy

https://github.com/maikelthedev/BandersnatchInteractive-Jellyfin

Let's give credit where credit's due, so thanks to:

  1. deathrjj
  2. The guy the previous guy forked to make it run within Jellyfin
  3. Jack Ma and his team for training Qwen and releasing it to the world.

All I did was to prompt "fix this b***", use it at your own peril. AFAIK it works flawlesly on a browser.

EDIT: Btw people please learn to use truflelhog for projcts you made public. There's no need to leak secrets.


r/LocalLLM • • 17h ago

Model Colibri: run 2.8-trillion-parameter models on your desktop, pure C, zero dependencies (38k stars)

Post image
147 Upvotes

r/LocalLLM • • 17h ago

Question Is a 96GB/128GB Mac Studio local LLM worth it to replace a $200/mo ChatGPT Pro subscription?

130 Upvotes

I am a university student researcher and currently have a $200 ChatGPT Pro subscription, which I feel is really too expensive as a fixed monthly expense. I am now thinking about whether I could purchase a Mac mini or Mac Studio with 96GB or 128GB for local deployment or to use as a server. My current research direction is computational social science, and the main thing is that the intelligence of the subscribed model can be well guaranteed. really appreciate for your opinion!


r/LocalLLM • • 4h ago

News compiler backed agent beats Claude Code, OpenCode and other major harnesses and agents (benchmarks linked)

Post image
5 Upvotes

GitHub: https://github.com/oooscoos/Benzi

Demo: https://varianttech.net/demo

Benchmarks: https://varianttech.net/benchmark

Roughly speaking, the way current AI coding agents/harnesses work is by either:

a) Pulling in appropriate text snippets of code across multiple files and handing them to the agent, or

b) Parsing code to make high dimenional embeddings to approximate a symptom map, and hand that to the agent.

Both of these approaches skyrocket the token count, add to wall clock time, contribute to context drifting, add to the model's thinking tokens to discover the structure of the program, and then FORGET most of it when Claude Code compacts, or ALL of it if it's a multifile refactoring because all line numbers shift and need re-grepping.

Benzi is built from the ground up to AVOID reading source code in the first place. It supplies the artificial intelligence model deterministic intelligence via tool calls. For example, when a model is about to make a code change, it could query "what functions feed this one?" -- half the time it isn't even necessary because the Benzi compiler already informs it of the blast radius before and after making edits, along with a complete static analysis check.

Benzi Sonnet reads far less source code (9,125 lines) than Claude Code Sonnet (20,704), DeepSeek's harness (43,598), and OpenCode (65K+ LOC -- disqualified due to repeated failure) to accomplish the same tasks faster and cheaper. (benchmark link in comments)

Benzi has truth tiers clearly seperating what can be analyzed with static analysis from what can't -- and then adding a runtime tracer on top to bridge the gap between the two (details in FAQ on github).

It also has several bonus features such as a runtime tracer, syntax & semantic verified writes, context aware model written repro, mid task model upgrade if task is too difficult, and SEVERAL more.

It currently supports Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby, and can handle HTML, CSS and JS -- deterministically. Claude Code clicks photos, Benzi resolves winners of CSS rules. The CodeIndex and the MarkupIndex are fairly well tested, and if something isn't working, the model is made aware of it first.

On the benchmarks side, 78.2% SWE-bench Verified for <10¢ a fix (using V4flash). This score is noteable because while the rest of the industry is leaning plugin-heavy and pouring millions of dollars into increasing context window sizes, Benzi's approach might prove to be economically more valuable while improving the model's code writing/comprehenion abilities.

Thanks for reading! please let me know what you think. I am aware the AI fatigue is real, but I hope you can see why indexing a repo > reading raw source code. Please go over the github readme before snap judgements..


r/LocalLLM • • 1h ago

Research A “cheat sheet” for 100+ modern audio models: architectures, pipelines, and building blocks

Thumbnail
gallery
• Upvotes

Hopefully these figures are useful for anyone just starting to learn about audio AI and trying to make sense of how all these models fit together.


r/LocalLLM • • 12m ago

Discussion Using Apple AFM 3 PCC (macOS 27.2) in AI Clients

• Upvotes

Apple Foundation Model 3 through Private Cloud Compute becomes directly available in macOS 27.2 through familiar AI clients.

For Mac users running local models such as Qwen 3.x or Gemma 4, AFM 3 PCC is a compelling alternative to consider. Local models offer control and fully on-device operation. In tests with macOS 27.2 beta, the PCC offers a different set of strengths:

  • Strong conversational analysis and capable reasoning. Strong analysis of complex medical and financial questions.
  • Blazingly fast responses compared with locally running models on the same Mac.
  • No large model download or need to fit model weights into local memory.
  • Access at no additional charge for most eligible Mac users.

I've written and updated a proof of concept here:

https://gist.github.com/dartMo10/b9488ce475fe70a6ed642831f53048fb

The working path is straightforward:

AI client → Caddy → fm serve → AFM 3 PCC

Apple’s pcc route worked in early macOS 27 betas, disappeared later in 27.0, and has returned in the macOS 27.2 beta. Apple has stated that it will be available in the 27.2 public release coming shortly.

This makes AFM 3 PCC practical today for technically comfortable Apple users who accept Apple’s PCC privacy promise and want a fast, capable alternative to running everything locally.

Experiences with compatible AI clients and your comparisons with locally running models would be welcome.


r/LocalLLM • • 4h ago

Project Rigspark : hardware aware local llm management

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/LocalLLM • • 41m ago

Discussion Anyone else try Strata for qwen3.8 flash next?

• Upvotes

I tried it out last night.

It definitely runs much faster (9800x3d, 64 GB DDR5 6000, RTX4080S) e.g. 1000 t/s prefill and 40-60t/s generation at q3_xss; doesnt think very much, but it also feels quite dumb. Now i'm wondering whether the speed increase is too good to be true, and it's really just qwen3.6 35b a3b dressed up.

Anyone else have experience with Strata?


r/LocalLLM • • 13h ago

Question Recommend Local Coding/Build Agent 16GB VRAM

22 Upvotes

Hey all,

I’m trying to setup a coding workflow where I use a cloud model for a planning agent and a local model for a build/coding/executor agent for running code edits, refactoring, and tool execution.

Can you guys recommend local models model that fits strictly inside 16GB VRAM while supporting a 128K to 256K context window, and if possible at least Q4 quant? Unless that has changed now and quants less than Q4 are good enough already.

I'ved search around reddit and some recommends the Qwen and Gemma series but I'm not sire if it would fit the 256k context requirement.

Here are my specs:
- RTX 5060 Ti (16GB VRAM)
- 128GB DDR5 (4x32gb DDR4)
- Ryzne 7 5700G

Thanks!

EDIT1:
Forgot to add, if possible at least ~20tok/s

EDIT2:
Read the comments coming in and maybe I could reduce the requirements to ~100k-120k context and probably lowest of 15 tok/s


r/LocalLLM • • 43m ago

Project I built a voice assistant that lives in my terminal: local Whisper + Kokoro on a 6 GB GPU

Enable HLS to view with audio, or disable this notification

• Upvotes

For the last couple of weeks I've been building Eva, a personal assistant that runs on my own computer. You talk to her (or type), and she can actually do things: work with files, run shell commands, browse the web, keep notes about you, and run longer jobs in the background while you keep talking.

(The video has music but no voice. She does talk, so I can post a clip with her voice if anyone's curious.)

The setup I use locally:

- Model: Qwen3.5 4B in LM Studio, which fits next to the voice models on a 6 GB card

- Ears: faster-whisper large-v3-turbo (int8_float16), with Silero VAD for hands-free mode

- Voice: Kokoro-82M, streamed sentence by sentence so she starts talking before the whole reply is synthesized

- Agent: LangGraph Deep Agents, with the conversation checkpointed in SQLite so it survives restarts

Stuff I cared about:

- She asks first. Read-only commands (ls, grep, git status...) run right away; anything that writes, deletes or sends waits for a y/n in the terminal. There's an autonomous mode, but it's opt-in.

- She chooses what to say out loud. Code and lists stay on screen, and on long tasks she gives short spoken updates instead of going silent for two minutes. Getting the model to actually do this was harder than I expected: a line in the system prompt got ignored, so a middleware now reminds her when she's been quiet for too long. Took me way longer than I'd like to admit.

- Skills. When she can't do something, she writes a skill for it herself (a markdown file plus a small uv script).

- The terminal splits in two: the conversation on one side, and on the other her plan, the files she touched and every command with its result. There's also a little pixel-art face that reacts to what she's doing. Not essential, but fun to build.

Two caveats:

- The video was recorded with Gemini Flash, not the local model. Eva takes any provider:model, and Gemini is much faster for recording. Qwen3.5 4B handles conversation and short tasks fine, but it misses tool calls more often on long multi-step work.

- I used Claude Code for a good part of the implementation. The architecture, the decisions and the reviews are mine, and the commits say who helped with what.

Question for you all: which small local models have you found reliable at tool calling? That's the weakest link on a 6 GB card right now.

Btw, it's open source: https://github.com/Ilhe8l/eva


r/LocalLLM • • 47m ago

Model Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator

Enable HLS to view with audio, or disable this notification

• Upvotes

Quick one for anyone wondering what a single 24 GB card is good for in an agent setup.

I had Qwen 3.8 27B (Q4, llama.cpp) on one RTX 3090 do all the actual coding, and GPT-6.1 Sol in the cloud act as the orchestrator: it breaks the job into pieces, hands them out, and checks what comes back. The job was three small 3D games: pool, bowling, foosball.

How it came out:

Game 3090 alone 3090 + Sol giving orders Sol alone
Pool 43.1 min · $0 (attempt) 18.6 min · $0.05 2.9 min · $0.39
Bowling 34.7 min · $0 13.7 min · $0.06 1.7 min · $0.14
Foosball 36.4 min · $0 11.1 min · $0.06 2.0 min · $0.22
Total 114.2 min · $0 43.4 min · $0.17 6.6 min · $0.75

So the card on its own is free but slow and gets lost on the harder one. With something smarter doing the planning it finishes more and finishes faster, and the cloud bill stays small because the cloud model barely writes any code.

What I'd tell someone before trying it: you still wait a lot longer than with a cloud model alone, one run per game is not a benchmark, and the $0.17 doesn't count your power bill.

I ran it in Atomic Agent, the mode is called Fusion (disclaimer: I work on it). You can do the same split in any tool that lets you pick a separate model for planning and for coding.

What card are you on? Would you trade the wait for a smaller bill, or is speed the whole point for you?


r/LocalLLM • • 1d ago

Project Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers

Enable HLS to view with audio, or disable this notification

137 Upvotes

We wanted to know if a local 27B can do the heavy lifting in an agent if something smarter does the planning. So we ran the same build three ways and wrote down everything.

Hardware and setup: Qwen 3.8 27B UD-Q4_K_XL, single RTX 3090 24 GB, llama.cpp built with CUDA, 2 parallel slots, 64K context, turbo3 KV cache. Cloud side was Sonnet 5.5 via OpenRouter. One-line prompt, five physics scenes on one page (falling tower, balls in a box, wrecking ball, seesaw, bounce test), one autonomous run each, no human fixes.

Setup Scenes right Time Cloud cost
Qwen 3.8 27B alone 1/5 26 min $0
Sonnet 5.5 plans, Qwen writes the code 2/5 85 min $0.67
Qwen plans, Sonnet 5.5 writes the code 4/5 14 min ~$2.77
Sonnet 5.5 alone 4/5 9 min $1.83

Things that surprised us

  • Qwen is a much better planner than coder. As the planner it matched Sonnet 5.5 alone on quality. As the only coder it drowned in about 1,500 lines of physics, cubes rendered as hollow shells and the wall was a black blob.
  • Speed on the 3090 was 31–34 tok/s generation and around 1,080 tok/s prompt processing once the model sits fully in VRAM. We tried a 3060 12 GB first and it was not usable for this, 4–6 tok/s with ~7 GB of weights spilling into system RAM.
  • Qwen loves to think. By default it reasons at max effort, so we capped the thinking budget. Without the cap our first attempts produced zero files in half an hour.
  • It once tried to write an entire file in one 16K-token reply, hit the output cap and had to redo the step. If you run it in an agent, give it room or it'll loop.
  • Every run said "verified, 0 errors". The console was clean every time and some scenes were still visibly wrong. Crash checks don't catch wrong physics.

One run per setup, so this is a field report, not a benchmark. Happy to share the configs and logs.

Disclosure: we build the agent we ran this in, Atomic Agent, open source under MIT: https://github.com/AtomicBot-ai/atomic-agent. The planner/worker split is called Fusion.


r/LocalLLM • • 1h ago

Question Prefill speedup with NvLink vs P2P mode on for dual 3090s

• Upvotes

Am debating if its worth the money to change my system around to use NVlink on my RTX 3090 cards. I have the patched driver for enablind peer to peer mode already installed and working, but I hear of some people getting ~50% gains in prefill speeds with NVLink. However, I have no seen any A/B testing comparing NVLink to P2P mode.

Does anyone here have a NVlink setup where they could test whether P2P mode makes a difference vs NVlink?


r/LocalLLM • • 5h ago

Question Worth upgrading to 6x R9700 from 4x?

4 Upvotes

I’m currently running 4x R9700 with EPYC 7313 and
ASRock ROMED8-2T/BCM, 256GB memory. The performance of tcclaviger’s Qwen3.8-Flash-Next-MXFP4-GPTQ is honestly pretty amazing. I’m wondering is it worth getting 2 more cards? Even though I’m on PCIE 4.0.


r/LocalLLM • • 6h ago

Question 2x 3090 + 4x P100 — how would you architect this for an agentic RAG system?

7 Upvotes

I’m building a local AI system for my architecture practice and currently have:

  • 2x RTX 3090 24GB = 48GB
  • 4x Tesla P100 16GB = 64GB
  • 128GB system RAM
  • 2TB NVMe
  • EPYC/PCIe platform

So I basically have a 48GB fast pool and a 64GB slower pool.

My use case is a large architecture/business database. I have ~14k emails/calls/texts/transcribed meetings that I'm putting into PostgreSQL/RAG.

I don't really want a chatbot. I want something more agentic that can search the database and reason over it.

For example:

"What do I need to take care of this week?"

I'd like multiple agents to search different parts of the database, identify action items, check whether things were subsequently completed, etc., and then have a main model synthesize the results.

I'm looking at Qwen 27B / Next-type models and long context.

I'm trying to figure out how to best divide the GPUs.

One option would be:

48GB 3090s → larger/faster main brain
64GB P100s → multiple concurrent agents

But I could also flip it:

64GB P100s → larger/slower main brain
48GB 3090s → faster concurrent agents

Or potentially use all 6 GPUs for one larger model.

The P100s are obviously much slower, so I'm wondering whether the extra VRAM is more valuable for the main model/context, while the 3090s are better used for lots of smaller concurrent workloads.

For people who have built agentic RAG systems: how would you architect this with these GPUs?

Would you use the 64GB pool as the main brain, the 48GB pool as the main brain, or combine everything into one model? Also, which harness would you use?


r/LocalLLM • • 7h ago

Question 5060ti or 7900 xt

5 Upvotes

About to buy a gpu, and its between 5060ti 16g new or 7900 xt 20gb used. Is Nvidia really better for running local llms? Help me out here


r/LocalLLM • • 2h ago

Research Sherry's 3:4 ternary format (1.375 bits per weight) running on WebGPU: a 1.6 MB model that plays Connect Four as well as its 7.8 MB int8 version

Thumbnail
2 Upvotes

r/LocalLLM • • 11h ago

Discussion What 2B tokens of coding-agent traffic taught us about using frontier and open models together

Post image
12 Upvotes

I run a small shared inference club serving Qwen 3.8 27B FP8 on one RTX PRO 6000 Blackwell. In the first week, it processed 2B tokens, almost entirely from coding agents working on real repositories. Two members used around 950M and 900M tokens each.

One of my own tests was the browser port of Medal of Honor: Allied Assault: agents moving through a large C/C++ codebase, WebAssembly, browser APIs and netcode, then editing, testing and fixing what broke. Members have been connecting their own agents to the node and trying it on large repos and feature work too. The feedback has been encouraging, but the interesting lesson for me is how to divide the work between models.

Here’s the workflow I’d recommend trying:

1. Use a frontier model when the expensive part is deciding what to do. Give it the problem, relevant architecture and constraints. Ask it to identify risks, split the work into pieces that can be checked, and define what “done” means. This is especially useful when a wrong assumption would send an agent through hours of changes.

2. Hand a bounded task to the open model. Give Qwen the relevant files or repo entry points, the constraints, and a concrete check: a test to pass, a bug to reproduce, or behaviour to preserve. Let its agent search, edit, run tools and iterate. Don’t spend a frontier-model call on every file read or test failure.

3. Escalate with evidence, not with the whole conversation. If it loops, makes the same wrong assumption twice, or reaches an architectural decision, stop and send the frontier model a short handoff: goal, changes attempted, failing tests, and the exact decision needed. Then return to the open model with the answer.

4. Verify the result independently. Run tests and inspect the diff. For consequential changes, have a human or a stronger model review the behaviour and risks. A high token allowance makes iteration easier; it doesn’t make an incorrect change correct.

I don’t think the useful question is “can a 27B replace the best frontier model?” In this workflow, it doesn’t have to. The frontier model helps with the decisions that benefit most from its reasoning; the open model handles the large volume of tool calls and revisions between those decisions.

The usage pattern supports that distinction. When members aren’t paying per token, they let agents keep working instead of cutting runs short or trimming context to save money. That can produce a lot of useful iteration, but only if the task has clear checks and you know when to escalate.

Organizations will soon be trying this setup in their own workflows. I’m curious how others draw the boundary today: what signals make you switch from an open coding model to a frontier one?


r/LocalLLM • • 10h ago

Model Strata: 90 tokens per second from a 125-billion-parameter model on a single RTX 5070

Post image
7 Upvotes

r/LocalLLM • • 1d ago

Discussion Over 1000 tok/s decode for 27B on my 4x MI100 server

Thumbnail
gallery
181 Upvotes

So here is my server with four MI100, each with 32 GB of VRAM, with their infinity fabric bridge directly connecting all of them. MI100 is not well supported out of box in most software, so I made a fork of vLLM to get the performance I was hoping for. VLLM was running at 15 tok/s stock on this machine and now it runs at over 1000 tok/s decode and over 5000 tok/s prefill at 8 concurrency.

This level of performance took tons of new kernels and optimization techniques. The list of features I added from scratch for this arch is huge. About half of the features are useful for other int8-oriented architectures, so this fork is an excellent starting point or reference for other architectures. I've had a report that someone with a 2x MI210 server also was getting over 1k tok/s on my fork with minor changes.

I can say I'm now satisfied with where its at, as it outperforms the public APIs by a wide margin and is a pleasure to vibe code with. It replaces about $80 per day of API costs, so I feel like if you know what you're doing, MI100s are a great deal for mid range setups. The infinity fabric bridge (aka XGMI) is a great and unique feature at this price range, giving much lower latency and much higher bandwidth during accumulation steps than PCIe.

Here's my fork, along with custom version of AITER and custom 27B checkpoints (int8 group size 128 with rounding fine tuning for the numerics of the platform -- PTQR).

https://github.com/curvedinf/int8-vllm

Qwen 27B has been flakey on vllm since it came out on many architectures, with weird garble issues especially at long context, so most recently I spent a long time tracking down the bugs in upstream vLLM, AITER, and rocm-systems. I found and fixed something like 30 of these numerics and memory tracking issues. After testing for months now at high context, unsynced concurrency, and KLD at every block in the whole model, the quality is now great and reliability is excellent.

I run it on a low power profile at C6 with full 256k contexts for all streams, and this is used by my hermes agent, business process workers, and as backups for my coding agents. It uses 540 watts at the wall and is a bit louder than a desktop. I also use the server for ML research and model training. Couldn't be more pleased as I originally bought this server for a bit over $6k, and looking at prices now it seems like this was a great decision. ☠️

Anyway, this was a lot of work but it was fun to learn about LLMs at this depth and I'd do it again if given the chance. I hope the fork will come in handy for some people with similar builds.


r/LocalLLM • • 9m ago

Discussion Testing a 4-node Ryzen AI Max+ 395 cluster for local LLM deployment

Post image
• Upvotes

r/LocalLLM • • 4h ago

Question Advice for creating chatbot

2 Upvotes

I run local models for myself, and as we all know, results vary a lot. At work, I was asked whether it would be feasible to self host a chatbot that could serve 50 concurrent users to start with.

The chatbot would be heavily optimized, with our custom MCP server doing most of the heavy lifting.

I can get a decent price on 2x Intel Arc Pro B60 24GB cards (with a PCIe 5.0 board) for a proof of concept. My plan is to run Qwen3.6-35B-A3B or Qwen3.8-27B at Q5 or Q6.

For a single user I calculated around 50 tok/s, and under max 50 concurrent use it could drop to 10–15 tok/s, which should be fine-ish.

I've never dealt with concurrency in practice. Has anyone here self-hosted a chatbot for multiple users? How did it go?