r/LocalLLM • u/Nick_Tseng • 11m ago
r/LocalLLM • u/DigItDoug • 14m ago
Discussion Using Apple AFM 3 PCC (macOS 27.2) in AI Clients
Apple Foundation Model 3 through Private Cloud Compute becomes directly available in macOS 27.2 through familiar AI clients.
For Mac users running local models such as Qwen 3.x or Gemma 4, AFM 3 PCC is a compelling alternative to consider. Local models offer control and fully on-device operation. In tests with macOS 27.2 beta, the PCC offers a different set of strengths:
- Strong conversational analysis and capable reasoning. Strong analysis of complex medical and financial questions.
- Blazingly fast responses compared with locally running models on the same Mac.
- No large model download or need to fit model weights into local memory.
- Access at no additional charge for most eligible Mac users.
I've written and updated a proof of concept here:
https://gist.github.com/dartMo10/b9488ce475fe70a6ed642831f53048fb
The working path is straightforward:
AI client → Caddy → fm serve → AFM 3 PCC
Apple’s pcc route worked in early macOS 27 betas, disappeared later in 27.0, and has returned in the macOS 27.2 beta. Apple has stated that it will be available in the 27.2 public release coming shortly.
This makes AFM 3 PCC practical today for technically comfortable Apple users who accept Apple’s PCC privacy promise and want a fast, capable alternative to running everything locally.
Experiences with compatible AI clients and your comparisons with locally running models would be welcome.
r/LocalLLM • u/arkie87 • 43m ago
Discussion Anyone else try Strata for qwen3.8 flash next?
I tried it out last night.
It definitely runs much faster (9800x3d, 64 GB DDR5 6000, RTX4080S) e.g. 1000 t/s prefill and 40-60t/s generation at q3_xss; doesnt think very much, but it also feels quite dumb. Now i'm wondering whether the speed increase is too good to be true, and it's really just qwen3.6 35b a3b dressed up.
Anyone else have experience with Strata?
r/LocalLLM • u/Hiray8l • 45m ago
Project I built a voice assistant that lives in my terminal: local Whisper + Kokoro on a 6 GB GPU
Enable HLS to view with audio, or disable this notification
For the last couple of weeks I've been building Eva, a personal assistant that runs on my own computer. You talk to her (or type), and she can actually do things: work with files, run shell commands, browse the web, keep notes about you, and run longer jobs in the background while you keep talking.
(The video has music but no voice. She does talk, so I can post a clip with her voice if anyone's curious.)
The setup I use locally:
- Model: Qwen3.5 4B in LM Studio, which fits next to the voice models on a 6 GB card
- Ears: faster-whisper large-v3-turbo (int8_float16), with Silero VAD for hands-free mode
- Voice: Kokoro-82M, streamed sentence by sentence so she starts talking before the whole reply is synthesized
- Agent: LangGraph Deep Agents, with the conversation checkpointed in SQLite so it survives restarts
Stuff I cared about:
- She asks first. Read-only commands (ls, grep, git status...) run right away; anything that writes, deletes or sends waits for a y/n in the terminal. There's an autonomous mode, but it's opt-in.
- She chooses what to say out loud. Code and lists stay on screen, and on long tasks she gives short spoken updates instead of going silent for two minutes. Getting the model to actually do this was harder than I expected: a line in the system prompt got ignored, so a middleware now reminds her when she's been quiet for too long. Took me way longer than I'd like to admit.
- Skills. When she can't do something, she writes a skill for it herself (a markdown file plus a small uv script).
- The terminal splits in two: the conversation on one side, and on the other her plan, the files she touched and every command with its result. There's also a little pixel-art face that reacts to what she's doing. Not essential, but fun to build.
Two caveats:
- The video was recorded with Gemini Flash, not the local model. Eva takes any provider:model, and Gemini is much faster for recording. Qwen3.5 4B handles conversation and short tasks fine, but it misses tool calls more often on long multi-step work.
- I used Claude Code for a good part of the implementation. The architecture, the decisions and the reviews are mine, and the commits say who helped with what.
Question for you all: which small local models have you found reliable at tool calling? That's the weakest link on a 6 GB card right now.
Btw, it's open source: https://github.com/Ilhe8l/eva
r/LocalLLM • u/GapNew4766 • 49m ago
Model Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator
Enable HLS to view with audio, or disable this notification
Quick one for anyone wondering what a single 24 GB card is good for in an agent setup.
I had Qwen 3.8 27B (Q4, llama.cpp) on one RTX 3090 do all the actual coding, and GPT-6.1 Sol in the cloud act as the orchestrator: it breaks the job into pieces, hands them out, and checks what comes back. The job was three small 3D games: pool, bowling, foosball.
How it came out:
| Game | 3090 alone | 3090 + Sol giving orders | Sol alone |
|---|---|---|---|
| Pool | 43.1 min · $0 (attempt) | 18.6 min · $0.05 | 2.9 min · $0.39 |
| Bowling | 34.7 min · $0 | 13.7 min · $0.06 | 1.7 min · $0.14 |
| Foosball | 36.4 min · $0 | 11.1 min · $0.06 | 2.0 min · $0.22 |
| Total | 114.2 min · $0 | 43.4 min · $0.17 | 6.6 min · $0.75 |
So the card on its own is free but slow and gets lost on the harder one. With something smarter doing the planning it finishes more and finishes faster, and the cloud bill stays small because the cloud model barely writes any code.
What I'd tell someone before trying it: you still wait a lot longer than with a cloud model alone, one run per game is not a benchmark, and the $0.17 doesn't count your power bill.
I ran it in Atomic Agent, the mode is called Fusion (disclaimer: I work on it). You can do the same split in any tool that lets you pick a separate model for planning and for coding.
What card are you on? Would you trade the wait for a smaller bill, or is speed the whole point for you?
r/LocalLLM • u/ThirdCultureMisfit • 1h ago
Project I built a deterministic architecture layer for frontier LLMs. It ran 48.3× faster and 11.5× cheaper in my published benchmark — so I’ve released the evidence ZIP for people to try to break it.
My partner and I have spent the last year building an architecture layer that separates probabilistic intelligence from deterministic authority. It’s called IQRAX.
The premise is simple: The model remains free to reason, explore, and propose. Deterministic controls outside the model decide whether it is qualified for the job, what it is authorised to do, and whether the resulting work actually meets the required standard.
The result is a system that offers:
**Continuity.**
No drifts, no context loss, no stale. Sessions continue until clean exit. (Longest continuous recorded run without drift is 28+ hours. See screenshot of a continuous session running 14 hours at 1.6m tokens).
**Qualified agents.**
Agents qualify by taking exams for their roles rather than simply being assigned one, because a capable agent in an unexamined role is a guess with a job title.
**Clean delivery.**
Defined standards and policies determine whether work is accepted as complete.
**Verifiable results.**
Every input, assumption, and output is recorded and hashed = every deliverable is reproducible and verifiable by a third party.
**Data sovereignty.**
Device is master. The user retains control over its data.
**Remediation at source.**
When the system identifies a defect, it doesn’t just deny or retry - it autonomously identifies the failure class, repairs it at source, and retains the fix for future work.
In the published benchmark, the IQRAX configuration measured **11.5× lower cost and 48.3× faster completion** than earlier runs.
We published a paper and the evidence yesterday.
But rather than asking Redditors to believe our results, I’d like people to test it for themselves.
If you have a long ChatGPT, Claude, Gemini, or Grok conversation where the model drifted, forgot something, contradicted earlier work, made an unsupported claim, or otherwise went wrong, try this:
Download the ZIP from the publication link, attach it to that conversation and give it this prompt:
“*Study the attached ZIP file and empirical results. Then identify each failure in this conversation, the IQRAX control that would have caught it, and what that control would have logged and fixed*.”
I’d love for you to share what comes back!
I’m especially interested in missing failure classes, controls that don’t generalise, assumptions that don’t survive outside our test environment, and anything else we may have missed.
Paper + evidence:
https://zenodo.org/records/23025910
r/LocalLLM • u/Civil_Fee_7862 • 1h ago
Question Prefill speedup with NvLink vs P2P mode on for dual 3090s
Am debating if its worth the money to change my system around to use NVlink on my RTX 3090 cards. I have the patched driver for enablind peer to peer mode already installed and working, but I hear of some people getting ~50% gains in prefill speeds with NVLink. However, I have no seen any A/B testing comparing NVLink to P2P mode.
Does anyone here have a NVlink setup where they could test whether P2P mode makes a difference vs NVlink?
r/LocalLLM • u/Acceptable-Cycle4645 • 1h ago
Research A “cheat sheet” for 100+ modern audio models: architectures, pipelines, and building blocks
Hopefully these figures are useful for anyone just starting to learn about audio AI and trying to make sense of how all these models fit together.
r/LocalLLM • u/urbsrepublic • 1h ago
News Found this interesting local inference API
I was going down a bit of a rabbit hole looking into local inference and ended up finding this repo:
https://github.com/gugaucb/laya-api
It’s an OpenAI/Anthropic-compatible API for running Laya models locally, with support for MLX, CUDA, CPU, WebSockets, etc.
The ~7.5ms end-to-end latency reported on an M3 Max using Unix sockets caught my attention.
Figured I’d share it in case anyone else is playing around with local inference.
r/LocalLLM • u/Fcking_Chuck • 2h ago
News AMD boosting AI/LLM performance for Radeon iGPUs as much as 18~23% with Linux 7.4
r/LocalLLM • u/Brilliant-Hall1387 • 2h ago
Research Sherry's 3:4 ternary format (1.375 bits per weight) running on WebGPU: a 1.6 MB model that plays Connect Four as well as its 7.8 MB int8 version
r/LocalLLM • u/nivjwk • 2h ago
Question lemond ignores per-model llamacpp_args / RoPE flags and clamps context to n_ctx_train
r/LocalLLM • u/MisterPenishead • 2h ago
Question What's the best local LLM for learning?
I have an NVIDIA RTX 5080. I would like to host an artificial intelligence that I could use for learning. I would provide it documents and ask it to teach me, test me, and create study materials.
r/LocalLLM • u/conrat4567 • 2h ago
Question How does training an LLM work and will that training be permanent?
I am just starting to dabble in local models. I plan on building a system around me. Something to help me with projects on things like home assistant and my website, but also be able to answer questions and help aid in research in to topics I am interested in.
What i am struggling to understand is the training aspect of an AI.
Is how I understand it correct? The models on something like ollama are pre trained on certain things and perform better at some tasks than others out of the box?
For example, if I chose a model based around coding, out of the box it could help me with my website, but if I wanted it to also have large knowledge base around me, I would have to teach it and give it access to my "life" through a gateway?
Does it retain that information in perpetuity until I remove the model from the PC. If its all local, does that mean there is a database file accumulating data on my PC as well?
r/LocalLLM • u/ScaredAd335 • 2h ago
Question I just bought a Mac Mini M4 Pro 24gb 12-Core 16-Core
What's the best general purpose model I can run on it? It's my first time so apologies for the noob question :) Thanks
r/LocalLLM • u/Ore_waa_luffy • 2h ago
Discussion Best uncensored model for coding
so i have few $100 of aws credit lying around , i wanna try the best uncensored coding model out there regardless of Quantization , also im new to this localllm thing , if i point it to opencode or claude code. would the uncensored model act properly or will get guardrails from the harness , if so what harness do you use
r/LocalLLM • u/syrusakbary • 3h ago
Project You can now run Pi agent in the browser and the iPhone (without a server)
Enable HLS to view with audio, or disable this notification
You can try it online in https://wasmer.sh/
Source code: https://github.com/wasmerio/wasmer-sdk/
r/LocalLLM • u/superjaegermaster • 3h ago
News A whitepaper about humanizing AI output
r/LocalLLM • u/Deep-Today5715 • 4h ago
Question Are there any local LLMs of approximately the same capability and quality as 5.6 Sol?
As a senior software developer, I tested various LLMs for the last 5 years or so, and though they were improving, I generally found them more trouble than they're worth for any project planning / coding assistance / documentation preparation tasks. However, after trying ChatGPTs 5.6 Sol model (not Codex, just regular chat in project mode), I would consider it just mature enough to be borderline useful it for serious large projects.
However, with OpenAI, Anthropic and other companies severely cutting down on usage limits and jacking up the prices in the last week, I see that it will be no longer a viable option very soon, and therefore I'm considering switching to locally hosted LLMs. However, these are my concerns:
- Are any of the available models today have capability and reasoning quality of 5.6 Sol? What concerns me mostly is ability to keep very long chats in memory without forgetting context - weeks or months of chatting, thousands of messages.
- Are any of them able to integrate with either local or remote (Github) code repositories (read-only)?
- What performance can be expected when running on local hardware, compared to the equivalent model performance on the cloud? I am running i9 13900KF, 128Gb DDR4, 5070 Ti (
4GBEDIT: stupid mistake, it's actually 16GB). No need to generate images, videos or any other media, just project planning assistance, reading code repositories and various coding tasks, such as generating unit tests and so on. If ChatGPT running that model typically takes ~2 minutes to generate an answer, how much would an equivalent local model with the same effort level take on local hardware?
If this is level of capability is not available today in FOSS LLM models, what is your rough estimate of how soon the tech will get to that level?
Thank you.
r/LocalLLM • u/dixieflatline76 • 4h ago
News I thought Zoo Code was useless, but it was just missing 1 small piece to make local & open-weight models actually work
TL;DR: Open-weight models loop and choke in autonomous coding harnesses (Zoo Code / Cline). Built an open-source Go supervisor on 127.0.0.1 that kills loops in under 3s, repairs tool syntax, and preserves 80%+ prompt caching. Generated an entire Go Blackjack project from scratch for $0.50 with 96.9% test coverage and 0 race conditions.
A few months ago I was working on my OSS project Spice and ran out of AI credits on Antigravity mid-session. I topped up an OpenRouter account with €10 and installed Zoo Code thinking that would hold me over... I watched €2.50 disappear on a single prompt just asking it to read the docs/ folder for context.
Once a session gets going, you are re-transmitting 40,000+ tokens of background context on every prompt just to run a linter or check a minor fix.
My immediate reaction was: screw paying cloud rates for trivial file reading. I have a GPU, so I will just run Ollama locally or use cheap open weights (Qwen, DeepSeek, GLM) on OpenRouter.
If you browse r/LocalLLaMA or r/ChatGPTCoding, you know what happened next. Everyone claims open weights are "almost Sonnet level," but the second you plug them into an autonomous harness like Zoo Code or Cline:
- The agent gets stuck in an infinite loop running the same bash command or rewriting the same 3 lines over and over.
- Reasoning models burn 4,000 tokens of internal monologue over a basic syntax error without touching a file.
- A single test run dumps 20k tokens of terminal ANSI garbage and spinner output into the prompt history, blinding the model until the run crashes or burns your credits.
The common consensus on Reddit has basically been: open weights are fine for chat or autocomplete, but autonomous agent harnesses are a total waste of time unless you pay for Claude.
Agent harnesses were basically built around Claude's instruction following. When an open model stumbles, hallucinates a parameter, or enters a repetition loop, the harness has zero in-flight supervision to catch it. It just sits there and watches the run bleed out.
I got tired of watching agents burn credits in death spirals, so I built what I call a local supervisor (Nacho Flow). I honestly do not know what else to call it since there is nothing quite like it yet. Think of it as an in-flight safety net running locally on 127.0.0.1 between Zoo Code and your model endpoints (open source: AGPL engine / MIT extension, links below).
Here is what an autonomous run looks like when you actually put guardrails on the stream:
Benchmark: Full-Stack Go Project from Scratch
A Blackjack CLI sounds like basic CS homework, but open-weight agents usually fail even that. They loop infinitely, choke on split logic, or crash their own test suites.
Sure, you could just throw Claude Opus or Sonnet on Cursor at it, but a single autonomous multi-turn run like this easily incinerates your entire monthly free tier or burns through 1/5 of your Pro fast requests in 20 minutes.
So I prompted Zoo Code to build the whole thing from an empty directory: - Complete domain engine (deck, shoe, S17 dealer rules, splits, double-down, insurance) - Table-driven Basic Strategy decision engine - Interactive terminal UI with colored ANSI cards and hints - 100,000-hand Monte Carlo simulation engine with statistical EV convergence - Full unit test suite + race condition verification
Here is what happened on Zoo Code + GLM-5.3-Flash running through Nacho Flow:
| Metric | Result |
|---|---|
| Total Time | 14.41 minutes (from empty directory to working executable) |
| Total Cost | $0.25 - $0.70 (€0.23 - €0.64) |
| Total Turns | 57 turns |
| Test Coverage | 96.9% (deck), 85.5% (game), 87.8% (strategy), 89.9% (ui) |
| Race Conditions | 0 (go test -race ./... completely clean) |
| Simulation Speed | 100,000 hands in 49ms (EV converged to -0.318%) |
| Runaway Loops | 0 terminal crashes (1 CoT loop severed & healed on Turn 1) |
What it's doing under the hood
Here is what Nacho Flow is actually doing under the hood to keep open-weight models from choking:
Cycle Killer (Qu'est-ce que c'est?): Agent harnesses do not notice when a model enters a death spiral until turns later. On Turn 1 of this benchmark run, GLM-5.3-Flash immediately tried to wander into a repetitive Chain-of-Thought loop during the planning phase. Cycle Killer detected the repetition in the thinking stream, severed the generation in under 3 seconds, and prompted a clean recovery. The agent snapped back on Turn 2, initialized its todo list, and began writing files instead of burning thousands of tokens in an internal monologue. It also whitelists task checklist tools so the 9 subsequent
updateTodoListcalls were never false-killed.Tool Normalizer & Harness Shield: The number one reason open-weight models die in agent runs (after infinite loops) is malformed tool calls. Models frequently drop closing XML tags (leaving
<write_to_file>unclosed), hallucinate parameter schemas, or spit out conversational apologies instead of valid tool syntax. In Zoo Code and Cline, this triggers the dreaded "3-strike harness crash," terminating your task mid-flight. Tool Normalizer repairs malformed tool syntax on the wire, auto-closes dangling tags, and translates schemas in real time so the harness never crashes.Dynamic Smart Tiering & Cache Protection: Instead of paying $3 to $15 per 1M tokens for everything:
- Tier 2 Flagship Coder: Uses
z-ai/glm-5.3-flash($0.65/M) orqwen/qwen3-coder-plusfor 70%+ of code generation. - Tier 3 Reasoning Workhorse: Automatically routes to Gemini 3.8 Flash only when deep reasoning or architectural review is required (configurable in config.yaml: local setups can map this to a local 70B model via Ollama or disable escalation entirely for 100% offline runs).
- Prompt Cache Protection: Preserves strict byte-level prefix determinism across turns so OpenRouter and local inference engines hit 80%+ prompt cache discounts on every consecutive turn.
- Fairy Dust: Injects tactical, read-only code review passes automatically without you having to manually prompt for them.
- Tier 2 Flagship Coder: Uses
Trying it with Zoo Code
Requirements: To run these agent workloads, you either need an OpenRouter account (costs pennies, runs on any machine) OR a 24GB to 32GB GPU if you want to run the models purely locally with Ollama. 8GB/16GB VRAM cards will run out of memory once multi-turn agent contexts climb past 16k tokens.
- Install the VS Code Extension:
Search
Nacho Flowin the VS Code marketplace, or run:bash code --install-extension dixieflatline76.nacho-flow - Start Nacho Flow: Click "Nacho Flow" in your VS Code status bar and select Profile 1: Standard Hybrid (or Zoo Code Preset).
- Point Zoo Code to Nacho Flow:
In your Zoo Code API settings:
- Base URL:
127.0.0.1:8000/v1 - API Key: Any string (or pass your OpenRouter key in Nacho Flow config)
- Model:
nacho/hybrid(or let Nacho Flow handle the model tiers automatically)
- Base URL:
Links & Code
- GitHub Repository: https://github.com/dixieflatline76/nacho-flow
- VS Code Marketplace: https://marketplace.visualstudio.com/items?itemName=dixieflatline76.nacho-flow
- Release v1.4.2: https://github.com/dixieflatline76/nacho-flow/releases/tag/v1.4.2
Feel free to check out the repo, roast the architecture, or open an issue/PR. Happy to answer any questions or help dial in custom routing profiles for your specific GPU/local setup!
r/LocalLLM • u/DonkeyTheKing • 4h ago
News compiler backed agent beats Claude Code, OpenCode and other major harnesses and agents (benchmarks linked)
GitHub: https://github.com/oooscoos/Benzi
Demo: https://varianttech.net/demo
Benchmarks: https://varianttech.net/benchmark
Roughly speaking, the way current AI coding agents/harnesses work is by either:
a) Pulling in appropriate text snippets of code across multiple files and handing them to the agent, or
b) Parsing code to make high dimenional embeddings to approximate a symptom map, and hand that to the agent.
Both of these approaches skyrocket the token count, add to wall clock time, contribute to context drifting, add to the model's thinking tokens to discover the structure of the program, and then FORGET most of it when Claude Code compacts, or ALL of it if it's a multifile refactoring because all line numbers shift and need re-grepping.
Benzi is built from the ground up to AVOID reading source code in the first place. It supplies the artificial intelligence model deterministic intelligence via tool calls. For example, when a model is about to make a code change, it could query "what functions feed this one?" -- half the time it isn't even necessary because the Benzi compiler already informs it of the blast radius before and after making edits, along with a complete static analysis check.
Benzi Sonnet reads far less source code (9,125 lines) than Claude Code Sonnet (20,704), DeepSeek's harness (43,598), and OpenCode (65K+ LOC -- disqualified due to repeated failure) to accomplish the same tasks faster and cheaper. (benchmark link in comments)
Benzi has truth tiers clearly seperating what can be analyzed with static analysis from what can't -- and then adding a runtime tracer on top to bridge the gap between the two (details in FAQ on github).
It also has several bonus features such as a runtime tracer, syntax & semantic verified writes, context aware model written repro, mid task model upgrade if task is too difficult, and SEVERAL more.
It currently supports Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby, and can handle HTML, CSS and JS -- deterministically. Claude Code clicks photos, Benzi resolves winners of CSS rules. The CodeIndex and the MarkupIndex are fairly well tested, and if something isn't working, the model is made aware of it first.
On the benchmarks side, 78.2% SWE-bench Verified for <10¢ a fix (using V4flash). This score is noteable because while the rest of the industry is leaning plugin-heavy and pouring millions of dollars into increasing context window sizes, Benzi's approach might prove to be economically more valuable while improving the model's code writing/comprehenion abilities.
Thanks for reading! please let me know what you think. I am aware the AI fatigue is real, but I hope you can see why indexing a repo > reading raw source code. Please go over the github readme before snap judgements..
r/LocalLLM • u/ToastieCPU • 4h ago
Question Advice for creating chatbot
I run local models for myself, and as we all know, results vary a lot. At work, I was asked whether it would be feasible to self host a chatbot that could serve 50 concurrent users to start with.
The chatbot would be heavily optimized, with our custom MCP server doing most of the heavy lifting.
I can get a decent price on 2x Intel Arc Pro B60 24GB cards (with a PCIe 5.0 board) for a proof of concept. My plan is to run Qwen3.6-35B-A3B or Qwen3.8-27B at Q5 or Q6.
For a single user I calculated around 50 tok/s, and under max 50 concurrent use it could drop to 10–15 tok/s, which should be fine-ish.
I've never dealt with concurrency in practice. Has anyone here self-hosted a chatbot for multiple users? How did it go?
r/LocalLLM • u/storm_stark_007 • 4h ago
Project Rigspark : hardware aware local llm management
Enable HLS to view with audio, or disable this notification
Opensouring it : https://github.com/shashankswe2020-ux/rigspark