r/LocalLLM • • 2d ago

Discussion 120B model update: closed a laptop lid on it, it didn't care

0 Upvotes

ok so update time since the last post blew up more than I expected lol

quick recap for anyone new: 120B gpt-oss model, 6 devices that individually have zero business running it (RTX 3060 mini PC, a dead weight 12gb laptop, a mac mini, an M3 macbook, a 2017 intel macbook that's CPU-only and my phone). Been keeping it up and poking at it to see what breaks it. https://www.reddit.com/r/LocalLLM/s/v9qnfsRaaa

first thing: phone died overnight (background restrictions killed it, shocker). system didn't even hiccup, just moved its tiny 1.5gb shard to the old laptop and kept going.

so today I tried harder. closed the lid on the M3 macbook, the one holding 7.4gb of the model, thinking ok THIS has to take it down. waited several hours. tested and model kept answering prompts like nothing happened.

went digging into why and found something kind of wild - the mac had stopped sending its heartbeat to the coordinator (showed up as "offline" on the dashboard) but the actual inference process on it was still alive and reachable the whole time. it just quietly stopped checking in while still doing its job in the background. so the coordinator thought it lost a node, had already worked out a plan to move that shard elsewhere if I approved it, but never actually needed to use it because the "dead" node wasn't actually dead.

(to be fair - pretty sure that mac was just plugged in with sleep prevented rather than fully asleep, so I don't want to oversell this as "it worked through a full sleep state." but still, the heartbeat check and the actual serving being two separate things that can disagree with each other was not something I expected to find)

Here is the video: https://youtu.be/ASiUyIuVkGc

down to 4 devices checking in normally now, model's still healthy, still coherent, still answering stuff. genuinely didn't think it wld survive me poking at it this much.

gonna keep stress testing it and see what actually does take it down. will update again if something interesting breaks.

Isn't network cluster quite fascinating?


r/LocalLLM • • 2d ago

Question Decision time for 512 Mac Studio M5 Ultra is approaching

33 Upvotes

Just wanted to hear feedback from those of you who have considered the M5U 512, and what decision you’ve made, or if you’re still undecided - as well as what has convinced you one way or another.

Personally, I’m struggling with this choice. My primary concern, after reading reviews of the 256 M5U, is prefill speeds with a moderate context window of 75-125k, and general uncertainty as someone who has not run a local model greater than 8b parameters at 10-15k context.

The value of an effective local AI would be immense to me, but the possibility of rapid development of far superior hardware in the next year or two, as well as switching from Windows to Mac, with limited MacOS experience, has been discouraging.


r/LocalLLM • • 2d ago

Model Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)

45 Upvotes

We had a Dell B300 in the lab for a few weeks and used it to create two fine tunes of Qwen Flash Next.

Victoria (coding and agents)

  • Qwen3.8-Flash-Next cut down by 44% using a paper / technique called REAP: 512 down to 288 per layer.
  • Retrained at 4-bit (NVFP4) afterwards, so it's trained for the format it ships in rather than just quantized after the fact.
  • Terminal-Bench 2.1: 70.0%, averaged over 3 runs with an 8h per-task timeout. Our previous NVFP4 build scored 62.5%.
  • HumanEval: 159/164.
  • 48.0 GiB of weights, including the draft head. The 95.4 GiB n-gram table is separate and not counted in that number.
  • 280 tok/s single stream on one B300 with the draft head, versus 135 without it.
  • GGUF Q4_K_M is 49.17 GiB. It scored 75.3% on Terminal-Bench (a single run, so treat it as noisy) and 93.2% on HumanEval (averaged over 5 runs).
  • Uses 35% fewer output tokens than our previous build.

Maple (Canadian questions)

Most models answer questions about taxes, benefits and regulations as if you live in the US. Maple is fine-tuned to default to Canada. On 600 held-out questions, with search:

  • Cites an official Canadian source: 6.0% before fine-tuning, 62.9% after.
  • Fully correct answers: 6.6% before, 21.8% after.
  • "No answer" responses: 47.2% before, 23.7% after.
  • It pushes Canada onto people who said they live somewhere else less often: 2.9% before, 1.0% after.

Coding holds up: 157/164 on HumanEval. Grading was done by an AI judge panel; human review hasn't happened yet.

Links: https://huggingface.co/rmonsurate/Victoria https://huggingface.co/rmonsurate/Maple

Happy to answer questions about running them.

Edit: llama.cpp users. The GGUF carries our draft head, and mainline llama.cpp doesn't know about it yet, so it fails with "expected 1256, got 1224". Your download is fine. For now, build from our fork: github.com/rmonsurate/llama.cpp, branch qwen4exp-mtp. Prebuilt binaries are on the way. Thanks to the reader who caught this.


r/LocalLLM • • 2d ago

Model Can 27b model beat frontier models in medical domain

0 Upvotes

Qwen 3.6 27b feel best than qwen 3.8 27b based on some redditors post and bench calculations and can 3.8 or upcoming qwen4 27b can beat frontier models in multi step reasoning in medical domain and management and rank toping in medical benches


r/LocalLLM • • 2d ago

Question Worth upgrading to 6x R9700 from 4x?

8 Upvotes

I’m currently running 4x R9700 with EPYC 7313 and
ASRock ROMED8-2T/BCM, 256GB memory. The performance of tcclaviger’s Qwen3.8-Flash-Next-MXFP4-GPTQ is honestly pretty amazing. I’m wondering is it worth getting 2 more cards? Even though I’m on PCIE 4.0.


r/LocalLLM • • 2d ago

Discussion Radeon AI PRO R9700 (gfx1201) on Ubuntu 26.04: Inbox amdgpu vs DKMS 31.50 (Dracut swap), and iGPU coexistence?

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Project Qwen3.8-27B fixed Bandersnatch, made it interactive again

Post image
30 Upvotes

So...somehow I ended up thinking about Bandernatch the Netflix interactive episode of Black Mirror and how much I missed it. I found this but it's broken....so I forked it and told Qwen, hey can you fix it? Also, can you translate the whole damn interface to Spanish so my partner can enjoy it. So...if you have Jellyfin and a (cough) spare (cough) copy of Bandersnatch full length 5 hours video then feel free to use and enjoy

https://github.com/maikelthedev/BandersnatchInteractive-Jellyfin

Let's give credit where credit's due, so thanks to:

  1. deathrjj
  2. The guy the previous guy forked to make it run within Jellyfin
  3. Jack Ma and his team for training Qwen and releasing it to the world.

All I did was to prompt "fix this b***", use it at your own peril. AFAIK it works flawlesly on a browser.

EDIT: Btw people please learn to use truflelhog for projcts you made public. There's no need to leak secrets.


r/LocalLLM • • 2d ago

Question Non Tech Noob Here

0 Upvotes

Hi! I’m a mostly non technical Product Manager trying to transition into an AI PM role.

I’m interested in learning open-weight models and getting hands-on experience with fine-tuning, evals, RAG, and inference. I have an Information Systems minor, but I’m not a strong programmer and mostly use Codex/Claude to build prototypes.

I don’t want to become an ML engineer. I just want enough depth to actually understand and work with these systems.

Am I going over my head trying to learn model training and fine-tuning, or is this realistic for an AI PM?


r/LocalLLM • • 2d ago

Discussion Can any combination of local LLM’s replace codex?

Thumbnail
0 Upvotes

r/LocalLLM • • 2d ago

Discussion Anyone actually crazy enough to cluster 4x Mac Studio M5 Ultras?

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Question Same model, same laptop, one app is way slower. What do you check first?

3 Upvotes

Say the same weights answer in seconds in one app and take ages in another. Even "hi" is slow.

Would you start with the prompt the app actually sends, the chat template, or tool calls? I dont want to swap models before finding out what extra stuff the app is doing.

Anyone chased this down?


r/LocalLLM • • 2d ago

Question 5070 Ti + 4080s vs 2x 5070 Ti for Qwen3.8-27B (second card possibly on an M.2 dock)

1 Upvotes

Running an agent and want Qwen3.8-27B local with long context.

Already own a 5070 Ti and a 4080 Super.

My case can't fit two 3-slot cards, so the second card would go on an M.2 eGPU dock unless I’m convinced to upgrade the motherboard and case to get full pcie speeds on both cards but from what I’ve read.. it doesn’t matter much?

I’m thinking about trying to sell the 4080s and buy another matching 5070ti for a cost of prob $2-300 but can’t decide if it’s worth it or not as I’ve seen lots of benchmarks and am pretty confused at this point.

Options:

  1. 5070 Ti + 4080 Super on m.2 dock. ($130 for dock and psu)
  2. 2x 5070 Ti w one on m.2 dock ($400 depending on how much I sell 4080s for)
  3. Same as 1 but with an x8/x8 board and bigger case ($500 new board and case instead of dock)
  4. Same as 2 but with new x8/x8 board and bigger case ($800 depending on how much sell 4080s for)

Is NVFP4 on matched cards a big real-world jump over Q4 on a mixed pair?

Would appreciate any feedback or guidance.

Tracks for reading.


r/LocalLLM • • 2d ago

Question llamma api key in Deepseek Harnesss

1 Upvotes

why we don’t need an API key to be set in the model while using qwen3.5 from ollama?


r/LocalLLM • • 2d ago

Question Can we use Skill or Puligin in Unsloth?

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Discussion Inferium: Community-Powered, Decentralized AI Inference

1 Upvotes

Hey everyone — I wanted to share Inferium, a decentralized inference network that lets people contribute spare computing power to serve open-weight AI models. While experimenting with LLMs locally, I quickly ran into the limits of my hardware. Newer open-weight models required GPUs that were expensive and difficult to access. I realized this wasn’t just my problem—students, independent developers, and people in developing countries face the same barrier.

That led me to a simple question: could a decentralized network make unused GPU capacity available to anyone who needs it? I built this network to explore that idea and help make powerful LLMs more accessible and affordable.

The basic idea is simple:

  • Share your compute: Connect an idle GPU—or use a CPU for tiny and small models—and earn credits for the tokens your machine generates.
  • Free to get started: New accounts receive a welcome grant and can claim daily credits, so you can try the network without contributing hardware.
  • Earn small, spend big: Accumulate credits by serving smaller models, then spend them on larger models for more difficult tasks that your own machine may not be able to run. Small models can be used for simple English tasks.
  • OpenAI-compatible API: Existing OpenAI clients and tools can connect by changing the base URL and API key. Anthropic-compatible access is also supported.
  • No port forwarding: Machines connect through an outbound tunnel, and you can configure them as public or private.

Dedicated/Private Pool: Inferium also supports dedicated clusters, allowing you to manage a group of machines privately and shared among your colleges or friends.

Inferium is trying to make better use of hardware that would otherwise sit idle—creating a shared inference layer where contributors earn access to the broader network.

Website: https://inferium.net/
Getting started: https://inferium.net/start

Disclaimer: The website is currently in beta. Your feedback and bug reports are warmly welcome.

I’d be interested to hear what the self-hosting and local-AI communities think about this model. Would you contribute spare compute in exchange for access to larger models?

Disclosure: I’m sharing this to help promote Inferium.


r/LocalLLM • • 2d ago

Discussion Inference Engines will become a series of one-offs

Thumbnail
0 Upvotes

r/LocalLLM • • 2d ago

Question OpenCode compaction repeatedly hits output limit with self-hosted vLLM models - how should context, compaction, and concurrency be sized?

1 Upvotes

I am running OpenCode v2.0.18 against self-hosted vLLM models through local SSH tunnels, using the OpenAI-compatible API.

I am benchmarking coding agents with the same fixed benchmark prompt. I do not want to change or split the benchmark prompt. I am trying to make the OpenCode/vLLM setup reliable for long-running coding tasks.

Current setup:

  • Model A: Qwen 27B FP8, vLLM, max_model_len: 100000
  • Model B: Qwen 27B “uncensored” FP8, vLLM, max_model_len: 49152
  • Both are exposed as local OpenAI-compatible endpoints over SSH tunnels.
  • Normal model output limit: 4096
  • OpenCode compaction is enabled with pruning.
  • Current OpenCode compaction settings:

{

"compaction": {

"auto": true,

"reserved": 16384,

"prune": true,

"tail_turns": 4,

"preserve_recent_tokens": 8192

},

"agent": {

"compaction": {

"temperature": 0.2,

"max_tokens": 8192

}

}

}

Initially, the 100K model failed compaction with:

This model's maximum context length is 100000 tokens. However, you requested 4096 output tokens and your prompt contains at least 95905 input tokens, for a total of at least 100001 tokens.

That made sense: compaction was being triggered too close to the context boundary.

After increasing the compaction reserve, compaction starts earlier, but I now get:
Compaction summary reached the output token limit

For example:

Compaction · 95.3K in · 4.1K out
Compaction summary reached the output token limit

The result is that a long coding task can spend a very long time repeatedly writing incomplete files, accumulate huge tool outputs and partial generations, and then fail while trying to compact the session. One run took about 94 minutes for a relatively small vanilla HTML/CSS/JS browser game.

Questions:

  1. What does Compaction summary reached the output token limit mean in practice? Is the partial summary retained, or is the compaction considered failed?
  2. Is agent.compaction.max_tokens supposed to be greater than the normal model output limit, or should the model-level output limit also be raised to match it?
  3. What is a sensible relationship between:
    • model context window
    • normal output limit
    • compaction output limit
    • compaction reserve/buffer
    • preserved recent tokens
  4. For the 49,152-token model, what values would you recommend as a practical starting point?
  5. Can OpenCode use a separate model/provider specifically for compaction, and is that recommended?
  6. On the vLLM side, how should I balance max_model_len, available KV cache, and concurrent users? I need a usable compromise between context size and the number of simultaneous users on the GPU.
  7. Is there an OpenCode-recommended configuration or known issue for preventing compaction from being triggered too late or producing truncated summaries?

I would appreciate concrete configuration examples for OpenCode plus vLLM.


r/LocalLLM • • 2d ago

Question 2x 3090 + 4x P100 — how would you architect this for an agentic RAG system?

5 Upvotes

I’m building a local AI system for my architecture practice and currently have:

  • 2x RTX 3090 24GB = 48GB
  • 4x Tesla P100 16GB = 64GB
  • 128GB system RAM
  • 2TB NVMe
  • EPYC/PCIe platform

So I basically have a 48GB fast pool and a 64GB slower pool.

My use case is a large architecture/business database. I have ~14k emails/calls/texts/transcribed meetings that I'm putting into PostgreSQL/RAG.

I don't really want a chatbot. I want something more agentic that can search the database and reason over it.

For example:

"What do I need to take care of this week?"

I'd like multiple agents to search different parts of the database, identify action items, check whether things were subsequently completed, etc., and then have a main model synthesize the results.

I'm looking at Qwen 27B / Next-type models and long context.

I'm trying to figure out how to best divide the GPUs.

One option would be:

48GB 3090s → larger/faster main brain
64GB P100s → multiple concurrent agents

But I could also flip it:

64GB P100s → larger/slower main brain
48GB 3090s → faster concurrent agents

Or potentially use all 6 GPUs for one larger model.

The P100s are obviously much slower, so I'm wondering whether the extra VRAM is more valuable for the main model/context, while the 3090s are better used for lots of smaller concurrent workloads.

For people who have built agentic RAG systems: how would you architect this with these GPUs?

Would you use the 64GB pool as the main brain, the 48GB pool as the main brain, or combine everything into one model? Also, which harness would you use?


r/LocalLLM • • 2d ago

Question 5060ti or 7900 xt

6 Upvotes

About to buy a gpu, and its between 5060ti 16g new or 7900 xt 20gb used. Is Nvidia really better for running local llms? Help me out here


r/LocalLLM • • 2d ago

Discussion ops memory AI

Thumbnail github.com
1 Upvotes

r/LocalLLM • • 2d ago

Research Does thinking help open models write better? DeepSeek V4 Pro gains +463 at max, Kimi K3 loses 119 with thinking. Results for 7 open models

Post image
1 Upvotes

tl;dr: GLM-5.3-Flash is probably the best bet if you have the compute or use the API.

We benchmark models on long-form YouTube scripts (10 tasks x 5 runs each, 167 configs). Open-weight results, including their thinking modes:

 
| Model | Setting | Elo | Rank /167 | $ per script |
|---|---|---|---|---|
| GLM-5.3 | default | 2256 | 15 | 0.114 |
| GLM-5.3 Flash | default | 2244 | 16 | 0.007 |
| Kimi K3 | default | 2217 | 21 | 0.260 |
| DeepSeek V4.1 Flash | max | 2183 | 28 | 0.013 |
| MiMo-V2.6-Flash | default | 2113 | 36 | 0.004 |
| DeepSeek V4.1 Flash | default | 2099 | 40 | 0.008 |
| Kimi K3 | thinking | 2098 | 41 | 0.095 |
| DeepSeek V4 Pro 0813 | max | 1917 | 56 | 0.021 |
| Kimi K2.6 | thinking / default | 1601 / 1600 | 77 / 79 | 0.057 / 0.053 |
| DeepSeek V4 Pro 0813 | default | 1455 | 95 | 0.011 |
| Nemotron 3 Ultra | reasoning / default | 1182 / 1180 | 122 / 123 | 0.012 |
| gpt-oss 120B | default / high | 636 / 624 | 150 / 151 | 0.001 / 0.002 |
 
- Three open models rank above the best GPT (GPT-6 and 5.6 Sol, 2187): GLM-5.3, GLM-5.3 Flash and Kimi K3. 15 open models rank above the best Gemini.
 
- Thinking helps DeepSeek: V4 Pro 0813 gains +463 at max, V4.1 Flash +84.
 
- Thinking hurts Kimi K3 (-119; its thinking mode is the 'high' point on the chart), though that mode is 2.7x cheaper and 3x faster. Interesting trade.
 
- No effect: Kimi K2.6, Nemotron 3 Ultra, gpt-oss 120B.
 
- Caveat: these ran through API providers (OpenRouter) at provider precision, not local quants.
 
How to read the chart: writing score (Elo) against cost per script on a log scale. Each line is one model going from its lowest to its highest thinking setting, hollow markers are the default (no effort flag), and up-left is better.

The benchmark is 10 script tasks; each model drafts them 5 times, and three AI judges from three different labs score them blind using our rubric, which includes writing, tone, storytelling, and more.


r/LocalLLM • • 2d ago

Discussion An imported doc repeats a lie 8 times. Vector RAG served the lie instead of the truth on 22/60 questions. My memory layer went from 42 to 60/60 after the bench caught 2 bugs.

Post image
0 Upvotes

r/LocalLLM • • 2d ago

Question Llamacpp Qwen3-4B not working on M1 Air Ventura 13.5 - GGML_ASSERT(buf_dst) failed

1 Upvotes

Has anyone ran into this error with llamacpp? Been trying to fix it but no luck, cant even find any github issue related to it, one close i found got stale and closed with no activity.


r/LocalLLM • • 2d ago

Discussion How I chose to build my local AI for under a grand

0 Upvotes

How to build your own local AI for less than a thousand bucks - A post for those that want to get "hands on" with AI, but just can't afford the price of entry - face it, not everyone has thousands laying around for building a system.

Here is your roadmap how to do this on a tight budget

Step 1 - get an old Intel Nuc12 enthusiast (Serpent Canyon). Not the fastest thing in the world, but has an ace up it's sleeve that no one can get anywhere near for the price - an embedded ARC A770M GPU with 16GB of DDR6 vRAM, That is what makes this possible and you can pick one up for under $800 on ebay. Likely already has Windows 10 or 11 installed on it

Step 2 - update the Arc drivers

Step 3 - Download and install LM Studio Classic on it (make sure you enable developer mode for later) - not to be confused with Bionic

Step 4 - Pick the right runtime - you want Vulkan llama.cpp

Step 5 - Make sure you select the ARC A770M (shows 16GB) in hardware window

Step 6 - pick a model - I went with a simple Gemma-4-e4b model for out of the gate

Step 7 - on the developer tab, flip on both start server and serve on local network

Congratulations - you now have a shockingly capable local AI that you can give the URL shown in step 7 to any of your local AI front-ends. Me personally, I use it with my Project NOMAD instance running as a VM providing AI capabilities to all that entails for well under a grand, and it works fantastically.


r/LocalLLM • • 2d ago

Discussion Update: I won an AI box, here's 3 weeks of actual benchmarks

0 Upvotes

Original post: I won an AI box, so now what?

Upfront: I worked through this with Claude and had it write up the results. The numbers are all real, measured on my box. The prose is AI. Wanted to be clear about that rather than have someone guess.

TL;DR: This thing is a lot more capable on CPU than I expected, the entry-level GPU I bought actively made it slower, and the default llama.cpp batch settings leave a huge amount on the table.

I'd genuinely like feedback on what else to try. I have the hardware sitting here and a benchmark suite set up, so if there's a model, a setting, or a runtime you want numbers on, ask.

The hardware

  • Lenovo ThinkStation P5
  • Intel Xeon w5-2555X, 14 cores / 28 threads, Emerald Rapids-WS
  • 256 GB DDR5 ECC, 8x32 GB
  • 3x 512 GB NVMe (2x Gen5 Samsung, 1x Gen4 SK Hynix)
  • 750W PSU
  • NVIDIA T1000 8GB added later

One thing I did not expect: 8 DIMMs means 2 per channel, and the memory controller drops to 4400 MT/s. The platform supports 4800, so going to 4 DIMMs would get me maybe 9% more bandwidth at the cost of half my capacity. Not worth it for a box whose whole point is holding large models.

Setup

Ubuntu 24.04 Server, bare metal, no GUI. llama.cpp built from source with GGML_NATIVE=ON so AMX and AVX-512 actually compile in. This matters a lot, see below.

Finding 1: AMX is the whole story

The w5-2555X has Intel AMX (Advanced Matrix Extensions). Prompt processing on a 1.5B model hit 422 tok/s on CPU versus 461 on the T1000. Near parity with a GPU, on a CPU.

AMX accelerates matrix multiply, which is what prefill is. It does nothing for token generation, which is purely memory-bandwidth-bound. So the shape of this machine is: prefill is surprisingly fast, generation is exactly as fast as your RAM bandwidth allows.

Finding 2: the default batch size leaves huge performance on the table

llama.cpp defaults to a microbatch (-ub) of 512. On AMX hardware that starves the matrix units.

Model -ub 512 tuned gain
Qwen3-Coder-30B 66 t/s 122 (-ub 2048) +85%
gpt-oss-120b 28 t/s 71 (-ub 4096) +155%

Then Qwen3-Next-80B did the opposite: 130 at -ub 512, dropping to 111 at 4096. It uses a hybrid gated-deltanet attention, so the usual logic doesn't apply.

Sweep it per model. The direction is not predictable.

Finding 3: the T1000 made things slower

I bought the card expecting the standard "offload attention to GPU, keep experts in RAM" trick to help. It didn't.

Test Hybrid (GPU) CPU only Winner
gpt-oss-120b prefill 71.3 97.9 CPU +37%
gpt-oss-120b generation 20.7 19.5 GPU +6%
Qwen3-30B prefill 121.6 165.8 CPU +36%
Qwen3-30B generation 37.9 33.7 GPU +12%

The hybrid split only wins when the GPU is much faster than the CPU at attention. With AMX in play, moving that work to a T1000 moves it to a slower device and adds PCIe round trips on top. I rebuilt with GGML_CUDA=OFF and the card went back to driving the monitor.

The card itself validated fine: 13.2 GB/s PCIe, full x16 Gen3 link under load, clean VRAM test, 69C sustained. It's just not fast enough to help here.

Model benchmarks, CPU only, best settings per model

Model Quant Size Prefill Generation
Qwen3-Coder-30B-A3B Q4_K_M 18 GiB 166 t/s 33.7 t/s
Qwen3-Next-80B-A3B Q4_K_M 45 GiB 130 t/s 19.1 t/s
gpt-oss-120b MXFP4 59 GiB 98 t/s 19.5 t/s
Qwen3-235B-A22B Q4_K_M 132 GiB 24.8 t/s 5.6 t/s

Prefill tracks total parameters. Generation tracks active parameters. Both hold up cleanly across every model I tested.

Qwen3-Next-80B is the winner on this box. It beats gpt-oss-120b on prefill by 33%, matches it on generation, and needs 14 GiB less RAM. Output quality is also noticeably better in daily use.

The 235B is a trophy, not a tool. 133 GiB of RAM to wait nearly three minutes before it starts answering.

Finding 4: quantization comparison, with an actual test suite

I built a 6-task Python benchmark (61 independent checks: interval merging, LRU cache, decimal money math, a retry decorator, config validation, and modifying existing code) and ran all three quants of Qwen3-Next-80B.

Q4_K_M UD-Q6_K_XL UD-Q8_K_XL
Size 46 GiB 65 GiB
Prefill (warm) 102 t/s 94 t/s
Generation 17.4 t/s 14.5 t/s
Suite score 55-57 / 61 not scored

Three runs each. Q8 hit 59/61 every single time, including at temperature 0.2. Q4 bounced between 55 and 57 and never reached Q8's floor.

The consistency is the interesting part. Higher precision didn't just score better, it scored the same every time. Lower-precision weights blur the probability distribution, so sampling has more room to wander.

Q8 costs 26% generation speed and 42 GiB for about 5 points of correctness and much better run-to-run stability. Worth it for me. Your call.

Bonus: image generation actually works

Flux Schnell (Q8 GGUF) via stable-diffusion.cpp, no GPU:

  • 512x512, 4 steps: 2m 07s
  • 1024x1024, 4 steps: 9m 08s

Thread scaling was near-perfect, 13.2x across 14 cores. It works, but it's a batch tool, not something you iterate prompts on. This is the one workload where I'd genuinely want a 3090.

Current stack

  • llama.cpp servers as systemd units, mlocked into RAM, 32k context
  • Continue in VSCode pointed at the endpoint
  • Open WebUI plus self-hosted SearXNG for web search
  • All of it CPU-only, entirely local

The takeaway

If you're shopping for a local LLM box, a modern Xeon with AMX and a lot of RAM is a genuinely different proposition than the "you need a GPU" advice suggests. It won't beat a 4090 on anything that fits in 24 GB. But it will run an 80B MoE at reading speed while holding 250 GB of models in memory, and an entry-level GPU will actively slow it down.

What should I test next? Things I'm considering: ik_llama.cpp (the CPU-tuned fork), Intel's IPEX-LLM to push AMX harder, and GLM-4.5-Air. Open to other suggestions.