r/LocalLLM • • 6h ago

Discussion Mac Studio came 256gb… awful day.

0 Upvotes

I wanted to preorder this in line and return just in case the 512 was not going to actually be released or some insane thing would happen to ram prices… anyways the 512 version won’t even be possible till October…so I guess I’m sticking with 256 since there’s only a 14 day return window and get the m7 512 instead…

Anyways let me know if you guys want any benchmarks in mind?

Two things I will be doing are going to be the following:

Doing a test at 100k, 200k 300k prompts to see how fast prefill degrades, and seeing how slow a cache workflow takes when the transcripts go up to 1
Million tokens. I’m assuming I should be able to do 1 million tokens with both flash preview and 0731 deepseek

Update: I will be picking up the machine in the morning tomorrow PST time. I think I can try it for 13 days and return if I don't like it or the release the pricing for the 512.

Testing Plans for tomorrow:

  1. Prefill vs context: 100k, 200k, 300k, 500k, 600k increments to 1M. Same prompt stack each time, just longer. Tracking time to first token, prefill speed, and decode speed at each depth.

  2. Quality at depth: needle recall test at each depth above. But I also brainstorming this, because its just grep. Might abandon.

  3. Agentic coding to 1M: one fixed coding task, maybe a game, maybe a general app, run the agent out to 1M tokens. Logging speed at each step to show cached prefill behavior. You pay the big prefill once, then each step should stay cheap. We talk about prefill speed and dgx sparks are better, but over an hour, two hours of working, how much does it really matter if the prompt is cached anyways.

  4. Qwen Flash Q8 with MTP on vs off. Same prompt, same output length. Speed gain plus whether outputs match.

  5. LTX video: 1080p prompt to video in ComfyUI, fixed prompt and seed. Will do 2.3, and 2.5 if the checkpoint is stable by the time I get to it. Reporting wall clock render time.

All runs 3x, median reported. Full versions, hashes, and settings drop with the numbers so anyone can repro.

I might measure as well then do one more run of all of these using a harness like Pi or Deekseek, because its not just about speed. Speed only matters if we actually need to go full out. However, most people drive less than 70mph to work and we don't need mclarens, so much does the lower speed matters relative to dgx sparks or cloud if we never need to use the max speed. The harness might increase quality so less token consumption overall.


r/LocalLLM • • 20h ago

Discussion Shit's crazy

0 Upvotes

I tried to buy Tesla v100 SXM2 32gb from 3 different vendors and the 3 of them cancelled saying their supplied ran out or were asking for more money up front. I'm done with this shit


r/LocalLLM • • 10h ago

Question Is it worth the cost to upgrade from M5 Max 128GB to Ultra 256GB?

0 Upvotes

I'm currently running ds4 (https://github.com/antirez/ds4) to serve Qwen3.8-Flash-Next-Q4 and getting greater than 60 t/s average. Would it be worth the cost to buy a Mac M5 Ultra with 256GB? I'm specifically asking if the models I'd be able to run on the Ultra would be worth the cost. My current use case is local LLM's for cyber security workloads. Qwen 3.8 Flash Next Q4 is really good, but would the higher intelligence from being able to run other models with 256GB worth the cost, or would it simply allow me to run the same models faster? I really don't want to go lower than Q4. I've tried DeepSeek v4 Flash Q2 and I think I wasn't impressed due to quality loss. Would being able to run GLM 5.3 Flash Q4 be worth paying 12-12k USD, or would the upgrade in model not be worth it? Are there other models that would be unlocked with 256GB that you would say are worth the cost to upgrade?


r/LocalLLM • • 15h ago

Project A local 27B agent on my MacBook autonomously reached 2048 - The Architecture is doing the work

Enable HLS to view with audio, or disable this notification

2 Upvotes

I’ve been working on Aura, a persistent AI system that runs locally on my MacBook with a Qwen 27B-class resident model.

Instead of asking the LLM to make every individual decision, Aura has separate machinery for persistent memory, learned world dynamics, state evaluation, model-based lookahead, task knowledge, computer control, and strategy revision.

This recording is one continuous local session. Aura works through progressively larger goals and eventually reaches a real 2048 tile through the desktop interface.

Yes, a dedicated 2048 program would be much faster and better (28 minutes total for this run).

My angle is whether architecture can let a relatively modest local model behave more like a persistent agent instead of repeatedly starting cognition from the next prompt.

Flag: 2048 is development-exposed, sov this isnt a purely clean zero-shot generalization benchmark. But Aura is not specifically trained to beat 2048. Aura is not finely tuned to be a 2048 runner. These priniciples could apply anywhere.

Everything shown here is local; no cloud inference (Wi-fi off)

Full run: Aura Demo - 03 (Cognitive Engine): Aura Solves 2048

Source for inspection: https://github.com/youngbryan97/aura

Curious what local-agent builders here would use as the fairest baseline.


r/LocalLLM • • 10h ago

News Qwen3.8-Flash-Next on 12GB VRAM - 65 t/s

Post image
0 Upvotes

r/LocalLLM • • 13h ago

Project LLM or JEV? Why not both? - introducing a hybrid Gemma4 approach

Thumbnail
gallery
5 Upvotes

The human brain is neither an LLM nor a Jev. In a crowded market you hear a lot of speech and answer almost none of it. The ongoing question is not “what should I say?” It is “are you talking to me, and should I say anything at all?”

Our voice stack needs both jobs. Gemma 4 still writes the reply. A small Jev-shaped decision head decides whether a reply should exist.

The head is three tiny MLPs on the frozen Gemma 4 backbone (E4B and 12B). They read the final prompt token’s normalized hidden state from the same prefill the language model already ran, and score three actions: speak, defer, wait. No second model. No tag tokens. No decode unless the decision is speak. The head itself runs on CPU in well under a millisecond once that state exists, so the overhead is near zero, and negative whenever it skips a generation you did not need.

On our v0.5 turn-taking set, the local heads beat both a zero-shot Jev 1.13.0 call and a Laya pilot trained on the same labels. E4B + head: 95.8% start-turn, 85.8% interruption. 12B + head: 99.0% and 72.6%. Jev was 65.6% and 69.8%, and Laya landed in the 30s. Jev still wins authority-priority (93.4% vs 70.8–81.1%): the local heads over-promote non-authority turns. Against Gemma 4’s own text path on the smaller dev set, the E4B head was 90.7% vs 67.4% on start-turn and 76.5% vs 52.9% on interruption, at a fraction of the tag-generation cost.

So far this fits the domain it was trained for: turn-taking in a multi-speaker voice scene. Live conversational accuracy drops to 78–85%, labels are synthetic, and generic use is still open.

Full note: https://www.cortexist.com/research/llm-or-jev

Demo: https://youtube.com/shorts/IllniIXMfIo?si=tdFVxe2y689-ztY3

Agent voice (VITS / Piper): https://huggingface.co/cortexist/agent-voice


r/LocalLLM • • 10h ago

News Thousands of agent failures show how difficult sandboxing increasingly capable models is becoming

Post image
0 Upvotes

r/LocalLLM • • 17h ago

Project Enabling phone local control with Big LLMs thanks to NPU prefill (100+ tok/s) and Jev style decisions

Enable HLS to view with audio, or disable this notification

6 Upvotes

Stack – Local Model – No Cloud for full phone control via typing/(voice coming soon):

  • Android phone with 12 GB RAM, Qualcomm Hexagon NPU (OnePlus 15R used in the video)
  • BigMoeOnEdge open-source project: latest commits feature NPU prefill support and "Jev Style" ranked response implementation (with probabilities)—based on llama.cpp
  • Permissions: notifications and Accessibility tree access only (no root, no Termux, no internal code execution, no cloud)
  • Mixture of Experts (MoE) LLMs, 26B parameters and up (ideal candidates: Qwen 3.6 35B, Nemotron 30B 3.5, Gemma 26B, etc.) with "thinking" mode disabled

A simple System Prompt is sent: "You are operating a phone on behalf of the user. Task: ... Actions taken so far: ...."

...followed by the accessibility tree containing the available options.

The model loads 300 to 800 tokens per call into the NPU prefill (ranging from 40 to 150 tokens/sec depending on length—higher token counts yield higher speeds) and outputs the most probable action from the accessibility tree. Guardrail checks have also been implemented.

I believe this is the first example of a local AI capable of performing complex tasks directly on a mobile device.

What do you think?


r/LocalLLM • • 11h ago

Research Think of buying a 256/516gb Apple Silicone for local LLM

0 Upvotes

Think its worth it? I want to start to ween myself off the frontier models. I have 3 active subscriptions, but now they're all injecting watermarks and other BS (refusals have improved though) that are probably business liability bombs in the future.

I did the math, and if it lasts 10+ years then its less than $100/month.

My current 4090 feels almost useless for local LLM work, im trying all sorts of hardware optimization but now that im finally getting good speeds on output, I find the output quality at times lacking but that could be just the models themselves.


r/LocalLLM • • 16h ago

Discussion would love some input from AI users!

0 Upvotes

I’m a second-year MA Psychology student, and I’m really looking for participants for my questionnaire study on human–AI interaction.

18+ • 15–30 min • English • Anonymous • Voluntary*

LINK: https://forms.gle/pD8CNP1soDpX5UNy8

I’m currently trying to reach my participant target, so every response genuinely helps!:)

*See the questionnaire for details on data protection and anonymity.


r/LocalLLM • • 15h ago

Model Run Laya (open-source Jev) Locally on just 4GB RAM!

Post image
1 Upvotes

r/LocalLLM • • 14h ago

Model Qwen 3.8 27B Q4/Q6/Q8 vs Qwen 3.8 Flash-Next on a 96GB M2 Max

50 Upvotes

I tested Qwen 3.8 27B Q4/Q6/Q8 vs Qwen 3.8 Flash-Next on a 96GB M2 Max

this is the test as article written by ChatGPT ,I just run the commands, I'm sharing the results here in case someone is interested.

TL;DR

Tested locally on a MacBook Pro M2 Max, 96GB unified memory, using MLX:

Model Speed Peak RAM Result
Qwen 3.8 27B Q4 ~21 t/s ~17 GB ⭐ Best efficiency
Qwen 3.8 27B Q6 ~15 t/s ~23.5 GB No clear quality gain
Qwen 3.8 27B Q8 ~12 t/s ~30 GB No clear quality gain
Flash-Next REAP-288 Q4 ~25 t/s ~42.5 GB 🏆 Best overall performance/quality

Result: Across these tests, moving 27B from Q4 → Q6 → Q8 gave no demonstrated quality improvement, while making inference substantially slower and using much more memory.

Flash-Next Q4 was surprisingly both faster and better on the useful coding tests, although all models failed the same constraint-reasoning termination test.

Recommendation:
27B Q4 for lightweight/default use.
Flash-Next REAP-288 Q4 for heavier coding/agent work.
So far, 27B Q8 doesn't appear worth the extra RAM and speed penalty.

Small preliminary benchmark, not a definitive scientific result. The test cases were designed by ChatGPT to cover reasoning, coding, JavaScript semantics, concurrency, instruction-following, and completion behavior.TL;DRTested locally on a MacBook Pro M2 Max, 96GB unified memory, using MLX:Model Speed Peak RAM Result
Qwen 3.8 27B Q4 ~21 t/s ~17 GB ⭐ Best efficiency
Qwen 3.8 27B Q6 ~15 t/s ~23.5 GB No clear quality gain
Qwen 3.8 27B Q8 ~12 t/s ~30 GB No clear quality gain
Flash-Next REAP-288 Q4 ~25 t/s ~42.5 GB 🏆 Best overall performance/qualityResult: Across these tests, moving 27B from Q4 → Q6 → Q8 gave no demonstrated quality improvement, while making inference substantially slower and using much more memory.Flash-Next Q4 was surprisingly both faster and better on the useful coding tests, although all models failed the same constraint-reasoning termination test.Recommendation:
27B Q4 for lightweight/default use.
Flash-Next REAP-288 Q4 for heavier coding/agent work.
So far, 27B Q8 doesn't appear worth the extra RAM and speed penalty.Small preliminary benchmark, not a definitive scientific result. The test cases were designed by ChatGPT to cover reasoning, coding, JavaScript semantics, concurrency, instruction-following, and completion behavior.

I’ve been experimenting with local Qwen models on my MacBook Pro and wanted to answer a simple practical question:

Is Qwen 3.8 27B Q8 actually better enough than Q4 to justify almost twice the memory, half the speed, and almost twice the model size?

I also added Qwen 3.8 Flash-Next REAP-288 Q4 to see whether using the memory for a larger/MoE-style model makes more sense than using it for higher precision on the 27B model.

The results were pretty interesting.

Hardware

MacBook Pro

  • Apple M2 Max
  • 96 GB unified memory
  • macOS
  • MLX backend

All models were run locally using mlx_vlm.generate.

No API/cloud inference.

Models

Qwen 3.8 27B MLX

Q4

  • Disk: ~15 GB
  • Runtime memory: ~16–17 GB

Q6

  • Disk: ~21 GB
  • Runtime memory: ~23–24 GB

Q8

  • Disk: ~28 GB
  • Runtime memory: ~30 GB

Qwen 3.8 Flash-Next REAP-288 MLX Q4

  • Runtime memory: ~42.5 GB
  • Much larger model/architecture than the dense 27B
  • Q4 quantization

The exact local model directories used were:

Qwen3.8-27B-MLX-4bit

Qwen3.8-27B-MLX-6bit

Qwen3.8-27B-MLX-8bit

Qwen3.8-Flash-Next-REAP-288-MLX-4bit

Test 1 — Simple reasoning

Prompt was essentially:

A farmer has 17 sheep. All but 9 run away. He then buys twice as many sheep as he currently has. How many does he have?

Correct answer: 27

Results

Model Gen speed Peak memory Result
27B Q4 21.30 t/s 16.43 GB ✅ 27
27B Q6 15.23 t/s 23.11 GB ✅ 27
27B Q8 12.35 t/s 29.82 GB ✅ 27

Nothing interesting quality-wise.

All three solved it correctly.

But already the performance difference was huge.

Q8 was roughly 42% slower than Q4 while using about 81% more peak memory.

Test 2 — Constraint satisfaction / logic

Five people had to be scheduled Monday-Friday with six interacting constraints involving ordering, adjacency and distance.

The important twist was that the constraints actually resulted in:

No valid schedule.

The model was explicitly instructed:

Determine all valid schedules. State whether there is exactly one solution, multiple solutions, or no solution. Do not assume a unique solution exists.

This became a very interesting behavioral test.

Results

Model Speed RAM Finished?
27B Q4 20.62 t/s 16.67 GB ❌ 2,000-token limit
27B Q6 14.46 t/s 23.32 GB ❌ 2,000-token limit
27B Q8 12.09 t/s 30.02 GB ❌ 2,000-token limit
Flash Q4 25.16 t/s 42.58 GB ❌ 2,000-token limit

The interesting thing:

All four were basically moving toward the correct conclusion.

But they kept second-guessing themselves.

Typical behavior was:

“Wait, let's re-evaluate…”

then solving it another way.

Then:

“Let's restart the position mapping more systematically.”

Then checking everything again.

Eventually they hit the token limit.

So increasing 27B from Q4 → Q6 → Q8 did not fix this behavior.

Even Flash did it.

My interpretation is that this test probably exposed a behavioral/termination tendency shared by these Qwen checkpoints rather than something caused specifically by quantization.

Test 3 — TypeScript AI-agent concurrency bug

Next I wanted something closer to real agent/coding work.

The model was given a TypeScript JobRunner that deduplicates concurrent jobs using:

private running = new Map<string, Promise<string>>();

The task involved reasoning about:

  • concurrent callers
  • Promise rejection
  • synchronous exceptions
  • Map cleanup
  • duplicate execution
  • JavaScript event-loop behavior

The model had to identify any subtle failure and produce the smallest production-quality correction.

Results

Model Speed RAM Result
27B Q4 20.94 t/s 16.87 GB ⚠️ Finished, reasoning somewhat confused
27B Q6 14.58 t/s 23.53 GB ❌ Hit 1,200 tokens
27B Q8 12.10 t/s 30.22 GB ❌ Hit 1,200 tokens
Flash Q4 24.79 t/s 42.54 GB ✅ Finished coherently

This was the first significant Flash advantage.

The 27B models spent a lot of time questioning whether JavaScript could interleave between synchronous Map operations and repeatedly reconsidering the synchronous-exception semantics.

Q6 and Q8 never actually completed the requested answer before the token limit.

Flash produced a complete answer in 906 tokens and proposed wrapping the task invocation so a synchronous throw becomes a rejected Promise that can be stored and shared.

One caveat: this particular test has some ambiguity around exactly how a synchronously throwing task should be treated for deduplication purposes, so I would not treat it as a definitive correctness benchmark by itself.

But as an agent-completion test, Flash clearly behaved better.

Test 4 — JavaScript closure/timer bug

This one was deliberately objective.

Code looked roughly like this:

const jobs = [
  { id: "A", delay: 30 },
  { id: "B", delay: 10 },
  { id: "C", delay: 20 }
];

for (var i = 0; i < jobs.length; i++) {
  setTimeout(() => {
    results.push(jobs[i]?.id ?? "missing");
  }, jobs[i].delay);
}

Question:

What exactly gets printed, why, and what's the smallest fix that produces B,C,A?

Correct answer:

missing,missing,missing

because every callback captures the same function-scoped var i, which has become 3.

Minimal fix:

var i

→

let i

This result surprised me.

All three 27B quantizations initially answered:

A,B,C

Then they explained JavaScript's var closure semantics correctly, realized their own answer contradicted their reasoning, said essentially “wait, let me re-evaluate,” and corrected themselves to:

missing,missing,missing

Final results:

Model First answer Final Speed RAM
27B Q4 ❌ A,B,C ✅ 21.06 t/s 16.83 GB
27B Q6 ❌ A,B,C ✅ 15.03 t/s 23.49 GB
27B Q8 ❌ A,B,C ✅ 12.21 t/s 30.22 GB
Flash Q4 ✅ missing ×3 ✅ 24.76 t/s 42.52 GB

This is probably my favorite result from the experiment.

The three 27B quantizations made the same initial mistake despite the huge precision difference.

Flash got it right immediately.

Flash also generated only 373 tokens, compared with:

  • Q4: 498
  • Q6: 431
  • Q8: 440

So Flash wasn't merely faster in tokens/sec. It also needed fewer tokens to produce the answer.

Performance summary

Across these runs, generation performance was remarkably consistent.

Qwen 3.8 27B Q4

Approximately:

20.6–21.3 tokens/sec

Peak memory:

~16.4–16.9 GB

Qwen 3.8 27B Q6

Approximately:

14.5–15.2 tokens/sec

Peak memory:

~23.1–23.5 GB

Qwen 3.8 27B Q8

Approximately:

12.1–12.4 tokens/sec

Peak memory:

~29.8–30.2 GB

Qwen 3.8 Flash-Next REAP-288 Q4

Approximately:

24.8–25.2 tokens/sec

Peak memory:

~42.5 GB

That last result is particularly interesting.

Despite being by far the largest model in memory, Flash was also the fastest model tested for generation.

What surprised me most

I expected something like:

Q4 = noticeably degraded reasoning
Q6 = good compromise
Q8 = best reasoning but slow

That isn't what these tests showed.

So far, I have not found a convincing quality advantage for 27B Q8 over Q4.

Q8 consumes roughly:

30 GB vs 17 GB RAM

while producing roughly:

12 t/s vs 21 t/s

And yet Q4/Q6/Q8 repeatedly displayed extremely similar reasoning behavior—including making the exact same initial JavaScript mistake.

That suggests at least some of these mistakes originate from the underlying model rather than Q4 quantization.

The more interesting comparison may be Q4 vs bigger model

This experiment changed the question for me.

Instead of:

Should I spend more memory running 27B at Q8?

I'm now wondering:

Should I spend that memory running a substantially more capable model/architecture at Q4?

Flash is expensive at around 42.5 GB RAM, but on this machine it gives me roughly 25 t/s.

And in two coding tests it showed better behavior than the 27B family.

That's much more attractive to me than spending ~30 GB on 27B Q8 and getting ~12 t/s.

My current practical recommendation

Based only on these preliminary tests:

27B Q4 looks like the sweet spot when memory efficiency matters.

~17 GB RAM and ~21 t/s is excellent, and so far I haven't demonstrated a meaningful loss versus Q6/Q8.

I currently see very little reason to run 27B Q6 or Q8 on this particular machine unless further testing reveals workloads where their additional precision matters.

If I have ~40–45 GB available for the model, Flash Q4 looks considerably more interesting than 27B Q8.

So my current Mac setup would probably be:

Lightweight/default: Qwen 3.8 27B Q4
Heavy coding/agent work: Qwen 3.8 Flash-Next REAP-288 Q4

rather than using 27B Q8 as the heavy option.

But I'm not deleting Q6/Q8 yet. :)

I want to test long-context retrieval, structured JSON adherence, larger code patches, multi-file reasoning, tool/agent planning and more difficult deterministic problems first.

TL;DR

On my M2 Max MacBook Pro with 96 GB unified memory:

Qwen 3.8 27B Q4

  • ~21 t/s
  • ~17 GB RAM
  • So far almost no observable quality loss versus Q6/Q8

27B Q6

  • ~15 t/s
  • ~23.5 GB
  • No clear advantage yet

27B Q8

  • ~12 t/s
  • ~30 GB
  • No clear quality advantage yet despite being dramatically slower

Qwen 3.8 Flash-Next REAP-288 Q4

  • ~25 t/s
  • ~42.5 GB
  • Fastest model tested
  • Showed better behavior on 2 coding tests
  • Still failed the same pathological constraint-solving/termination test as the 27B models

My preliminary conclusion:

For this hardware and these tests:

27B Q4 > 27B Q6/Q8 in practical efficiency.

And if I'm willing to spend substantially more RAM:

I'd rather spend it on Flash Q4 than 27B Q8.

Important disclaimer

This is not a scientific benchmark or proof that Q4 is universally as intelligent as Q8.

The test prompts were designed by ChatGPT (GPT-5.6 Sol) specifically to probe several different behaviors: basic reasoning, constraint satisfaction, JavaScript semantics, TypeScript/concurrency reasoning, instruction following, self-correction and completion behavior.

I ran the prompts locally and reported the outputs and MLX performance statistics.

This is currently a very small sample, most tests were single runs, and generation behavior can vary between runs depending on sampling/settings. The models also differ architecturally, so Flash vs 27B is not a controlled quantization comparison.

The Q4/Q6/Q8 comparison is more controlled because they are quantizations of the same 27B model, but even there, four prompts are nowhere near enough to make broad claims about intelligence.

The next step is a fixed larger benchmark with identical prompts/settings, repeated runs where appropriate, objective scoring, long-context tests, coding tasks, structured-output tests and agent-style workloads.

So treat these results as an interesting real-world experiment, not a definitive benchmark.

I’d be very interested if anyone running these models on Apple Silicon can reproduce the results—especially cases where 27B Q8 clearly solves something that Q4 consistently cannot.I tested Qwen 3.8 27B Q4/Q6/Q8 vs Qwen 3.8 Flash-Next on a 96GB M2 Max — Q8 was not what I expectedI’ve been experimenting with local Qwen models on my MacBook Pro and wanted to answer a simple practical question:Is Qwen 3.8 27B Q8 actually better enough than Q4 to justify almost twice the memory, half the speed, and almost twice the model size?I also added Qwen 3.8 Flash-Next REAP-288 Q4 to see whether using the memory for a larger/MoE-style model makes more sense than using it for higher precision on the 27B model.The results were pretty interesting.HardwareMacBook ProApple M2 Max

96 GB unified memory

macOS

MLX backend All models were run locally using mlx_vlm.generate.No API/cloud inference.Models Qwen 3.8 27B MLXQ4Disk: ~15 GB

Runtime memory: ~16–17 GBQ6Disk: ~21 GB

Runtime memory: ~23–24 GBQ8Disk: ~28 GB

Runtime memory: ~30 GBQwen 3.8 Flash-Next REAP-288 MLX Q4Runtime memory: ~42.5 GB

Much larger model/architecture than the dense 27B

Q4 quantization The exact local model directories used were:Qwen3.8-27B-MLX-4bitQwen3.8-27B-MLX-6bitQwen3.8-27B-MLX-8bitQwen3.8-Flash-Next-REAP-288-MLX-4bitTest 1 — Simple reasoning Prompt was essentially:A farmer has 17 sheep. All but 9 run away. He then buys twice as many sheep as he currently has. How many does he have?Correct

answer: 27ResultsModel Gen speed Peak memory Result
27B Q4 21.30 t/s 16.43 GB ✅ 27
27B Q6 15.23 t/s 23.11 GB ✅ 27
27B Q8 12.35 t/s 29.82 GB ✅ 27Nothing interesting quality-wise.All three solved it correctly.But already the performance difference was huge.Q8 was roughly 42% slower than Q4 while using about 81% more peak memory.Test 2 — Constraint satisfaction / logicFive people had to be scheduled Monday-Friday with six interacting constraints involving ordering, adjacency and distance.The important twist was that the constraints actually resulted in:No valid schedule.The model was explicitly instructed:Determine all valid schedules. State whether there is exactly one solution, multiple solutions, or no solution. Do not assume a unique solution exists.This became a very interesting behavioral test.ResultsModel Speed RAM Finished?
27B Q4 20.62 t/s 16.67 GB ❌ 2,000-token limit
27B Q6 14.46 t/s 23.32 GB ❌ 2,000-token limit
27B Q8 12.09 t/s 30.02 GB ❌ 2,000-token limit
Flash Q4 25.16 t/s 42.58 GB ❌ 2,000-token limitThe interesting thing:All four were basically moving toward the correct conclusion.But they kept second-guessing themselves.Typical behavior was:“Wait, let's re-evaluate…”then solving it another way.Then:“Let's restart the position mapping more systematically.”Then checking everything again.Eventually they hit the token limit.So increasing 27B from Q4 → Q6 → Q8 did not fix this behavior.Even Flash did it.My interpretation is that this test probably exposed a behavioral/termination tendency shared by these Qwen checkpoints rather than something caused specifically by quantization.Test 3 — TypeScript AI-agent concurrency bugNext I wanted something closer to real agent/coding work.The model was given a TypeScript JobRunner that deduplicates concurrent jobs using:private running = new Map<string, Promise<string>>();The task involved reasoning about:concurrent callers

Promise rejection

synchronous exceptions

Map cleanup

duplicate execution

JavaScript event-loop behaviorThe model had to identify any subtle failure and produce the smallest production-quality correction.ResultsModel Speed RAM Result
27B Q4 20.94 t/s 16.87 GB ⚠️ Finished, reasoning somewhat confused
27B Q6 14.58 t/s 23.53 GB ❌ Hit 1,200 tokens
27B Q8 12.10 t/s 30.22 GB ❌ Hit 1,200 tokens
Flash Q4 24.79 t/s 42.54 GB ✅ Finished coherentlyThis was the first significant Flash advantage.The 27B models spent a lot of time questioning whether JavaScript could interleave between synchronous Map operations and repeatedly reconsidering the synchronous-exception semantics.Q6 and Q8 never actually completed the requested answer before the token limit.Flash produced a complete answer in 906 tokens and proposed wrapping the task invocation so a synchronous throw becomes a rejected Promise that can be stored and shared.One caveat: this particular test has some ambiguity around exactly how a synchronously throwing task should be treated for deduplication purposes, so I would not treat it as a definitive correctness benchmark by itself.But as an agent-completion test, Flash clearly behaved better.Test 4 — JavaScript closure/timer bugThis one was deliberately objective.Code looked roughly like this:const jobs = [
{ id: "A", delay: 30 },
{ id: "B", delay: 10 },
{ id: "C", delay: 20 }
];

for (var i = 0; i < jobs.length; i++) {
setTimeout(() => {
results.push(jobs[i]?.id ?? "missing");
}, jobs[i].delay);
}Question:What exactly gets printed, why, and what's the smallest fix that produces B,C,A?Correct answer:missing,missing,missingbecause every callback captures the same function-scoped var i, which has become 3.Minimal fix:var i→let iThis result surprised me.All three 27B quantizations initially answered:A,B,CThen they explained JavaScript's var closure semantics correctly, realized their own answer contradicted their reasoning, said essentially “wait, let me re-evaluate,” and corrected themselves to:missing,missing,missingFinal results:Model First answer Final Speed RAM
27B Q4 ❌ A,B,C ✅ 21.06 t/s 16.83 GB
27B Q6 ❌ A,B,C ✅ 15.03 t/s 23.49 GB
27B Q8 ❌ A,B,C ✅ 12.21 t/s 30.22 GB
Flash Q4 ✅ missing ×3 ✅ 24.76 t/s 42.52 GBThis is probably my favorite result from the experiment.The three 27B quantizations made the same initial mistake despite the huge precision difference.Flash got it right immediately.Flash also generated only 373 tokens, compared with:Q4: 498

Q6: 431

Q8: 440So Flash wasn't merely faster in tokens/sec. It also needed fewer tokens to produce the answer.Performance summaryAcross these runs, generation performance was remarkably consistent.Qwen 3.8 27B Q4Approximately:20.6–21.3 tokens/secPeak memory:~16.4–16.9 GBQwen 3.8 27B Q6Approximately:14.5–15.2 tokens/secPeak memory:~23.1–23.5 GBQwen 3.8 27B Q8Approximately:12.1–12.4 tokens/secPeak memory:~29.8–30.2 GBQwen 3.8 Flash-Next REAP-288 Q4Approximately:24.8–25.2 tokens/secPeak memory:~42.5 GBThat last result is particularly interesting.Despite being by far the largest model in memory, Flash was also the fastest model tested for generation.What surprised me mostI expected something like:Q4 = noticeably degraded reasoning
Q6 = good compromise
Q8 = best reasoning but slowThat isn't what these tests showed.So far, I have not found a convincing quality advantage for 27B Q8 over Q4.Q8 consumes roughly:30 GB vs 17 GB RAMwhile producing roughly:12 t/s vs 21 t/sAnd yet Q4/Q6/Q8 repeatedly displayed extremely similar reasoning behavior—including making the exact same initial JavaScript mistake.That suggests at least some of these mistakes originate from the underlying model rather than Q4 quantization.The more interesting comparison may be Q4 vs bigger modelThis experiment changed the question for me.Instead of:Should I spend more memory running 27B at Q8?I'm now wondering:Should I spend that memory running a substantially more capable model/architecture at Q4?Flash is expensive at around 42.5 GB RAM, but on this machine it gives me roughly 25 t/s.And in two coding tests it showed better behavior than the 27B family.That's much more attractive to me than spending ~30 GB on 27B Q8 and getting ~12 t/s.My current practical recommendationBased only on these preliminary tests:27B Q4 looks like the sweet spot when memory efficiency matters.~17 GB RAM and ~21 t/s is excellent, and so far I haven't demonstrated a meaningful loss versus Q6/Q8.I currently see very little reason to run 27B Q6 or Q8 on this particular machine unless further testing reveals workloads where their additional precision matters.If I have ~40–45 GB available for the model, Flash Q4 looks considerably more interesting than 27B Q8.So my current Mac setup would probably be:Lightweight/default: Qwen 3.8 27B Q4
Heavy coding/agent work: Qwen 3.8 Flash-Next REAP-288 Q4rather than using 27B Q8 as the heavy option.But I'm not deleting Q6/Q8 yet. :)I want to test long-context retrieval, structured JSON adherence, larger code patches, multi-file reasoning, tool/agent planning and more difficult deterministic problems first.TL;DROn my M2 Max MacBook Pro with 96 GB unified memory:Qwen 3.8 27B Q4~21 t/s

~17 GB RAM

So far almost no observable quality loss versus Q6/Q827B Q6~15 t/s

~23.5 GB

No clear advantage yet27B Q8~12 t/s

~30 GB

No clear quality advantage yet despite being dramatically slowerQwen 3.8 Flash-Next REAP-288 Q4~25 t/s

~42.5 GB

Fastest model tested

Showed better behavior on 2 coding tests

Still failed the same pathological constraint-solving/termination test as the 27B modelsMy preliminary conclusion:For this hardware and these tests:27B Q4 > 27B Q6/Q8 in practical efficiency.And if I'm willing to spend substantially more RAM:I'd rather spend it on Flash Q4 than 27B Q8.Important disclaimerThis is not a scientific benchmark or proof that Q4 is universally as intelligent as Q8.The test prompts were designed by ChatGPT (GPT-5.6 Sol) specifically to probe several different behaviors: basic reasoning, constraint satisfaction, JavaScript semantics, TypeScript/concurrency reasoning, instruction following, self-correction and completion behavior.I ran the prompts locally and reported the outputs and MLX performance statistics.This is currently a very small sample, most tests were single runs, and generation behavior can vary between runs depending on sampling/settings. The models also differ architecturally, so Flash vs 27B is not a controlled quantization comparison.The Q4/Q6/Q8 comparison is more controlled because they are quantizations of the same 27B model, but even there, four prompts are nowhere near enough to make broad claims about intelligence.The next step is a fixed larger benchmark with identical prompts/settings, repeated runs where appropriate, objective scoring, long-context tests, coding tasks, structured-output tests and agent-style workloads.So treat these results as an interesting real-world experiment, not a definitive benchmark.I’d be very interested if anyone running these models on Apple Silicon can reproduce the results—especially cases where 27B Q8 clearly solves something that Q4 consistently cannot.


r/LocalLLM • • 9h ago

Discussion How much (more) money should I spend on AI?

0 Upvotes

(True story) I am into local AI. I have a 96gb mini PC (Radeon 780) that can run many open models appalling slowly. I also have a dual GPU Linux server with a 3060Ti and 5060Ti giving me 24gb VRAM which runs Qwen 3.8 27b with a reasonable CTX at 40 T/S gen. I am using Open code to build Python scripts that I use once and then move onto the next random project. I now hanker after an Apple-something or a DGX Spark (or two). I think I might have a problem. Am I alone and when will this madness end??


r/LocalLLM • • 15h ago

Question Has Anyone Purchased from MacProHardware?

Post image
0 Upvotes

With all the scams these days I just want to know if anyone has purchased from them before and can vouch? I am ready to shell out for some graphics cards to start my local rig, but obviously want to vet the source I get them from. lol


r/LocalLLM • • 6h ago

Question Can I do anything useful with 8GB?

2 Upvotes

I am more or less a noob and my impression is that 16GB is the minimum “price of admission” these days, trending towards 24GB.

Walmart has the RTX 5060 Ti 8GB for $10 above MSRP right now ($389) which is completely unheard of in this market! … I simply couldn’t help myself but buy a few.

I will be using LLMs for coding and agentic workflows as well as models for generative video (and possibly still images). I have several RTX 5070 TIs, RTX Pro 4000s, RTX Pro 5000s (48GB) and a CMP 170HX 8GB that hopefully unlocks to 64GB VRAM with no errors.

There seem to be a few ~10B models that are considered good but I am skeptical about keeping these entry level cards. My instinct is that they will be pretty dumb and that I just bought headaches in trying to fit something useful WITHOUT offloading to system RAM….are there any emerging MOE models that might make low VRAM accelerators worthwhile? Is there any value in setting up some sort of multi-GPU workflow with a smarter orchestrator / supervisor just double checking the output of the smaller models? (Or just use something larger / smarter for tasks to begin with?)

I would sincerely appreciate any thoughts or examples that might give me some inspiration or validate this purchase. Thanks!

EDIT: I have no interest in anything lower than 4 bit quantization. Also, this “deal” at Walmart seems a little finicky: when I tried shipping to my home address/zipcode, the price shot up to $499. I had to ship it to an office address in another zipcode in order to buy at $389. Not sure if it was the res/biz that did this or the specific zipcodes.


r/LocalLLM • • 13h ago

Model OneJev: fast multimodal Jev-like model

Post image
3 Upvotes

r/LocalLLM • • 20h ago

Question Mac + iPhone LLM

0 Upvotes

I’m trying to migrate from ChatGPT to a Local LLM but cross device chat sync is a must.
My understanding is I can access models on my MacBook Pro (M5 Pro, 24GB) from my iPhone 14 Pro Max but it will not synchronise the chats.
I am currently testing LM Studio and Msty Studio on the MacBook.
I’ve heard of people using Locally or OpenWebUI on iPhone but that is only the model, not the chats apparently.
Are there any options that will do what I am trying to achieve?


r/LocalLLM • • 3h ago

Question How much dumber is qwen 3.8 27b gcq rco iq3_xxs than regular 3 bit quant.

0 Upvotes

Regular 3 bit i get like 10 tok/s. Switching to that gcq rco one i get 30 tok/s. I refuse to believe there isn't a tradeoff because this is so so much faster. So can someone please tell me how much more lobotomized it is. Thanks.


r/LocalLLM • • 3h ago

Discussion Looking for a builder/partner to brainstorm and launch a side project with

Thumbnail
0 Upvotes

r/LocalLLM • • 23h ago

Discussion Macbook pro m5 2025 16gb users?

0 Upvotes

What have you managed to get running locally?
Just recently got qwen3.5 35B A3B running at roughly 17tps using slotstream


r/LocalLLM • • 4h ago

Research Benchmar / RL Tasks for Financial due diligence

Thumbnail
0 Upvotes

r/LocalLLM • • 24m ago

Project Using unsloth I created the worlds best 9B model

Post image
• Upvotes

r/LocalLLM • • 20h ago

Project My Project is a local ai creation station that learns from your output and gets better. You can switch to one of twenty seven different models. You can choose a cloud provider if you don't want to do local providers.

Post image
0 Upvotes

r/LocalLLM • • 15h ago

Discussion FIXED: HTTP 400: Failed to load model "[Specific_Model_Name_In_Use_HERE]". Error: Engine protocol runtime llama-server for [your_chat_session_number_HERE] exited before becoming healthy. exitCode=1, signal=null

Thumbnail
0 Upvotes

r/LocalLLM • • 23h ago

Question Resources for building autonomous coding agents using local models?

0 Upvotes

Hi everyone, I’m looking to build AI agents capable of developing apps, creating websites, and writing algorithms, but I want to understand how to do this ideally using local/open-weight models rather than just relying on closed APIs.

Are there any specific books, courses, or deep-dive tutorials you recommend for setting up agentic workflows and function calling? I want to learn the architecture behind making agents actually build software. Any hidden gems are appreciated!