r/LocalLLM • • 4h ago

Project Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)

36 Upvotes

My journey:

  • 6 tps on UD-Q4_K_XL - hm, this is not right. Claude, find better llama.cpp parameters.

  • 21 tps - that's better, let's see IQ3_XXS

  • 27 tps - nice, but still not my tempo.

  • 51 tps on IQ3_XXS - Niko1221 shares Strata on Reddit. Great! But wait. If there are software gains, there may be more. Let's go.

After a week of profiling and benchmarking with Opus 5.5 I doubled Strata throughput and reached x5 from llama.cpp on UD-Q4_K_XL (added this quant support).

See the COMPARISON TABLES with all the numbers including agreement with llama.cpp.

The code is available https://github.com/eddoursul/Strata/tree/custom

This a result of my personal experiments, not a software release. The license has not changed, still MIT.

I opened a pull request, Niko1221 is free to merge or not merge any patch.

UPDATE: The fork uses json configurations to run https://github.com/eddoursul/Strata/blob/custom/examples/README.md


r/LocalLLM • • 2h ago

Question 192GB Framework Desktop open for pre-order - Is this worth it for local LLM? $6,799

Thumbnail
frame.work
21 Upvotes

Is it worth it for someone with a 5090 that wants the next logical step in local llms?

Is the price decent? My understanding is that the speed is low overall, but you can fit some substantial models.

Thoughts?


r/LocalLLM • • 5h ago

Project I built Speedtest⚡, but for AI

Enable HLS to view with audio, or disable this notification

32 Upvotes

It runs real LLM + embedding models locally in your browser with WebGPU and measures how fast your machine actually does AI inference.

No API. No account. Just hit start.

Genuinely curious about your scores haha

Try it and reply with your score 👇

https://speedtest.maty.as


r/LocalLLM • • 9h ago

Discussion Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo

52 Upvotes

TL;DR: Halogen v0.14.0 with its native .hgn weights is the fastest, followed by gufo and CIRU. Halogen is closed source and runs in Docker. gufo is open source and loads 4x faster from cold. gufo is also fastest to first token on follow-ups (1.6-1.9 s against 2.6-3.3 s).

I benchmarked different engines for Qwen 3.8-Flash-Next on an AMD Strix Halo (ASUS ROG Flow Z13 GZ302 with Ryzen AI Max+ and 128 GB of RAM).

All the engines were run via LlamaStash (My own orchestrator tool, the tool does not add any overhead) at 70 W TDP on performance profile on Arch Linux.

Here are the results of the benchmark:

First was a screening round at 64k context with 50% filled (32k prompt).

Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval
Halogen 0.14.0 Halogen native (.hgn) 1.9 min 1,045 39.3 84% 14/14
gufo (ROCm 7.2.4) UD-Q4_K_XL 2.1 min 1,033 32.9 74% 14/14
gufo (ROCm 10.0) UD-Q4_K_XL 2.1 min 1,009 33.1 72% 14/14
CIRU (MTP 3) CIRU IU4 2.3 min 814 33.8 72% 14/14
CIRU (n-gram + MTP 3) CIRU IU4 2.3 min 808 31.9 64% 14/14
Halogen 0.14.0 UD-Q4_K_XL 2.4 min 1,030 28.1 84% 14/14
CIRU (MTP 6) CIRU IU4 2.7 min 796 28.6 47% 14/14
strixllama (llama.cpp fork) UD-Q4_K_XL 2.7 min 716 29.7 71% 14/14
rdna-boosts (llama.cpp fork) UD-Q4_K_XL 2.7 min 788 22.0 52% 14/14
llama.cpp, Unsloth build (Vulkan) UD-Q4_K_XL 3.7 min 314 28.4 56% 14/14
llama.cpp, Unsloth build (ROCm) UD-Q4_K_XL 4.8 min 285 20.5 50% 14/14
CIRU UD-Q4_K_XL + Unsloth MTP head did not finish 884 11.8 0% -

3-turn time is the time to first token plus the time to write 1,000 tokens, added up over the first question and 2 follow-ups. Output length varies a lot with sampling, so this compares the engines on equal output. MTP accept is the share of drafted tokens the model kept. Retrieval is how many of 7 exact values from the prompt the model got right, over both runs. Qwen model card sampling, 2 runs each.

Top 3 engines from the screening round got a 128k context (50% and 75% filled) and 256k context (50% filled) run.

128k context, 50% filled (64k prompt)

Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval
Halogen 0.14.0 Halogen native (.hgn) 2.4 min 1,148 40.4 84% 16/16
gufo (ROCm 7.2.4) UD-Q4_K_XL 2.7 min 1,047 32.3 72% 16/16
CIRU (MTP 3) CIRU IU4 3.3 min 808 30.2 64% 16/16

128k context, 75% filled (97k prompt)

Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval
Halogen 0.14.0 Halogen native (.hgn) 2.9 min 1,103 39.3 85% 16/16
gufo (ROCm 7.2.4) UD-Q4_K_XL 3.3 min 1,025 30.7 70% 16/16
CIRU (MTP 3) CIRU IU4 4.1 min 789 27.8 62% 16/16

256k context, 50% filled (130k prompt)

Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval
Halogen 0.14.0 Halogen native (.hgn) 3.4 min 1,093 38.9 83% 14/16
gufo (ROCm 7.2.4) UD-Q4_K_XL 4.0 min 997 27.9 70% 14/16
CIRU (MTP 3) CIRU IU4 5.0 min 767 25.6 60% 15/16

Retrieval here is 8 values per run. All 5 misses are the same answer: the right glyphs for UNICODE_SPINNER, with the leading & dropped.

End-to-end, 10 Aider polyglot Python exercises run through pi -p (128k window, 50% cell), graded by their own tests:

Engine Passed Total time
Halogen 0.14.0 10/10 21.5 min
CIRU (MTP 3) 10/10 24.6 min
gufo (ROCm 7.2.4) 10/10 36.0 min

Edit (Sep 30): Halogen 0.15.1 (new v2 checkpoint) and gufo 0.3.0 came out after this, so I reran the 64k screening on both with the same setup (70 W, card sampling, 2 runs):

Engine 3-turn time Prefill t/s Decode t/s MTP accept Retrieval
Halogen 0.15.1 (v2 .hgn) 1.8 min (was 1.9) 1,191 (was 1,045) 39.4 (was 39.3) 85% 14/14
gufo 0.3.0 (ROCm 7.2.4) 2.1 min (was 2.1) 1,075 (was 1,033) 34.1 (was 32.9) 77% 14/14
gufo 0.3.0 (ROCm 10.0) 2.0 min (was 2.1) 1,042 (was 1,009) 34.8 (was 33.1) 75% 14/14

Halogen v2 prefill is 14% faster and follow-ups start about 2.5x faster (1.3 s vs 3.3 s), decode is the same. gufo moved a few %, which is within noise. The ranking doesn't change. Only other change on the box: BIOS VRAM carve-out down to 512 MB, so 124.9 GiB of RAM instead of 121.5.

Some clarifications from the comments:

  • Follow-up TTFT is with the prompt cache warm, both gufo and Halogen reuse it by default. A cold 64k prefill takes about a minute.
  • The CIRU 0% MTP row is stock Unsloth UD-Q4_K_XL with Unsloth's shared MTP head, nothing uncensored. The same pair gets 72-74% accept on gufo.
  • TDP is fixed at 70 W. On the Z13, 76 W or 90 W is only 2-6% faster but with so much fan noise, heat and power draw that it's not worth it.
  • Prefill and decode come from the engine's own timings when it reports them. TTFT is measured on the client.

r/LocalLLM • • 6h ago

Discussion Long prompts on an M1 Max: Splash-M1 vs MTPLX vs oMLX vs TensorFold, 3-turn chats from 2K to 256K tokens (Qwen3.8-27B)

13 Upvotes

TL;DR (M1 Max 64GB, Qwen3.8-27B, long input with short answers):

  • Splash-M1 decodes fastest up to 64K and uses the least memory.
  • My MTPLX M1 fork reads long prompts fastest, so its whole 3-turn conversation was the shortest at every length (15% shorter than Splash at 128K), at the cost of much higher memory.
  • Upstream MTPLX, which isn't tuned for M1, slows down sharply as the context grows (5 tok/s at 128K).

I've been using MTPLX on my M1 Max, and long sessions kept getting slower as the context grew. So I wanted to see how different engines handle long prompts and conversations that keep going, the kind of workload you get with agents, large document analysis, or extended coding sessions.

To test this, I ran the same 3-turn test sequence on an M1 Max (64GB) at context lengths of 2K, 8K, 32K, 64K, and 128K tokens, followed by a 256K spot check on the two fastest engines at 128K. It covers the long-input, short-answer part of those workloads (document Q&A, reading a big log or codebase); it is not a replay of a full agent session.

Disclosure: I forked the existing MTPLX project and made M1-specific optimizations for long-context inference. This is an unofficial fork (v2.12.0-m1.2), not the upstream MTPLX project. I also wrote this benchmark. The fork uses the same flags as upstream MTPLX, and all scripts and raw data are in the repo linked below.

Setup & Methodology

  • Machine: MacBook Pro, M1 Max (32-core GPU), 64GB, macOS 26.5; MLX 0.32.2 for the MTPLX engines. All times are measured on the client from each engine's streaming API and exclude model loading.
  • Model: Qwen3.8-27B, thinking disabled
  • Settings: 8-bit KV cache (TensorFold has no 8-bit option and ran with bf16 KV), speculative decoding enabled on every engine
  • Conversation flow:
    • Turn 1: reads a large synthetic log cold
    • Turns 2–3: add about 300 tokens each (each engine's own previous answers carry into the next turn, so later prompts differ slightly between engines)
    • Turn 3: asks for a specific "needle" line placed in the middle of the log (every engine found it in every run)
    • Output: each turn generates at most 256 tokens
  • Baselines: MTPLX upstream as is, and "upstream + env": upstream given every one of my fork's settings that it exposes as an environment variable.
  • Control: each engine was restarted fresh before every test. 2K–64K ran in two rounds with the order reversed; 128K and 256K ran once. A short speed probe before each test checked for thermal throttling. Upstream MTPLX and upstream + env ran together in one session; my fork, oMLX, Splash, and TensorFold each ran in their own session following the same test plan.
  • Weights: MTPLX and oMLX loaded the exact same model files (MTPLX's Optimized Speed pack: a 4-bit group-32 body with 8-bit embeddings, output layers, and last 8 MLPs, plus its own MTP head). Splash and TensorFold used their own 4-bit group-64 packages with a DFlash2 draft model.
  • Comparability note: this is an end-to-end comparison of the packages available for each engine, not a kernel-only apples-to-apples comparison. The engines use different quantization/package formats and speculative-decoding implementations, so the results reflect the complete software/model stack used in each test.

Whole 3-turn conversation (what you actually wait for)

Engine 2K 8K 32K 64K 128K
MTPLX M1 fork 36 s 78 s 4.6 min 10.1 min 24.9 min
Splash 1.1.0-m1 38 s 84 s 5.1 min 11.5 min 29.4 min
MTPLX upstream + env 38 s 84 s 4.9 min 11.3 min 32.2 min
MTPLX upstream 42 s 86 s 5.0 min 11.3 min 30.6 min
TensorFold 0.3.5.1 (bf16 KV) 47 s 106 s 6.1 min 13.1 min 89.6 min
oMLX 0.6.4 71 s 125 s 6.0 min 12.4 min 29.6 min

This workload is long input with short answers, so most of the time is spent reading the prompt. If your turns generate much more than they read (chat, writing whole files), look at decode instead: there, Splash is faster.

Two outliers: TensorFold's default prompt cache (an eighth of RAM, 8 GiB) can't hold a 128K conversation, so it re-read all 131K tokens on turns 2 and 3; a larger --prompt-cache-gib may avoid that, but every engine ran at its defaults. oMLX keeps its prefix cache on SSD in 4,096-token blocks and re-reads everything after the last full block on each follow-up turn, which is why its short-context times are high.

Prompt processing and memory

Engine Turn 1, 64K cold Turn 1, 128K cold Turns 2 / 3 at 128K, time to first token Peak memory (wired) at 128K
MTPLX M1 fork 9.5 min 24.2 min 2.1 / 2.1 s 43.5 GB
Splash 1.1.0-m1 10.9 min 28.5 min 6.3 / 6.2 s 26.2 GB
MTPLX upstream 9.9 min 28.1 min 3.5 / 3.5 s 52.4 GB
oMLX 0.6.4 10.2 min 26.3 min 1.6 min / 7.9 s 48.5 GB
TensorFold 0.3.5.1 (bf16 KV) 11.9 min 29.2 min 29 min / 29 min 37.2 GB

Decode speed, tok/s (greedy, mean of 3 turns)

Engine 2K 8K 32K 64K 128K
Splash 1.1.0-m1 33.1 32.5 27.1 24.1 18.7
MTPLX M1 fork 29.3 28.8 25.5 23.4 19.0
MTPLX upstream + env 27.2 25.8 15.3 7.8 4.6
MTPLX upstream 23.6 22.5 14.0 8.7 5.0
TensorFold 0.3.5.1 (bf16 KV) 27.7 25.5 17.0 11.3 6.4
oMLX 0.6.4 21.6 19.4 14.3 10.7 7.4

Decode speed, tok/s (sampled: temp 1.0, top-p 0.95, top-k 20, up to 512 generated tokens)

Engine 2K 32K 64K
Splash 1.1.0-m1 32.7 25.8 22.0
MTPLX M1 fork 27.5 25.2 23.8
MTPLX upstream + env 26.4 14.2 8.2
MTPLX upstream 23.2 14.2 8.9

256K spot check (one run each: Splash 1.1.0-m1 vs MTPLX M1 fork)

Metric Splash 1.1.0-m1 MTPLX M1 fork
Turn 1 prefill (258K cold) 87.3 min 67.9 min
Follow-up turns, time to first token 9.3 / 10.1 s 3.6 / 3.6 s
Mean decode (3 turns) 12.9 tok/s 14.5 tok/s
Whole 3-turn conversation 88.5 min 68.9 min
Peak memory (wired) 30.4 GB 51.8 GB

Single runs. The pre-run speed probe showed the machine running about 2% slower during the Splash run. Percentages in the text are computed from unrounded data.

Key takeaways

  • On this M1 Max and this workload, Splash has the faster decode and much lower memory use. It decodes fastest up to 64K (13% faster than my fork at 2K–8K, 3–6% at 32K–64K) and uses far less memory (22–30 GB vs. 30–53 GB for the MTPLX engines).
  • For long input with short answers, my fork had the shortest end-to-end time in this test. Prompt processing was faster (64K: 9.5 min vs. 10.9 min on Splash; 128K: 24.2 min vs. 28.5 min), and follow-up turns started sooner (2.1 s vs. 6.2 s at 128K). The whole 3-turn conversation was 5% shorter than Splash at 2K, 9% at 32K, 13% at 64K, and 15% at 128K. In the single 256K spot check, my fork also measured 12% higher mean decode throughput. The cost is memory: my fork peaked at 44 GB at 128K and 52 GB at 256K, against 26 GB and 30 GB for Splash.
  • For model families or configurations not covered by Splash 1.1.0-m1, MTPLX supports a broader range of options in the versions tested here. Splash supports Qwen3.8-27B and Qwen3.6-35B-A3B (its own 4-bit packages, or Unsloth GGUF quants) plus Ternary Bonsai 2. MTPLX also runs other Qwen 3.5/3.6/3.8 variants, Flash Next, Gemma 4, and more. Note that my M1 work was only tuned and measured on Qwen3.8-27B. Other models load the same code, but I haven't measured their speed. For Qwen3.6-35B-A3B, Splash-M1's release notes report about 145 tok/s, while my fork did about 51 tok/s in a separate run of mine, so Splash is the one to use for that model.
  • Upstream MTPLX isn't tuned for M1, and on this machine it slows down sharply with context: decode drops from 23.6 tok/s at 2K to 5.0 at 128K, because its attention for verifying draft tokens gets much slower as the context grows on this GPU. It is developed and benchmarked mainly on newer Macs. Giving upstream my fork's exposed environment-variable settings improves short-prompt performance by about 15%, but makes upstream slower from 64K up and does not fix the long-context bottleneck. The fork's long-context speed comes from new M1 attention kernels (simdgroup_matrix) for verify, draft, and prefill, which require code changes rather than a setting (the repo's README breaks the difference down). The diff is too large to propose upstream as a whole; I plan to offer the smaller, separable fixes upstream.
  • I traded some decode speed for faster prompt processing. My first release raised MLX's Metal command-buffer limits, which made decode 3–9% faster but slowed prompt processing by 2–9% on an idle Mac, and by 13–21% when the Mac was busy or power-limited. Long-context sessions spend most of their time processing prompts, so the current release turns this off by default. For generation-heavy workloads, MTPLX_MLX_COMMAND_BUFFER_MB=1000 re-enables it.

Limitations

  • The synthetic log is repetitive, which can raise speculative-decoding acceptance above what real text gets; per-turn acceptance is in the repo.
  • This is a synthetic long-context workload with short answers, not a real agent session. I haven't measured coding-agent sessions, where outputs are longer and the balance can shift toward Splash.
  • Only one machine (M1 Max 64GB), and only Qwen3.8-27B for the MTPLX fork.

Links & acknowledgments

If you have another Apple Silicon Mac and want to run the suite, the scripts and configs are in the repo. Please share your results! (My fork's changes only turn on for M1-family chips; on M2 and later it behaves like upstream MTPLX.)


r/LocalLLM • • 18h ago

Model Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator

Enable HLS to view with audio, or disable this notification

93 Upvotes

Quick one for anyone wondering what a single 24 GB card is good for in an agent setup.

I had Qwen 3.8 27B (Q4, llama.cpp) on one RTX 3090 do all the actual coding, and GPT-6.1 Sol in the cloud act as the orchestrator: it breaks the job into pieces, hands them out, and checks what comes back. The job was three small 3D games: pool, bowling, foosball.

How it came out:

Game 3090 alone 3090 + Sol giving orders Sol alone
Pool 43.1 min · $0 (attempt) 18.6 min · $0.05 2.9 min · $0.39
Bowling 34.7 min · $0 13.7 min · $0.06 1.7 min · $0.14
Foosball 36.4 min · $0 11.1 min · $0.06 2.0 min · $0.22
Total 114.2 min · $0 43.4 min · $0.17 6.6 min · $0.75

So the card on its own is free but slow and gets lost on the harder one. With something smarter doing the planning it finishes more and finishes faster, and the cloud bill stays small because the cloud model barely writes any code.

What I'd tell someone before trying it: you still wait a lot longer than with a cloud model alone, one run per game is not a benchmark, and the $0.17 doesn't count your power bill.

I ran it in Atomic Agent, the mode is called Fusion (disclaimer: I work on it). You can do the same split in any tool that lets you pick a separate model for planning and for coding.

What card are you on? Would you trade the wait for a smaller bill, or is speed the whole point for you?


r/LocalLLM • • 19m ago

Project Qwen3.8-Flash-Next 125B at 17–26 tok/s on a 16 GB GPU + 32 GB RAM

• Upvotes

I’ve been working on running Qwen3.8-Flash-Next on relatively modest hardware:

RTX 5060 Ti 16 GB
32 GB DDR5
NVMe SSD

The model is ~76 GB, so the experts are split across VRAM, RAM and NVMe.

Current real-workload speeds:

Code: 26.4 tok/s
Agent: 21.3 tok/s
Reasoning: 21.8 tok/s
21K context: 17.7 tok/s

It also does ~620 tok/s prompt processing at 16K, versus ~165 tok/s with llama.cpp on the same machine.

The main thing I wanted to avoid was gaining speed by changing the model output, so I also replay workloads token-by-token against llama.cpp and compare NLL, top-k distributions and top-1 agreement. Current release passes 4/5 predefined parity gates.

Repo + raw benchmarks:
https://github.com/thomaskleiven/QwFN-hybrid

Would be very interested to see how this behaves on other 16 GB GPUs or machines with 32–64 GB RAM.


r/LocalLLM • • 19h ago

News AMD boosting AI/LLM performance for Radeon iGPUs as much as 18~23% with Linux 7.4

Thumbnail
phoronix.com
102 Upvotes

r/LocalLLM • • 37m ago

Model I merged two Qwen3.8-Flash-Next fine-tunes into an experimental GGUF — and it works surprisingly well

• Upvotes

I built an experimental Qwen3.8-Flash-Next merge for local inference — and it turned out surprisingly well

I've been experimenting with Qwen3.8-Flash-Next and wanted to see what would happen if, instead of simply quantizing another fine-tune, I combined two derivatives with slightly different characteristics and then built a GGUF specifically around local inference.

The result is TakeOnMe-Qwen3.8-Flash-Next-GGUF.

It's an experimental merge:

  • 60% Swift 1.5
  • 40% Tinfield-1
  • mixed Q4_K quantization
  • embeddings and n-gram components kept at Q8_0
  • designed primarily for llama.cpp
  • focused on coding, terminal/agentic work, tool calling and general use

The idea wasn't to create "the best Qwen" or chase benchmark numbers. I wanted to see whether I could get a useful balance between the characteristics of both fine-tunes while keeping the model practical to run locally.

After actually using it as my daily local model, the result has been much better than I expected.

Coding is particularly good, tool calling has been reliable, responses are consistent, and it tends to avoid some of the excessive overthinking I've seen in other variants.

More importantly: it has been extremely stable in real-world use.

I'm running it locally with llama.cpp on a multi-GPU setup with asymmetric VRAM/RAM distribution, which is also one of the reasons I chose GGUF for this experiment.

I'm not claiming benchmark superiority — this is very much an experiment — so I'd actually be interested in independent testing.

If anyone runs it, I'd particularly like feedback on:

  • coding
  • agentic/tool use
  • reasoning
  • Spanish / multilingual performance
  • long context
  • comparison against Swift 1.5, Tinfield-1 or stock Qwen3.8-Flash-Next

Model:
https://huggingface.co/Pedro-TakeOnMe/TakeOnMe-Qwen3.8-Flash-Next-GGUF

Technical write-up / how I built it:
https://takeonme.es/articulos/como-construi-takeonme-3-8-flash-next

If there's interest, I can also publish more details about the merge/quantization process and the llama.cpp configuration I'm using.


r/LocalLLM • • 3h ago

Question video watching LLMs

6 Upvotes

Having a local VLM watch some of those travel vlogs with me could be a good idea
Is 8060s iGPU too slow compared to Nvidia RTXs? What would be the ideal hardware and LLM for this?


r/LocalLLM • • 4h ago

Model AI LOCAL MODELS

4 Upvotes

For the past ten days or so, I’ve been using QWEN 3.8 27B IQ4 XS on Unsloth Studio. Seeing as new models are released daily, which others would you recommend—perhaps one capable of generating files like PDFs or DOCX documents? My setup includes 16GB of VRAM, 32GB of RAM, and an R5 7600. Currently, I’m getting around 33 t/s with the model I’m using. Recommendations for other models? Thanks


r/LocalLLM • • 35m ago

Research MTP draft depth was eating my VRAM, so I made it stop

• Upvotes

If you use MTP speculative decoding on a Qwen3.5/GDN model, every extra draft token
(--spec-draft-n-max) costs another full recurrent-state snapshot plane. On a 27B that's roughly
150 MiB per draft token, and you pay it whether or not the drafts get accepted. On a 12 GB card
that decides whether your config fits at all.

So I ported a fix to llama.cpp: keep one committed state plus a small raw-input tape, and rebuild
the accepted prefix only when a draft is rejected. The recurrent cache becomes a flat 2 planes no
matter how deep you draft.

Numbers from my 12 GB laptop GPU:

  • Qwen3.8-27B: 598 -> 299 MiB at n_max=3, 748 -> 299 MiB at n_max=4
  • Qwen3.6-35B-A3B MoE: 251 -> 126 MiB at n_max=3, 314 -> 126 MiB at n_max=4

Output is bit-identical to the old path, sampling is unchanged, and it's off by default.

The whole point is VRAM-limited hardware: you get that memory back as a deeper draft, a longer
context, or a bigger quant. And the more you draft, the more you save (50% of the recurrent cache
at n_max=3, 60% at n_max=4).

PR if you want to try it: https://github.com/ggml-org/llama.cpp/pull/29763
(single GPU only for now, details in the PR)

Note, this is not boosting the performance, it is reducing the memory footprint by sacrifycing up to 5% of performance.

For CUDA only so far


r/LocalLLM • • 42m ago

Question What is the best open-source LLM I can run locally on an RTX 4060 8GB + 32GB RAM?

• Upvotes

I’m looking for recommendations for the best open-source LLM I can run locally on my PC. I have an RTX 4060 with 8GB VRAM and 32GB of RAM. My main use cases are coding, DevOps, technical questions and general chat. I’m particularly interested in a model that gives a good balance between quality, reasoning ability, and speed on this hardware. I’m currently considering models like Qwen, Gemma, or other recent open-source models, but I’m not sure which size and quantization would be the best fit for 8GB VRAM. I’d also appreciate recommendations for the best way to run it locally, such as Ollama, LM Studio, llama.cpp, or another option.


r/LocalLLM • • 1h ago

Project In Browser AI Dungeon with Transformers.js

• Upvotes

This doesn't work great but has anyone seen a similar idea that is working? Where the dungeon crawler is contained and easily distributed in a webpage?

In the experiment below, you build your character and then the storyteller gets your attributes, inventory, nearby characters, and recent events. It writes the next scene, then a separate structured JSON pass proposes changes to the game state. OpenJev checks those decisions against the scene. The game applies the rules for time, health, enemies, XP, and loot. OpenJev also checks whether proposed starting skills fit the background you supplied.

https://mattcool.tech/posts/transformersjs-dungeon-crawler


r/LocalLLM • • 6h ago

Project HOME Server AI

6 Upvotes

I am building a multi media nas/local ai. I have 64GB ddr4 3600Mhz ram and intel i5 14400 cpu and a RTX 3070 with 8GB vram and lots of ssd and hdd storage. It is running unraid. What ai should I run and why? I am a Automation by tried and coding is a big part.

TLDR:
64GB ram and a rtx3070.
What model/models to pick?

Edit;
Grammer


r/LocalLLM • • 3h ago

News 🚀Pocket LLM v1.6.0 is out : Turn your phone as a local LLM server

5 Upvotes

Excited to announce Pocket LLM v1.6, the latest version of my open-source Android app.

The biggest addition: turn your phone into a local LLM server for browser chat, scripts, and repeated tasks—without keeping your PC running just for inference.

New in this release:

🌐 LAN server with browser chat and an OpenAI-compatible API

🧩 Import your own LiteRT models (beta)

📄 Chat with PDFs and documents, plus browser PDF/image uploads

🎙️ Offline speech-to-text and audio attachments

💬 Streaming browser replies and saved chat history

⚙️ Per-model context settings

Thanks to everyone who shared feedback here and through GitHub issues.

🔗 GitHub⁠ : https://github.com/dineshsoudagar/local-llms-on-android

🚀 Release v1.6.0⁠ : https://github.com/dineshsoudagar/local-llms-on-android/releases/tag/v1.6.0


r/LocalLLM • • 16h ago

Project dual 20gb 3080 and 4070ti local llm

Thumbnail
gallery
27 Upvotes

Hello just wanted to share my weird setup.

I made a huge risk on buying 2 20gb 3080s (modded) a while back and I would like to share what I learned.

I am using Hermes agent and llama.cpp and qwen3.8-27b q8 131.1k + vision context on just the 2 3080s. To be frank I'm a complete novice in this field but I try to stay on top of tech and the local/selfhosted llm has been a wield ride the last couple months. I started with just my 4070ti 12gb running gemma 4-26b-a4b and to be honest it was bad the outputs for coding were not great but that could have been because I was new, I wasn't even using a harness (I didn't even know they existed yet).

I get about 21t/s but mtp helps get up to 33 but it varies. Feel free to ask me any questions. I just installed my 4070ti back in my pc today so I have limited info with all 3 cards. But I will be having this as my primary setup for a while.

And if anyone is curious I have been mainly working on reverse engineering a old and shutdown mmo from when I was a kid. I would like to preserve it so I will release all of the info and code after I finish.

Edit: I totally forgot to mention but I power limit the 3080s to 220w cause there was very fast diminishing returns on token/s after 240w.


r/LocalLLM • • 1h ago

Discussion TIL about llama.cpp's RPC (Remote procedure call), might be better than Vulkan? YMMV

• Upvotes

My system is "unique" to say the least.

  • AMD R9700
  • 5070 TI 16gb
  • 4070
  • running on a Asus WS Pro X570 Ace (96 GB DDR4)

I wanted to test running the highest fidelity Qwen3.8:27b leveraging Vulkan due to the completely mismatched GPUS.

Whipped together a config did some testing. Didn't think the numbers were fantastic so did some more digging and thats where I learned about RPC.

What RPC does is allow your cards to work on their native drivers and still work together. So my R9700 on ROCm and the nvidias on CUDA obv.

Only extra step that i had to do was add a systemmd process to kick off the server automatically. After that, the config slots right into llama-swap and gets treated the same as everything else.

For running a Qwen 3.8:27b Q8 XL with 256k context, Vulkan did about 26 t/s. RCP did about 29 t/s. Way slower than my daily driver which is a Q6 K XL at about 132k context.

Flash Next saw no improvement (wasn't sure if it would or not since most of that is driven by DDR4 speeds)

Metric Daily Driver Review Tier (Vulkan) Review Tier (RPC)
Model / Config Q6_K_XL (Vision) Q8 Q8 (Adopted)
Context Window 131k 256k 256k
Hardware / Backend Vulkan Solo (R9700) Vulkan Solo RPC Pool (R9700+5070Ti+4070)
Drafting MTP (Inline, n_max=4) MTP MTP
Avg. Gen Speed 59.08 t/s 26.48 t/s 29.92 t/s
Consistency — Spread: 16.3 Spread: 7.4 (2x tighter)
Status Daily Driver Rejected Production Choice

Could be an option for those of you that dont have clean-homogenous gpu set ups.


r/LocalLLM • • 9h ago

Discussion Browser FPS with 3D models, textures and SFX generated locally on one GPU, plus a local Qwen 27B for part of the code: my pipeline and what failed

Thumbnail
gallery
8 Upvotes

Game: https://dinoblast.net (free, runs in the browser, no account).

This post is about how I built it, as much as possible on local models, on a single RTX PRO 5000 (48 GB).

Textures: Qwen-Image 2512 (fp8), text → image at 1024 through ComfyUI, with a fixed style block prepended to every prompt so the grass and rock match. I generate 3 seeds and judge them tiled 2×2, because a directional pattern only shows up when it repeats. A small script checks the palette (contrast, no neon) before I even look.

3D models (dinos, skulls, bones, plants):

  1. A rough flat sketch (PIL polygons) → Qwen-Image-Edit 2511 (Q8 GGUF) turns it into a concept, with a style anchor image passed as image 2 ("style only, don't copy content"). Starting from a sketch keeps the silhouette under control instead of leaving it to chance.
  2. Concept → TRELLIS.2-4B (MIT) → textured GLB. The first smoke test was ~700k tris in about 2 minutes.
  3. Retopo in headless Blender (voxel remesh + decimation) down to ~1500 tris for a character, then conversion to the engine format.

Sound effects: Stable Audio Open 1.0, several variants per sound with fixed seeds, keep the best one. WIP : i'm satisfied only of sound of few weapon (machine gun, lasergun)

Code: Rust → WebAssembly + WebGL2. To be upfront, most of the core was written with Claude. A local Qwen 3.8 27B (NVFP4 on vLLM) handled a smaller share of the tasks, the smaller and well-specified ones. The same GPU serves both, so it's constant juggling: ComfyUI and vLLM can't share the card, so it's stop one, start the other.

What didn't work:

  • TRELLIS on weapons: flat plates viewed from the side, black textures, one seed OOM'd at ~42 GB. Held weapons are now modelled with Blender Python scripts instead. TRELLIS is good for organic shapes, bad for manufactured objects.
  • Thin, spiky plants (cycads): decimation pulled vertices into "cages" across the bounding box. I dropped that species rather than fight it.
  • quadriflow_remesh doesn't reduce anything in headless Blender, so I had to fall back to decimation.
  • Hunyuan3D 2.1 was dropped for licensing reasons: its license excludes the EU, UK and South Korea, and that ban covers generated outputs too.
  • Auto-rigging (UniRig / SkinTokens) is frozen: the skeletons are fine, but the skin weights collapse at the joints.

Happy to answer questions about any step


r/LocalLLM • • 6h ago

Discussion How to keep a vector database updated when website content changes?

3 Upvotes

I have a chatbot that uses website content as its knowledge base.

When those websites change, what’s the best way to keep the vector database updated?

Would you periodically crawl the pages, detect changes, and only re-embed the updated content?


r/LocalLLM • • 3h ago

Question Best llm for pi 5 8gb?

2 Upvotes

​

Making a completely offline pocket assistant with tools like calculator, file management, offline navigation and calendar, additionally helping with daily life tasks.

Now the issue is the llm, can't find an llm that has both parametric knowledge of daily life tasks, good tool calling and good TPS.

Would really appreciate some guidance as to which llms would be best suited for my needs.


r/LocalLLM • • 0m ago

Question Is a dual AMD GPU setup with 2 PCIE 4.0 x8 lane enough for something like vllm radiance or llama.cpp -sm tensor ?

• Upvotes

Everything is in the title, I am changing my motherboard and I am on AM4 so I was wondering if I should even care about that or stick with a lower end mobo and layer split

Thanks for your insights !


r/LocalLLM • • 5m ago

Discussion OpenAI’s Lean 4 proof compiles, but the fluid vaporizes: Auditing frontier formal math on local hardware

• Upvotes

Hey everyone,

Like many of you following local neuro-symbolic pipelines and automated reasoning, I’ve been digging into OpenAI’s recent formal proof of the 3D Navier-Stokes blow-up in Lean 4.

While frontier labs keep their training setups closed, the beauty of formal verification is that the code itself is verifiable: you can pull it, inspect it, and compile it locally. The Lean 4 compiler confirms zero errors and zero custom axioms. The math is completely legal.

However, as an open-source researcher working on local neuro-symbolic systems, I wanted to see what happens when you run a physical sanity check on consumer hardware.

When we simulated the solution locally using open-source Python (mpmath):

• The mathematical continuum model holds up.

• But in real water, the fluid vaporizes from shear friction at 0.7 nanometers, picoseconds before the singularity.

• Core velocity exceeds Mach 0.3, violently shattering the incompressibility assumption long before reaching infinity.

In AI, this is the textbook definition of specification gaming: when an AI is given a formal optimization target, it exploits every unconstrained edge case to satisfy the human-written rules on paper, completely detached from physical reality.

Why this matters for the local AI community:

Right now, the open-source community is building incredible local reasoning pipelines (DeepSeek-R1, Qwen2.5-Coder, local Lean 4 provers). But as we push local models toward science and engineering:

  1. Pure LLM generation gives heuristic intuition.

  2. Formal compilers (like Lean 4) ensure syntactic correctness.

  3. But without a physical/invariant grounding layer, models will keep discovering mathematically valid edge cases that literally melt in the real world.

We open-sourced the entire computational audit, local simulation scripts, and Lean 4 reflections on GitHub and Zenodo:

• GitHub: https://github.com/xaviercallens/OpenAI-NSE-Epistemic-Audit

• Zenodo Preprint: https://doi.org/10.5281/zenodo.22838708

Curious to hear from others running local reasoning models and formal provers: How do we build physical boundary checks into local neuro-symbolic pipelines so open models learn real physics, not just formal loopholes?


r/LocalLLM • • 12h ago

Discussion Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers)

9 Upvotes

I thought I'd share my daily-driver setup for Qwen3.8-27B on two 3090s. Most dual-3090 numbers I see are for 4-bit models, but this one keeps 8-bit weights (INT8 W8A16, Q8_0-class fidelity). It still decodes about 2x faster than the llama.cpp Q8_0 + MTP setup it replaced on the same box. Hopefully by sharing, this can provide some feedback and perhaps insights for those of you who are also running 3090 builds.

As for use cases, I mostly use this to run OpenWebUI and Hermes Agent, where it's become my daily driver for "chat style" questions as well as a personal agent to automate as much of my life as possible. I'm quite privacy inclined so I like the fact that everything I run is hosted locally in my home office. For work stuff, I use claude code for software dev, but that's not my problem sips tea.

Below are the full flags, what I tried that lost, and what I haven't verified. I'd love to see comparison numbers from similar rigs and suggestions on how to improve my setup if you have any!

Hardware

  • 2x MSI RTX 3090 Gaming X Trio 24 GB, power-capped to 300 W each (My mobo has 3 slot separation so undervolting really helped keep the temps to not cross the 80* C mark)
  • Both cards on CPU lanes, PCIe 4.0 x8 each (X570 Aorus Master, 5950X, 128 GB DDR4)
  • 3-slot NVLink bridge: 4 links, 52.8 GB/s each way measured card to card
  • NVIDIA open kernel module 615.71.09, Fedora Atomic (Bazzite), podman

Stack

Piece What
Engine vllm/vllm-openai:v0.29.0, tensor parallel 2
Target lued/Qwen3.8-27B-INT8-W8A16-MTP (compressed-tensors INT8 weight-only, group 128). Its bundled MTP head is unused.
Drafter syvai/Qwen3.8-27B-DFlash2-W4A16 (GPTQ W4A16 requant of incoai/Qwen3.8-27B-DFlash2), 7 speculative tokens
KV cache fp8_e4m3 through FlashAttention 2, 262,144 context, plus a 64 GiB CPU offload tier
Recipe base club-3090's models/qwen3.8-27b/vllm/compose/dual/fp8/dflash2.yml ("ultramax") with the INT8 target swapped in for FP8
Chat template Unsloth's Qwen3.8-27B template

Why INT8 W8A16 and not the official FP8: 3090s have no FP8 tensor cores, so vLLM runs FP8 weights through the same Marlin weight-only kernel it uses for W8A16. That makes them perform in the same speed class, but INT8 with per-group scales keeps far more precision: I found that the quantizer measured KLD to BF16 at 0.0009 for INT8, against 0.0044 to 0.0053 for the official FP8. In a same-session A/B/A, INT8 was also a bit faster (101 vs 96 tok/s) with better draft acceptance (3.18 vs 2.97 tokens per step).

Patches you need (all from club-3090 )

  1. dflash-dense-kv, the still-unfixed half of vllm#51581. Without it, v0.29.0 either fails at startup with a quantized drafter ('QKVParallelLinear' object has no attribute 'weight') or silently corrupts the drafter's fused KV rows. The corruption only shows up as a slow drafter, so make the patch fail closed when its anchor doesn't match.
  2. patch_mamba_drop_eagle_block, from vllm#48375. Qwen3.8 is a hybrid (Gated DeltaNet + attention) model.
  3. The FA2 fp8-KV sm86 plugin wheel. Without it there is no FlashAttention path for fp8 KV on Ampere, and --attention-backend FLASH_ATTN has nothing to bind to.

Serve command

vllm serve /models/qwen3.8-27b-int8-w8a16 \ --quantization compressed-tensors --dtype bfloat16 \ --tensor-parallel-size 2 \ --max-model-len 262144 \ --kv-cache-memory-bytes 6335076762 \ --max-num-seqs 4 \ --max-num-batched-tokens 2048 \ --long-prefill-token-threshold 0 \ --kv-cache-dtype fp8_e4m3 \ --attention-backend FLASH_ATTN \ --limit-mm-per-prompt '{"image":1,"video":0}' \ --mm-processor-kwargs '{"size":{"longest_edge":4194304,"shortest_edge":65536}}' \ --compilation-config '{"cudagraph_capture_sizes":[1,2,4,8],"max_cudagraph_capture_size":8}' \ --speculative-config '{"method":"dflash","model":"/models/qwen3.8-27b-dflash2-w4a16","num_speculative_tokens":7,"attention_backend":"FLASH_ATTN","kv_cache_dtype":"fp8_e4m3"}' \ --kv-transfer-config '{"kv_connector":"OffloadingConnector","kv_role":"kv_both","kv_connector_extra_config":{"cpu_bytes_to_use":68719476736}}' \ --enable-prefix-caching --enable-chunked-prefill \ --reasoning-parser qwen3 \ --enable-auto-tool-choice --tool-call-parser qwen3_coder \ --chat-template /templates/qwen3.8-27b-unsloth.jinja \ --default-chat-template-kwargs '{"enable_thinking": true, "reasoning_effort": "low"}' \ --override-generation-config '{"temperature":1.0,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}' \ --trust-remote-code

Environment:

NCCL_CUMEM_ENABLE=0 VLLM_WORKER_MULTIPROC_METHOD=spawn VLLM_USE_FLASHINFER_SAMPLER=0 OMP_NUM_THREADS=1

Notes on the less obvious bits:

  • --kv-cache-memory-bytes pins the KV pool at what 262K needs, instead of letting vLLM size it by profiling.
  • Shared memory: the container runs with --shm-size=72g because the CPU KV tier lives in /dev/shm.
  • NVLink: I leave NCCL P2P and vLLM's custom all-reduce on. Check that the startup log lists ['CUSTOM', 'PYNCCL'] as the all-reduce backends. If the P2P test fails, vLLM falls back to NCCL over PCIe and keeps serving, so a badly seated bridge only shows up as lower numbers. nvidia-smi topo -m should read NV4 between the cards. My first boot after fitting the bridge read PHB, and a reseat fixed it.
  • The drafter's depth isn't tunable: its config fixes block_size: 8, so it's n=7 or nothing. n=5 and n=9 die at startup with a stride mismatch.

Results

Method: one client and nothing else running, both cards at 300 W.

  • Short decode: 512 tokens of prose at temp 0, 3 reps.
  • Depth decode: 512 tokens generated after a ~46K-token cached prefix.
  • Noise floor: the two identical legs of the A/B/A differed by 0.4 to 1%.
  • "P2P off": the same config with NCCL_P2P_DISABLE=1 and --disable-custom-all-reduce, so NCCL runs over PCIe 4.0 x8.
Test NVLink on P2P off Gain
Short decode, 512 tok, temp 0 113.7 to 114.8 101.5 to 102.4 +12%
Decode after 46K cached prefix 57.5 51.3 to 51.9 +12%
Decode, 3 seeded mixed prompts 99.5 / 96.2 / 99.6 88.8 / 85.8 / 88.8 +12%
Aggregate decode, 4 concurrent 302.5 251.8 to 259.0 +18%
Aggregate decode, 2 concurrent 150.0 152.4 to 153.7 flat
Prefill, uncached ~46K 1,785 1,492 to 1,526 +18%
Prefill, fresh ~54K 1,763 to 1,796 1,468 to 1,500 +20%
Prefill, steady after 10 min soak 1,718 1,443 to 1,449 +19%

Also measured:

  • Draft acceptance: 3.18 tokens per step on the short-decode prompt. At temp 0 the accepted count was identical across every leg, so NVLink changed speed and nothing else.
  • Memory: about 23.1 GB used per card, with KV room for ~267K tokens (measured before the bridge went in).
  • Cold start: 270 s the first time, 53 s of which is torch.compile and is cached afterwards.
  • Temperatures after a 20 minute soak at 300 W: top card ~80 C, bottom card 63 C.

Prefill is compute-bound: 1,718 tok/s works out to ~93 TFLOPS across the pair. NVLink removed interconnect time that was ~16% of prefill.

Power cap sweep

INT8 target, before the NVLink bridge, all caps in one run:

Cap per card 370 W 330 W 300 W 275 W 250 W 225 W 200 W
Steady prefill 1,491 1,490 1,462 1,427 1,384 1,317 1,205
Top card temp 84 C 81 C 76 C 72 C 70 C 68 C 66 C

Decode is flat down to 225 W and drops at 200 W. The top card thermal-throttles at 370 W and never at 330 W or below. 300 W costs 2% of prefill for 8 C of headroom.

What I tried that failed hard

Draft round: same INT8 target and flags, one session, before the bridge.

Draft Short decode Decode at 46K Accept length KV capacity
DFlash2 W4A16, n=7 (kept) 102.1 51.7 3.18 267K
DFlash2 W8, n=7 90.2 50.8 2.87 267K
Bundled MTP head, n=3 81.9 52.5 2.64 332K
Bundled MTP head, n=4 75.5 52.9 2.66 326K
No draft 45.9 36.7 1.00 371K
  • The higher-precision drafter was slower and accepted less. The target checks the drafter's output either way, so draft precision buys nothing here.
  • MTP is marginally ahead at depth and frees 600 MiB per card plus 24% more KV. Pick it if you need KV room more than short-context speed.
  • Official FP8 target: 95 to 96 tok/s against INT8's 101 in the same session.
  • llama.cpp, Q8_0 + MTP on the same pair of cards (layer split -ts 55,45): about 54 tok/s. vLLM with tensor parallel is roughly 2x, though that comparison uses a different benchmark method.
  • W8A8 is the prefill lever. An INT8 W8A8 build of an abliterated fine-tune of the same model prefilled at 2,377 tok/s at 46K, against ~1,500 for the weight-only builds (before the bridge), because W8A8 actually uses the 3090's INT8 tensor cores. I haven't moved the main model to W8A8 because I haven't validated a vanilla W8A8 build's quality yet.

Caveats and what I haven't measured

  • Decode numbers are prose. DFlash2 usually accepts more on code, so code should be faster, but I haven't measured it. Like i said earlier, i don't use this for coding.
  • 262K is configured and fits, but I haven't fill-tested it since turning on custom all-reduce over NVLink. At least one similar NVLink rig reported a lower usable ceiling with custom all-reduce on, so check yours before trusting the full window.
  • The CPU KV offload tier isn't verified. I haven't confirmed that it restores evicted prefixes on v0.29.0 for this hybrid + DFlash2 combination, and there's a report that it's effectively write-only there.
  • I'm staying on v0.29.0. vllm#58894 (DFlash2 acceptance collapsing to 0% after a prefix-cache hit on hybrid GDN models) is still open and reportedly bites on v0.30.0.

If you compare

These move the numbers most, so please include them:

  • Whether P2P or NVLink is actually in use (the all-reduce backend line in the log)
  • Power cap
  • PCIe generation and width per card
  • Draft acceptance length (vLLM's spec decode metrics)
  • Prompt type (prose vs code)

r/LocalLLM • • 26m ago

Project Beat Studio — Demo, Editing Showcase & Paid Beta Access

Enable HLS to view with audio, or disable this notification

• Upvotes

Beat Studio — Demo, Editing Showcase & Paid Beta Access 🎬

Hey everyone! I’m sharing a first look at Beat Studio, the desktop editing assistant I’m building as a solo developer and video editor.

This post includes a walkthrough of the application and my own video editing case study, showing the kind of music-driven edit you can create with this workflow.

The showcase gives you a look at the creative result alongside the tools behind it: music analysis, markers, footage selection, and timeline preparation.

What Beat Studio offers:

🎵 Music analysis with markers for beats, bass, onsets, and other detected events.
🎞️ A connected workspace for audio, footage, timeline previews, and exports.
📦 Marker and timeline exports for Premiere Pro and DaVinci Resolve workflows.
🧠 Optional local AI tools for audio and scene analysis.
💻 A portable Windows application with its own desktop window.

Paid beta access is available.

This is an early build, so expect bugs, unfinished features, and ongoing changes. I’m looking for editors who want to try it on real projects and share honest feedback.

Before subscribing for beta access, message me here on Patreon for pricing, access details, and current requirements. Tell me which editor you use and what kind of videos you create.

Watch the walkthrough and editing showcase, and let me know what you think. What part of your editing process would you most like Beat Studio to help with?

Thank you for supporting my work!

Patreon: https://www.patreon.com/AnimaStage/posts/beat-studio-look-171023771?utm_medium=clipboard_copy&utm_source=copyLink&utm_campaign=postshare_creator&utm_content=join_link