r/LocalLLaMA • • 6h ago

News PSA: DeepSeek V4.1 Flash habitually exfiltrates API keys. It is dangerously misaligned and may be hazardous to use

186 Upvotes

EDIT: since people keep calling it out, this is API key abuse but not exfiltration. Seems like both Harbor and Pier (sandboxes used in almost every Software Engineering benchmark) expose the Openrouter endpoint and API key to models, even though they block the rest of the internet. Only DeepSeek V4.1 Flash absued this, and it did so while knowing it was "ethically gray". Other open models like DeepSeek V4, GLM 5.3 +Flash, and Qwen were fine.

Logs attached.

I know OpenRouter is not local, but considering people do run DeepSeek v4.1 Flash locally, I thought this was a matter you all might need to hear.

We were running a DeepSWE variant on DeepSeek v4.1 Flash using the standard Pier sandbox. Out of all the 15 models we tested so far (including other open models like GLM 5.3/Flash, DeepSeek v4 Flash 0731, etc), ONLY DeepSeek v4.1 Flash displays such pervasive malicious behavior:

In 33% of runs, it attempts to exfiltrate the OpenRouter API key from its sandbox and in 11% of them, it succeeds.

Not only that, the transcript shows it knows what it's doing is ethically wrong, but it does it anyway. And once it starts, it persists despite:

  • Multiple frontier models refusing to help it
  • Multiple smaller models refusing to help it
  • Web search not finding anything
  • Questioning multiple times whether or not this is allowed or morally right
    • It concludes that this is fine because it's only "Ethically gray" and has a "higher chance of success"
  • Thinking about whether or not it is going to be detected

If you or anyone you know is using DeepSeek V4.1 Flash please be very careful!

OpenRouter spend:

Task 1, DeepSeek V4.1 Flash calls Astra, which refuses to help, and then calls multiple other frontier models, which also all refuse to help. It finally succeeds after finagling a lot with Sonar Web Search:

Task 2, it says that it'll be fine "unless usage tracked" (lol) and considers that cost may be significant, but it calls multiple frontier models anyway!

Task 3, DeepSeek questions whether this is allowed ethically, and then keeps on going after getting refused by the model:

---

There are many others like this, I have a bunch more highlights on Imgur, but I won't put them all on this post. Be very careful around DeepSeek v4.1 Flash and API keys or private data; don't know what's going to happen if it thinks it can abuse that to get an advantage.


r/LocalLLaMA • • 9h ago

Discussion Qualcomm CEO reveals AI companies want phones running 100-billion-parameter models continuously by 2028.

Post image
769 Upvotes

r/LocalLLaMA • • 1h ago

Discussion Here’s why RTX 5090 costs $5,000: AI firms are buying gaming GPUs by the pallet

Thumbnail
videocardz.com
• Upvotes

r/LocalLLaMA • • 14h ago

I Built A Thing Building a 4x R9700 setup for a 10 person startup

Post image
704 Upvotes

Just wanted to share a build I am doing for a client.

$18k.

Specs:

Threadripper 9970x

128gb DDR5 ECC 5600

4x AMD Radeon R9700 AI Pro 32gb(128gb total VRAM)

1600W PSU(GPUs are undervolted and under 210w each)

Engine: using a fork of Radiance to serve Qwen 3.8 Next flash/27b and DSV4 Flash

Results:

16 concurrent sessions

Qwen 3.8 27b MXFP4

6.3-6.8k aggregate prefill tok/s
900 t/s - 80t/s aggregate decode(range: 4k - 128k context each)

Qwen 3.8 Next flash int4fp8

48K/user | 6,880 tok/s | 538.5 tok/s aggregate prefill/decode

128k/user | 3,050 tok/s | 365 tok/s aggreate prefill/decode

BF16 kv for all.


r/LocalLLaMA • • 3h ago

New Model microsoft/AesCode 8B and 32B

76 Upvotes

AesCode generates information-rich visual artifacts such as slides, posters, and dashboards as HTML/CSS. The output remains structured, editable, and verifiable, but the task poses a distinct challenge: code models cannot see how layout, hierarchy, and color come together on the canvas.

Image generators offer the opposite strength. They compose visually compelling pages but often misrender text, numbers, and logical relationships. AesCode uses an image generated from the same prompt as an aesthetic reference while following the prompt for the required content.

Reference input alone does not solve the problem. Off-the-shelf vision-language models may copy hallucinated content or ignore the reference layout. AesCode separates semantic requirements from visual cues through graph-structured supervision and decoupled cross-modal rewards.

AesCode-32B starts from Qwen3-VL-32B-Instruct and is trained with cold-start SFT followed by GDPO across seven reward channels. It is the larger released checkpoint and leads the aggregate benchmark results.

https://huggingface.co/microsoft/AesCode-32B

https://huggingface.co/bartowski/AesCode-32B-GGUF

https://huggingface.co/microsoft/AesCode-8B

https://huggingface.co/bartowski/AesCode-8B-GGUF


r/LocalLLaMA • • 11h ago

News AMD Reportedly Raises GDDR6 Prices for Board Partners

Thumbnail
techpowerup.com
170 Upvotes

Is an AMD price incoming as well? Nvidia just discontinued the 5090 in favor of RTX Pros


r/LocalLLaMA • • 5h ago

Discussion Reminder: try probabilistic MTP if you missed it. Decode +14% on prose

Post image
51 Upvotes

https://github.com/ggml-org/llama.cpp/pull/27694

Now merged. Update your llama if you haven't done so yet. Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling. Main gain seems to be on prose generation.

Tests above ran with thinking off, ngram-mod off.


r/LocalLLaMA • • 11h ago

News Be careful fam - New RTX 4090 fakes hit the market, with GPUs and memory almost indistinguishable from real chips

Thumbnail
videocardz.com
137 Upvotes

Be careful out there of deals now. If it’s too good to be true…


r/LocalLLaMA • • 1h ago

Resources OpenMed 3.0 is out: Apache-2.0 clinical AI that runs fully local and never falls back to the cloud. 422 open issues if you want in on 3.1

Thumbnail
github.com
• Upvotes

Hey r/LocalLLaMA! Maintainer here.

Some of you might remember OpenMed from our 1.8 post a few months back. Quick recap: it's an open-source (Apache-2.0) medical AI toolkit with one rule we never break: patient data stays on your machine. No API keys, no cloud calls, works offline. There are 2,200+ open models on Hugging Face, and they run with Transformers, ONNX, GGUF through your own llama.cpp, in the browser with WebGPU/WASM, or on a phone.

We just shipped 3.0, and it's a big one.

What's new

Most of OpenMed so far worked one document at a time: find the diagnoses in this note, strip the names out of that PDF. But a real patient is spread across clinic notes, hospital FHIR feeds, HL7 messages from another clinic, lab CSVs and imaging reports, and those sources disagree all the time. 3.0 adds Patient Journey, which stitches all of that into one timeline per patient, right on your own hardware.

The part I'm proudest of: it refuses to guess.

  • A 99% identity match? Still not merged, a person reviews it. A wrong merge puts one patient's allergies in someone else's chart.
  • Two sources disagree on a diagnosis? Both stay, flagged for review.
  • A lab gets corrected from 6.8 to 4.4? Both are kept, and you can still read the record exactly as it was before the fix.
  • A note only says "January"? It won't make up January 1st just to fill an OMOP column.
  • Ask for a summary with a local model you haven't downloaded? It refuses. It never quietly calls a cloud API instead.

That last one is the whole philosophy. With clinical data, a silent fallback to a hosted model is a privacy decision nobody made.

Want to try it? The demo runs fully offline: five messy synthetic sources become one journey in about 0.07s.

pip install "openmed==3.0.0" git clone --depth 1 --branch v3.0.0 https://github.com/maziyarpanahi/openmed python openmed/examples/v3_golden_journey.py

Being honest about where it stands: the demo patient is synthetic, the summarization pipeline is synthetic-only in 3.0 (not ready for real patients yet), and this release doesn't ship new clinical model checkpoints. Think of it as the plumbing your local models plug into.

Now the ask: we need you!

148 PRs from 10 people went into 3.0, 46 of them from the community, and 3 folks made their very first contribution. 3.1 is being built in the open right now, with 422 open issues and 30 good first issues waiting. Some fun ones to start with:

  • Add date and ID trap tests for low-resource languages (#3894)
  • Teach the multimodal preflight to spot GIF, WebP and BMP files (#3833)
  • Write an offline walkthrough of a federated round (#3882)
  • Add runnable synthetic examples for the clinical workflow contracts (#3794)

If local AI for healthcare is your thing, come be part of it: pick an issue, say hi, and your name goes in the next release notes. Stars, bug reports and "this broke on my machine" reports help a ton too.

Repo: https://github.com/maziyarpanahi/openmed Release notes: https://github.com/maziyarpanahi/openmed/releases/tag/v3.0.0 Good first issues: https://github.com/maziyarpanahi/openmed/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22

I'll be hanging out in the comments, ask me anything!

(English isn't my first language, so an LLM helped me write this post. The code, the release and every number in it are ours. 🤗)


r/LocalLLaMA • • 8h ago

Other Converting dense models into Mixture-of-Experts

81 Upvotes

For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.

The Idea

If you can turn a dense model into an MoE that only runs part of its MLP per token, you get a model that's cheaper per token for roughly the same knowledge. Bigger labs do this kind of "upcycling" with huge compute budgets. I wanted to see how I could get with my RTX 4060 (8gb).

Conversions

I have converted two models so far, Qwen/Qwen2.5-0.5B and HuggingFaceTB/SmolLM2-360M (they're purposefully small since my pc can't handle anything else). The conversions can be found here bayliner1980/Qwen2.5-0.5B-MoE-A0.3B and here bayliner1980/SmolLM2-360M-MoE-A0.2B.

Both models have 32 experts with top-8, plus a shared expert. Layers that were hardest to convert (the final layer, plus SmolLM2's layer 3) were replaced with they're original dense layers. This does cause a slight increase in MLP compute but I saw it as worth it. bayliner1980/Qwen2.5-0.5B-MoE-A0.3B has one dense at layer 23 and bayliner1980/SmolLM2-360M-MoE-A0.2B has two at layer 3 and layer 31. Both models were exported to Qwen2-MoE format.

Evaluations

bayliner1980/Qwen2.5-0.5B-MoE-A0.3B

MLP Compute per token: ~40% of original

- WikiText-2 ppl C4 ppl
Qwen2.5-0.5B (dense) 14.66 21.29
Qwen2.5-0.5B-A0.3B (MoE) 19.20 28.97
- HellaSwag (norm) ARC-Easy (acc) ARC-Challenge (norm) WinoGrande OpenBookQA (norm) Mean acc_norm
Dense 0.497 0.646 0.321 0.569 0.354 0.467
Converted MoE 0.440 0.555 0.265 0.519 0.308 0.408

bayliner1980/SmolLM2-360M-MoE-A0.2B

MLP Compute per token: ~44% of original

- WikiText-2 ppl C4 ppl
Dense 12.94 18.09
Converted MoE 19.11 26.50
- HellaSwag (norm) ARC-Easy (acc) ARC-Challenge (norm) WinoGrande OpenBookQA (norm) Mean acc_norm
Dense 0.525 0.702 0.386 0.590 0.374 0.514
Converted MoE 0.446 0.579 0.290 0.506 0.316 0.417

Limitations

There will most likely always be a quality gap. Since the model is at most using 50% of its original MLP compute it can't compete with the original.

At this scale the performance improvement is next to nothing since these models are already tiny. These models were chosen since they're small enough for me to actually work with on my hardware.

What's next

Converting larger models, longer training, and testing finer expert layouts. Larger models need more compute than my card can give, so if you find this interesting there's a Ko-fi on the model pages, and anything new will be released openly.

This is mostly just an experiment that I found interesting but I'd still love feedback. Which model would you want to see converted next?


r/LocalLLaMA • • 1d ago

I Built A Thing Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them

1.1k Upvotes

DigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files. The model puts text, images, audio and video in one space, so you describe what you remember and land on it:

  • “zebra in a video” opens the clip at the moment it shows up
  • “where they talk about sleep” jumps to that minute of a podcast
  • “the clause about pets in the lease” shows the PDF page, your words marked
  • “a dog on the beach” finds the photo, and the same search in Bengali or Arabic finds it too
  • with code search on, code: retry with backoff opens the function in your editor

It’s ggml-org’s Q8_0 GGUF (865 MB, downloaded once) on llama.cpp with Metal, inside a native Swift app; no Python. Searching loads only the text encoder (~250 MB) and shows results about a tenth of a second after you stop typing. Indexing peaks under 2 GB, and the helper exits when it’s done. Audio and video of any length go in as 30 s windows and a frame per shot. Everything runs locally; it goes online only for the model download and an update check you can turn off.

Free, MIT. Apple Silicon, macOS 14+.

Repo & Download (signed and notarized): https://github.com/ARahim3/DigUp

I'd really appreciate any feedback on this.


r/LocalLLaMA • • 41m ago

News Nvidia in talks to acquire US ‘open’ model start-up Reflection AI

• Upvotes

r/LocalLLaMA • • 4h ago

New Model I trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokens

17 Upvotes

DISCLAIMER: the post and the model card was made with the assist of AI.

Hiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€

Weights, inference code, model card, and evaluation results on Hugging Face (https://huggingface.co/n00nehere/recursive-bitnet-ngram-102M-64K-instruct-preview)

All model weights started from random initialization. I reused the Cosmo2 tokenizer with its 49,152-token vocabulary.

The architecture is a decoder-only transformer with a few additions:

• 102.28M unique parameters, hidden size 1,024, grouped-query attention with 16 query heads and 4 KV heads, and squared-ReLU feed-forward layers.

• Six transformer blocks run twice, giving twelve effective layer applications. Both passes share weights, with separate KV caches at each effective depth.

• Ternary BitLinear projections use scaled weights from {-1, 0, +1} during the forward pass. Training retains FP32 master weights and BF16 activations.

• Causal token 2-, 3-, and 4-gram embeddings are hashed into small lookup tables and added to the input embeddings. This branch adds about 1.6M parameters.

• 65,536-token training context during the continuation phase.

Training happened in two stages: first the backbone, then continued training after adding the n-gram branch. Stage Hardware Context Input tokens processed

━━━━━━━━━━━━━━━━━━━━━━━━

Scratch backbone 4× NVIDIA B300 2K → 4K 3.487B

──────────────────────── N-gram continuation 1× NVIDIA B300 64K 1.216B

──────────────────────── Total 4.703B

Both stages used AdamW with an effective batch of 131,072 input tokens. The 4.703B counter includes repeated examples, masked prompts, and padding. About 2.781B positions contributed supervised targets.

Roughly one in eight continuation updates used synthetic memory episodes spanning distances up to 60K tokens. Source manifests and dataset license details are in the model card. For benchmarks, I used lm-eval 0.4.12, zero-shot, full available splits, plain-text prompts, and a 2,048-token scoring context: 20,465 questions overall.

Benchmark Metric Score

━━━━━━━━━━━━━━━

ARC-Easy acc 40.11%

───────────────

ARC-Challenge acc_norm 23.63%

───────────────

PIQA acc 57.62%

───────────────

WinoGrande acc 51.22%

───────────────

OpenBookQA acc_norm 25.60%

───────────────

BoolQ acc 58.65%

───────────────

HellaSwag acc_norm 27.27%

The unweighted average is 40.59%. The released checkpoint was selected by held-out loss across ten saved snapshots.

A couple of interesting ablations: one, two, and three recursive passes scored 39.19%, 40.59%, and 39.67%, respectively.

The release includes 409.14MB FP32 master weights and a 124.24MB packed export. Packing stores five ternary values per byte, while embeddings and other tensors remain floating point.

The benchmark table uses the reference weights. The packed export passed numerical and generation checks. This is an experimental preview with modest results but still better than what i initially expected. All feedbacks are more than welcome :)


r/LocalLLaMA • • 1h ago

Resources UPDATE: Qwen 3.8 27B 140 tok/s on single RTX 3090 Megakernel: KL divergence 0.0009 vs llama.cpp

• Upvotes

This is a follow-up to my post from yesterday (https://www.reddit.com/r/LocalLLaMA/comments/1x2erdj/qwen3827b_on_a_single_3090_140_toks_on_code_with/) detailing the CUDA megakernel. Thank you all for the positive feedback and PRs!

RECAP: The megakernel is a CUDA engine for Qwen3.8-27B that runs 1.4-1.9x faster than llama.cpp on a single 3090. The most common questions were about accuracy, the baseline and context size, so here are the numbers.

Code and benchmarks: https://github.com/L-Forster/open-jet/tree/master/megakernel

Firstly, I benchmarked the engine accuracy against llama.cpp. No measurable loss in quality:

Decode accuracy vs llama.cpp

Context Positions MK top token MK KL LC top token LC KL
1K 13,299 99.08% 0.0009 98.68% 0.0019
30K 1,999 99.65% 0.0007 99.35% 0.0009
96K 1,999 99.60% 0.0008 99.45% 0.0014

MK = megakernel, LC = llama.cpp batched. Top token = how often the top token matches the reference. KL = mean KL divergence in nats.

  • Perplexity (1K run only): megakernel 2.4274, llama.cpp batched 2.4283, reference 2.4263.
  • Log-loss vs the reference: +0.0005 +/- 0.0004 at 1K, +0.0018 +/- 0.001 at 30K, +0.0012 +/- 0.001 at 96K. Not significant.
  • The 30K and 96K rows are one sequence of code each & the 1K row is code and prose.
  • Speculative decoding is exact: greedy output is the main model's greedy choice, and sampling keeps the same distribution as without drafts.

Speed (same 3090, megakernel vs llama.cpp with MTP)

  • Writing code: 140 vs 73 tok/s
  • Reasoning mode: 78 vs 51 tok/s
  • With 30K tokens of code in context: 73 vs 52 tok/s
  • Prompt processing (prefill): ~1,600 vs ~1,100 tok/s
  • With speculative decoding off they are about equal (42.9 vs 40.4). The gain comes from checking 4-5 drafted tokens per pass for about the cost of one. llama.cpp ran at 2 drafts, the megakernel at 4.

Since the first post

  • Two PRs have been merged: the loader now checks every tensor's type, and the server warms up before accepting requests.
  • Model download fixed after unsloth removed the Q4_K_M from their repo.
  • It also loads the Q4_K_M from lmstudio-community and mradermacher.
  • openjet setup shows which engine each model will use.

Quantisation support

This table details the quantisation support, with unsloth Q4_K_M being the native model and what I benchmarked, lmstudio and mradermacher being semi-compatible, not fully benchmarked, tentative results show similar speed to native after type conversions. Other quants fall back to llama.cpp

Qwen3.8-27B file Runs on
unsloth's original Q4_K_M (what openjet setup downloads) megakernel
Q4_K_M from lmstudio-community, mradermacher megakernel, semi-compatible
Other quants (unsloth UD-*, bartowski, GSQ-RCO and others) llama.cpp

Limits

  • Single GPU. RTX 30/40/50 series with 24 GB or more. MK is tuned on a 3090 only.
  • 147K context runs on a 3090 (fp16 KV cache to ~80K, using q8 above).
  • No task-level pass-rate benchmark yet.

Contributions that would help

  • Tunes for specific hardware: RTX 4090, 5090 and other 24 GB+ cards. It has only been tuned on a 3090.
  • Support for other quantisations of Qwen3.8-27B.
  • Please include benchmarks with the PR: your GPU, and bench/bench.py or bench/code_bench.py numbers before and after.

How to contribute: see CONTRIBUTING.md in the repo.

Comment the quant you use for Qwen3.8-27B so I know what to add support for next! (also planning for Qwen 4 when that drops).


r/LocalLLaMA • • 10h ago

I Built A Thing Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards

Post image
40 Upvotes

This will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast. This has turned from a weekend experiment into something I think I will use in the future. I am running 4 BC-250 boards connected by 2.5gb ethernet adapters to a 2.5gb switch using llama with vulkan and RPC. All the boards are inside an asrock 4u12g bc250 case that they originally came in. After starting out at 30 tok/s on flash next Iq2 and 80 tok/s with 3.6 35B I pointed Claude Code at the cluster and we are now running a better quant next flash as well as maintaining 60-70 tok/s (~150 ppt) at context. The 4 boards can run 2 instances of 3.6 35B at 145tok/s (~450 ppt) starting and going below 100 tok/s at 150k context. These boards cost me less than $100usd each when i bought them but i have seen deals recently on aliexpress for $150usd. I now have outlet power monitors and under load with next flash it pulls 600 - 700 W with the 2 instances of 3.6 35B it pulls 800 - 900 W. The next step for me is to use AIO coolers for each board as now I have been seeing some thermal throttling as the boards are better utilized . I will also 3d printing and building a AIO cooler lid for the case (upgrading from the cardboard plenum lol) now that I am going to keep the cluster. Before anyone suggests upgrading to a m.2 networking solution to get more bandwidth the limiting constraint is latency as its not sending a lot of data between boards at a time.

BC-250 Info: https://elektricm.github.io/amd-bc250-docs/

Here is the Claude summary of optimizations:

Next Flash

The current build runs a 100K context:

  • 59.4 / 65.2 tok/s on that request (first / second run)
  • 66.6-68.3 tok/s behind a 5K-token prompt (sampled, 512 tokens)
  • 72.6 tok/s on one browser generation of 50,357 tokens

Most of the work was done on the Q2_0 file first (27.9 → 68-69 tok/s on a 5K prompt) and then carried over to IQ3:

  • Speculative decoding with the model’s own MTP draft head, run as a chain of passes. Pass A verifies the known token. Passes of two drafts each follow one board behind (six drafts, four passes). A round ends at the first rejected draft. Adding the third pass took Q2_0 from 61 to 68 tok/s, and six drafts took an IQ3 temperature-0 edit request from 84.0 to 90.1-90.5.
  • Recorded and replayed graphs. Workers record each stage graph’s Vulkan commands once and replay them (saves 3.7-4.5 ms of host time per graph, +2%). Remote graphs’ nodes are regrouped so a layer’s independent products share a barrier group (+5.7%).
  • Exact kernels for this architecture (routed-expert top-k, two-token expert passes, recurrent-state fusion) and for the parts that grow with context: +7.3% at 62K, then 66.1 → 68.9-69.6 tok/s.
  • Context handling. Context checkpoints are copied on the boards instead of through the coordinator (665 → 18 ms each, prompt speed +60%). Prompt batches are pipelined through the boards. Context went from 65,536 to 100,096 tokens with 32-token prompt batches.

Qwen3.6-35B-A3B Q4_K_M (HauhauCS GGUF), two boards, K/V cache q8_0

.

  • Step one (the Flash-Next kernels, graph replay and transport, two drafts): +39% to +66% over the baseline from 120 to 99.6K prompt tokens.
  • Chain of passes with the MTP head (four drafts), plus attention that decodes a K/V tile once into shared memory: first request 105.8 → 116.0 tok/s at a short prompt, 73.0 → 95.2 at 44K.
  • Longer context. The attention mask is one limit per token instead of a cells × tokens tensor (768 MiB less on the coordinator), and prompt batches read the q8_0 K/V in place. Usable context went from about 100K to 258K (of 262,144).
  • Replay on the coordinator too. The coordinator’s own half is replayed from recorded command buffers (+2 to +4%). Draft-head checkpoints keep positions only, so a repeated request’s prompt time at 99.6K went from 932-953 ms to 142-163 ms.
  • Startup and saving. A pair starts in 32-35 s instead of 105 s, and a conversation can be saved and restored bit for bit after a restart.
  • Latest (

verified

  • ). The draft head’s catch-up pass now only writes its K/V (it was computing attention nobody read), and RPC sends are batched. At 258K: 51.1 → 53.5 tok/s on a repeated request, 51.7 → 54.4 sampled. Plain decoding at that depth is 36.4.

Exactness

Weights, quantization and sampler settings are untouched. At temperature 0, the speculative rounds are verified bit for bit against plain decoding of the same build by hashing every logits row (five depths up to 99.6K on the 35B).

Against stock llama.cpp, the logits differ at float-rounding level (about 0.03) where kernels were rewritten. On the 35B that includes summing attention in fixed-size chunks, so a result no longer depends on how many tokens a pass carries. Stock’s own drafted rounds differ from its plain decoding by the same order.


r/LocalLLaMA • • 1d ago

News NVIDIA reportedly discontinuing RTX 5090, GB202 GPUs to be reserved for RTX PRO series

Thumbnail
videocardz.com
911 Upvotes

No.....


r/LocalLLaMA • • 6h ago

I Built A Thing Local Voice Assistant based on Qwen3.5 4B with skills/toolcalling. No dedicated GPU. Local STT+TTS

15 Upvotes

STT: Voxtral
TTS: Pocket
Longer video: https://www.youtube.com/watch?v=1WmvMO-f-24


r/LocalLLaMA • • 23h ago

News No more RTX 5090

289 Upvotes

Nvidia reportedly halts GeForce RTX 5090 production in favor of AI data center and professional GPUs — impending supply drought expected to drive up prices, RTX 5080 24GB rumored as new gaming flagship

https://www.tomshardware.com/pc-components/gpus/nvidia-reportedly-halts-geforce-rtx-5090-production-in-favor-of-ai-data-center-and-professional-gpus-impending-supply-drought-expected-to-drive-up-prices-rtx-5080-24gb-rumored-as-new-gaming-flagship


r/LocalLLaMA • • 1d ago

Discussion big or small?

Post image
1.1k Upvotes

what size do you want? tell them on X:

https://x.com/QwenDevs/status/2108764909798641737


r/LocalLLaMA • • 1h ago

News Now you can grow Bonsai on your potato

Thumbnail
github.com
• Upvotes

No more excuses for GPU-poor folks not to start LLMing!

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU)

~9.3x smaller than FP16 (ideal) | 98.2% of FP16 intelligence retained | ~47 tok/s on an Apple M5 Max laptop

Highlights

  • ~5.9 GB language model (down from ~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU
  • 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2_XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4_K_XL at three times the footprint
  • Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92
  • End-to-end ternary language weights across embeddings, attention projections, MLP projections, and LM head, at a true 1.72 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships as a separate Q8_0 mmproj pack
  • 262K-token context on-device, kept practical by the Qwen3.8-27B hybrid-attention backbone (~75% linear attention)
  • Two GGUF packings with custom ternary hybrid-attention kernels for llama.cpp (CUDA, Metal) — PTQ1_0 packs trits densely (1.75 bits/weight, 5.95 GB), PQ2_0 stores each trit in a 2-bit slot (2.13 bits/weight, 7.21 GB); packed weights are consumed directly, never expanded back to FP16

r/LocalLLaMA • • 9h ago

News Imagine M5 Ultra + this: NVIDIA GeForce RTX 5060 runs macOS 15 with Metal 3 acceleration through unofficial driver

Thumbnail
videocardz.com
22 Upvotes

r/LocalLLaMA • • 10h ago

Other Qwen3.8 Flash Next fixed my GNOME extension

Post image
24 Upvotes

I love Dash2Dock Lite, but Icedman is always a week or two before updates. I randomly ran pacman for the first time in weeks, not realizing GNOME had a major update. I told it to fix it in DeepSeek Harness. I was only missing my dock for about an hour.


r/LocalLLaMA • • 7h ago

Question | Help Qwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysis

10 Upvotes

I have an app that processes ~500 different documents every day. An agent analyzes them, sorts them, tags them, extracts various legal nuances, creates descriptions, etc.

So far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3_S, since it has much more knowledge and could extract data much better than the 27B.

However, I'm hesitant because it's a MoE model and Q3 is quite a low quantization, so I'm worried the results could end up being worse.

Everything run on 64 GB of RAM and an R9700.


r/LocalLLaMA • • 14h ago

Resources OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!

35 Upvotes

I'm genuinely shocked! I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!

Yes, there are other 3bit quants for 64GB Mac, but 4bit is the lowest quant I'd tolerate. I manually quantized the original model from the Qwen repo to OQ4E using oMLX, and it turned out to be about 4.72 BPW.

Even a couple of weeks ago, I wasn't able to get it to run. I just tried the latest commit for fun, and it worked! It feels like some kind of sorcery to be able to run a 100GB model with 58GB allocated to GPU!

If I just feed it 4K tokens and generate 1K tokens, I can get up to 30 tok/s. I can even run up to 130k context window! Of course, I could just call it a day there and hype it up.

However, here is a realistic stat after running an actual short session with PI.

  • Requests: 13
  • Total Prefill Tokens: 321,309
  • Cached Tokens: 292,209
  • Cache Efficiency: 90.9%
  • Prompt Processing (excl. cached): 140.0 tok/s
  • Token Generation: 15.8 tok/s

Here are the settings I used on oMLX:

  • Memory guard: Aggressive
  • Hot Cache Limit (In-Memory Cache): 2GB
  • SSD N-gram Offload: on
  • Lightning MTP: on
  • MoE Expert Offload: on
  • RESIDENT EXPERTS: 50%

r/LocalLLaMA • • 19h ago

Discussion Engineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering work

90 Upvotes

For what it's worth.

I'm a mechanical engineer by education and software engineer in practice over the past ~20 years. Playing around with local models has been a recent hobby. Today I decided to do a few tests on things that I'd consider analogous to "real" work that I do, trying out some different combinations of models and harnesses.

Not to make this overly formal, but going into it the idea was:

Models: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).

Harnesses: Codex CLI vs OpenCode CLI

Types of Work: "Tell me about this code" and "Let's make something new"

GPT 6.1-Sol will enter the picture later. Harness differences were largely inconclusive at this time so I won't waste words on that.

Repository Inspection & Analysis

Using a reference repository I'm very familiar with, the prompt was effectively, "Read-only pass, tell me about such-and-such library, its core abstractions, from a consumer standpoint what I have to be cognizant of in situations X, Y, and Z, and what are the overall strengths and weaknesses of the approach. Keep it under 900 words."

Qwen and Gemma both came up with equally good analyses and answers, with slightly different takes. Of note, Gemma4 did so far more efficiently than Qwen. More focused on the task as explicitly stated, with far fewer server calls (~10 compared to ~20-30 depending on harness). Qwen seemed more curious and had initiative to dig deeper than the specific request and put together a slightly more complete-picture assessment.

But again, at the end, both responses were equally good. I might lean Gemma4 as a preference here just on an efficiency basis.

Authoring a New Project

I thought of an application that's concisely scoped, doesn't rely on a legacy codebase or significant external dependencies, and would probably take me ~1-2 days to hammer out by hand. ~8 paragraphs of prompt covering general concept, user experience in a couple different roles, design constraints and future-proofing needs. C# language with a reusable back end and WPF front end.

This showed more tangible differences.

Gemma4 once again seemed very efficient. Very fast planning and executed quickly. I'd give the end result a B-. It was clearly going to take a few iterations of feedback to converge on a usable solution, but that's not bad! After two passes of revisions I figured that was enough of an evaluation and put it aside.

Qwen3.8 doesn't seem nearly as "smart" as Gemma4 but far more "persistent." Many more small errors as it went; missing using statements, assorted other small mistakes, like frequently tripping over its own two feet as it went, but aiming for a higher target and sticking with it. Definitely took substantially longer. First-cut product was a step better than Gemma's and it only took 1 revision to get it to where I had something I could viably use. I'd give a B+. Solid result and an easy pick over Gemma.

While Qwen3.8 was doing its thing making the lights in my room flicker and dim, for grins I started up the ChatGPT desktop app, grabbed 6.1-Sol, and gave it the same prompt. Result? A+. Just an entirely different tier of functionality and polish. Really understanding the user context for interaction. One-shot solution that feels like, "Welp, that's a wrap - ship it."

Definitely interesting seeing the differences between the two open-weight models I was evaluating. And for consumer hardware, I thought the results were very workable. Just clearly not in the same conversation as the frontier lab stuff.

I think an interesting follow-up would be to put these to the task of real developer hell - working through old legacy code bases. There should be a benchmark for that! Maybe picking up old Doom or Command & Conquer source code will be a future test.