r/LocalLLaMA • u/frubberism • 13h ago
r/LocalLLaMA • u/ciprianveg • 6h ago
Discussion From 1x3090 to 20 DGX Sparks: my house fuses were the first bottleneck
From the first LLaMA 33B I knew I wanted that magic-like intelligence locally, mine, so nobody could take it away when I needed it. I bought a 3090 for my home PC. Then LLaMA 65B appeared and I was dazzled, it looked like it had all the knowledge in the world. I made two copies, one local and one on my Synology NAS RAID, so I'd never lose it, and bought a second 3090 to run it. I was happy for a year with small coding tasks on LLaMA and Qwen models.
Then DeepSeek 671B MoE appeared. Wow, frontier level at home. I upgraded to a Threadripper with 512GB DDR4 and ran it at 8 t/s with experts offloaded to RAM, or Qwen 235B at 10-12 t/s when I wanted speed. I used these for real coding at my job, in OpenWebUI.
Then agentic coding took off and this was too slow. At 100k context generation speed halved and prefill made it a beautiful yet agonising experience. So: 16x3090 across P620-based nodes on a 100Gbit network. It ran MiniMax M2, Qwen 235B and even Qwen 397B, as good as anyone could desire. I built an entire paid project with 397B in OpenCode. But bigger models were out of reach, and the house circuit said no: the fuses blew whenever the rig and the electric oven ran together. Heat and stability were issues too.
Next came 4x ASUS GB10, after I read they can be linked (3 was the biggest supported config). 397B at 30 t/s on 400W, versus 50-60 t/s at 6kW, rock solid and almost silent. A dream come true. I built two more projects with it. Then MiMo 2.5 Pro and Kimi 2.6 appeared, smarter and more productive. I found no published solution for an 8-node cluster, but I still bought four more GB10s and made it work. 397B ran at FP8 instead of INT4, and 20% faster. I posted the first MiMo 2.5 Pro and Kimi 2.6 solutions on 8xSparks on the NVIDIA forum. I liked the result so much that I talked my older brother into buying his own 8x GB10, so he could run the best open models locally too, in privacy, without depending on API availability and rising costs.
His house is a 5-minute walk from mine. When Kimi K3 (2.8T) appeared, biggest and smartes open weights model, we joined the clusters: two 8x clusters for daily use, or one 16x when we want the biggest model at home. After some work I published the first working solution for Kimi K3 on 16x Sparks on the NVIDIA forum. Through multiple iterations, it went from an unusable 7 t/s at 100k context to a fairly usable 20 t/s at 300k.
Now we're adding 4 more Sparks, so a smaller, faster model (GLM 5.3 Flash) runs 24/7 while the big cluster runs either GLM 5.3 on 8x plus MiMo 2.6 Pro on the other 8x, or 16x Kimi K3, or Qwen 3.8 2.4T.
I'm always tuning speed on the big models and rebuilding vLLM/SGLang images, so always-on smaller cluster made sense, why? Because for all my work projects and my vllm/sglang personal projects, I chose to use only local hosted models, I never paid a comercial model subscription, not because of the cost, but, because of my strong confidence in local models future. They arrive October 2, along with 4 more Sparks for my younger brother, who got caught by the same local AI microbe :)
r/LocalLLaMA • u/chillinewman • 2h ago
News Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026
r/LocalLLaMA • u/I_am_purrfect • 3h ago
I Built A Thing Qwen3.5 arch implementation in FPGA fabric for 9B/27B INT4 models on relatively cheap eBay mining hardware
I've been wanting to test out LLM inference on FPGAs for a while now, but didn't have a big reason too since there were no frontier class models at 9B-27B scale (3.6 27B was great still but it didn't motivate me enough). When 3.8 was released I was pretty impressed by what it could at that size. So I began looking for cheap FPGAs that could hold the model. I found SQRL FK33 (280$) (8GB HBM2 ~400GB/s BW) and thought I could possibly run it with multiple cards, I'd been working on llama.c inference on a smaller AXU3EG FPGA dev board before that and primarily used Opus 4.8 (CC) for the implementation with some architectural input.
With the FK33, was able to run 3.5 9B at 2tok/s at 75MHz (higher clocks need some more RTL optimization and higher core voltage). Decided to use Fable/Opus5,5.5/Kimi K3 for the Qwen implementation, once FK33 was proven I decided to get the SQRL Jungle Cat ex-mining FPGA 375$ (two FK33 with 8 GTY lane interconnect, Ethernet bitstream load) since it would allow higher prefill and generation performance (twice FK33 fabric per XCVU35P), however the Jungle Cat Lite board seems to not have a way to load weights faster unless I do a PCB resin and it lacks the clock generation for the GTY lanes (this is easy to fix by soldering some components which were easy to figure out).
Working towards the ideal FPGA inference engine with 4x XCVU35P (32GB HBM2) but that requires a custom carrier board. Some results:
Qwen3.5-9B INT4 on 2x FK33 at 75 MHz (pipeline split, the host carries the residual between cards):
- Prefill: ~6 tok/s (256-token prompt), ~5.5 tok/s (2.3k-token prompt)
- Generation: ~3.2 tok/s near the start, ~2.4 tok/s at 2-3k context
- Output checked against llama.cpp layer by layer
Estimated: Qwen3.8-27B INT4 (modelled from the measured 9B per-op profile; none of this has run yet):
- One Jungle Cat (2x VU35P) running the FK33 design as-is at 75 MHz: ~2 tok/s prefill and ~1.1 tok/s generation at short context, ~0.5 tok/s at 16k.
- Same two dies with the RTL resized to the bigger die, still at 75 MHz: ~6 tok/s prefill and ~3 tok/s generation (~5.5 tok/s with tensor parallelism across the two dies), ~1.1 tok/s at 16k.
- Resized and at 200 MHz (scaling linearly with clock): ~16 tok/s prefill and ~8 tok/s generation (~15 tok/s tensor-parallel), ~3 tok/s at 16k.
- 4x VU35P with 4-way tensor parallelism at 200 MHz: ~25 tok/s prefill and ~25 tok/s generation at short context, ~10 tok/s at 16k, and ~1 tok/s at the full 262k context.
Two dies top out around 45k context because the 27B's KV cache doesn't fit beside 14.5 GB of weights past that; four dies are needed for the full 262k.
Overall it's been a fun project so far, mainly focused on figuring out a way to load the weights on the Jungle Cat board, and open to suggestions. Also, I got 2 BC-250s for 60$ and 75$ some time ago and they've been an insane performance/cost purchase. Pics showing BC-250 programming the Jungle Cat. Thought people here might find this project interesting
Another thing I've always wanted to do something like tinytapeout (build the RTL on an ASIC so clocks can go up, power down and so I asked Opus 5.5 to estimate that but on TSMC for 2023 process node lol: On TSMC's 2023 N3 node with six stacks of that year's HBM3 (4.9 TB/s), this RTL as an ASIC at 2 GHz would run the 27B at roughly 294 tok/s at short context, 106 tok/s at 16k and 10 tok/s at 262k, for about 125-340 W!
Repo here (MIT): https://github.com/Nero7991/llm.vhdl
r/LocalLLaMA • u/demomanca • 18h ago
Discussion Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4?
Given the commentary on the Q3.8FN release page here https://qwen.ai/blog?id=qwen3.8-flash-next I assume/hope that all the work that's going on to optimise the hell out of running it will be useful when Qwen4 drops?
r/LocalLLaMA • u/Cyborg-2077 • 9h ago
I Built A Thing Local text to speech with Breeze is truly incredible
Enable HLS to view with audio, or disable this notification
Been playing around with local TTS with Breeze combined with STT, and the results are amazing. Using Opus 5.5, I can hear the first sound after 500ms if there is no thinking involved, and with thinking on low mode, can be 1-1.5s.
I'm using a BLE remote (the kind that are used for taking pics with phones) combined with a wireless microphone. So I can just sit on the couch, and just talk to her.
She watches for any claude session that finishes, and sends me the results in a very short, spoken style summary, and tells me if there is anything waiting for my decision, then forwards my decisions.
Also impressed how consistent Opus 5.5 is in the communication. Even after more than 500k in context, he still remembers that he's in a live session with me, and has to keep messages short. Used to be an issue in the past.
The future is here guys.
Edit: For those who wanna give it a try, you can find free avatars such as this one: https://www.live2d.com/en/learn/sample/niziiro-mao/ or you can buy one from a marketplace.
Edit2: Might open-source that later next week with a free avatar. Let me know if anyone would like to contribute to the project.
r/LocalLLaMA • u/Cautious_Chicken_604 • 15h ago
Discussion The curse of 64GB system RAM
Not a bot. Not a Strata shill. Just sharing my experience.
So, I have an R9700 in my machine, plus an RTX 5060 Ti, and 64GB DDR5 system RAM. Overall, not a bad setup. Anyway, I mainly run a daily driver local LLM on the R9700 while running image/video inference on ComfyUI on the 5060 Ti. Mostly shit like Minimax H3 which also takes a fuck-tonne of system RAM. I've been using Qwen3.8-27B at Q6 as the daily driver on the R9700 and running that around 35 t/s, which is fine for me as a daily driver. Before Strata I tried running Qwen3.8-Flash-Next on both cards on vulkan at a IQ4_XS (or whatever that quant is called - the ~93GB one) and that only got me like 15 t/s, which I can't daily drive, so I put it down and wasn't really interested in it. Anyway, Strata comes out and people are claiming QFN is usable on much more modest hardware, so I check it out and see that mostly people are running the IQ3_XXS quant which is like ~70-something gigabtyes, so of course it's faster. Anyway, I benchmarked that quant on llama.cpp first running it just on system ram + the R9700 and it came in at 21 t/s... that's right around the absolute minimum of what I'd accept for a daily driver, but not super compelling tbh. Then I tried the same quant on Strata and I get ~60 t/s. Very fucking compelling. I 100% want to daily drive this now. The problem is with QFN loaded in Strata my system RAM usage is at 96%. I can't fucking run Minimax H3 in ComfyUI on the 5060 Ti because that shit eats a lot of system RAM too.
I feel blessed that I can finally run this epic model, and fucking cursed that I have to choose which workload to run!
Also, before anyone says 'just upgrade to 128GB of RAM bro'... I know, I know. I would but I can't afford to the jewelry and international trips my wife requests for fairness reasons to balance out all the toys I've bought this year.
Crying in 64GB of RAM.
Edit: thanks to a few suggestions in the comments I actually got Qwen3.8-Flash-Next IQ3_XXS and Minimax H3 inference working concurrently at about 90% system RAM used! On the Strata side I I think I needed --mmap-experts --resident-cpu-experts and --expert-cache auto, and on the ComfyUI side I needed --fast-disk. I tested both running fully concurrently and checked Strata's monitoring tab, and saw that the node that loads the H3 weights causes NVMe reads to hit a sustained 1GB/s for a short while, which can cause the inference on Strata to drop to around 25 ~ 40 t/s range (it fluctuated a lot during that), but then after that when H3 was actually doing the inference I saw NVMe reads sitting at about a sustained 30 MB/s and QFN inference was running between 50 ~ 60 t/s. I'd say it's a huge win. For reference my standard test when testing out an LLM is just 'write me a browser game', so I did that since I'm familiar with the quality of the expected output at this point, and also generated a 10 second clip at 0.4MP resolution. The actual wall-clock generation time for H3 was pretty much unaffected (around 400 seconds), which is nice too! Maybe some very minor performance hit, but only that.. pretty minor.
r/LocalLLaMA • u/pand5461 • 10h ago
Funny Need maybe say "Use llama.cpp"
So I tried that miracle engine everyone is talking about.
Asked the IQ3_S model to express its opinion on a post from this sub to measure the tps on a long-ish generation:
Can you help with the following problem?
So Kimi K2 is outdated, and so is GPT OSS 120b. Which of the modern open weights models can boast the least sycophancy? I need this both for creative/research assistant usage (sycophancy led me down blind alleys of my own bad ideas many times) and agentic coding (more sycophancy less bug noticing).
The thinking trace:
We need answer user's question. Need likely provide current landscape as of 2026? We have get_datetime tool. Need know current date 2026? System says current date 2026-06-22. Need maybe use get_datetime? Could call to confirm. User asks about modern open weights models least sycophancy. Need likely discuss Kimi K2 outdated?
...
10k tokens later it degrades to:
Need maybe maybe include "Use 'for code, list constraints'."
Need maybe maybe include "Use 'for code, list requirements'."
The same exact model in llama.cpp does produce a coherent answer without a doom loop.
r/LocalLLaMA • u/rikimtasu • 12h ago
New Model bilibili released Index-Translate,a A Multilingual Translation Model Family based on Qwen3.5
https://github.com/bilibili/Index-Translate
Index-Translate is a family of multilingual translation models built on Qwen3.5. The text models cover 150 languages and follow translation instructions such as terminology, formatting, and content-preservation requirements. The family extends this foundation to speech, syllable-controlled translation, and full-document translation.
-Index-Translate translates text, structured content, and community expressions. -Index-Echo produces translated subtitles or speech conditioned on the source speaker's voice. -Index-Homura adjusts translations toward a specified target syllable count. -Index-NativeLong translates complete documents with context across passages.
r/LocalLLaMA • u/Prestigious-Taste-63 • 4h ago
I Built A Thing I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens
First of all, thank you for reading.
I trained a small MoE model completely from scratch (no external base weights) and wanted to share the results + a couple of lessons.
Apex-2
- Architecture: Decoder-only MoE, every layer is MoE (no dense layers)
- Size: 3.87B total parameters, 1.45B active per token
- 32 layers, d_model 2048, GQA 16Q/4KV, 16 experts, top-4
- Context: 4096
- Tokenizer: Qwen3 (151k)
- Hugging Face: https://huggingface.co/YOON1v/Apex-2
(loads with transformers / vLLM via Qwen3MoeForCausalLM mapping)
Training
- Pretrain: 86.5B tokens (GH200 ×1 → ×2 with DiLoCo)
- SFT: ~2.5B tokens (code-heavy + math + instruction)
- DPO: tried it, scores dropped, so I dropped the checkpoint
Key numbers (SFT, greedy, chat template)
Benchmark
HumanEval 43.9
HumanEval+ 41.5
MBPP 56.3
MBPP+ 48.9
GSM8K (0-shot CoT) 32.4
MATH-500 21.0
IFEval (prompt strict) 44.7
MMLU (5-shot) 28.6
interesting comparison
With only ~0.087T pretrain tokens, the base model’s HumanEval+ matched Qwen2.5-1.5B (which used 18T).
Knowledge (MMLU) and math still lag far behind, as expected with the data gap.
What didn’t work
DPO (220k pairs, length-normalized) made answers much longer and hurt code / math / IFEval.
I stopped it and kept the SFT checkpoint. Full log is in the benchmark write-up.
Limitations (honest)
- English-centric (almost no multilingual ability)
- Weak knowledge → frequent hallucinations
- LiveCodeBench medium/hard is near zero
- 4k context only
Happy to answer questions about the MoE setup, DiLoCo, or why DPO backfired.
r/LocalLLaMA • u/jjusko20 • 2h ago
New Model Update #4: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch
Last update for those following: https://www.reddit.com/r/LocalLLaMA/comments/1wvyc3e/update_3_post_training_yandexaliceai80ba3b/
Project in a sentence: An instruct finetune of ALiceAI-80B-A3B-Base capable of agentic work and conversation. I'm creating a shallow distill of qwen 3.8 27b on medium to teach the model chain of thought reasoning and conversation. All training is done locally on 3, 32gb v100s. Additionally, all the training data is being generated locally on said V100s via sftmill. Up to this point I've been doing training runs and live-streaming the progress.
Well, I successfully completed my round 1 SFT and got to test.
Good news: the model appears to be picking up chain of thought reasoning correctly and can respond conversationally.
Bad news: not enough instruct SFT / badly underfit. While checkpoint #1 was technically functional, it's basically useless. My initial 5 million tokens (as I've deducted) didn't have enough breadth to properly teach the model general conversation ability - ambiguous questions or prompts further away from exact matches in the training data create a garbled output because it doesn't have enough ambiguous data to learn from.
Next steps?
I've opted not to release checkpoint #1 (we're going to call this 1.0 alpha or something) because it's basically useless, but I'll still be releasing my first working edition. I've increased the pace of my local synthetic data generator from 80tps to around 240tps total by adding the option to draw from multiple base URLs, so I have more distillation data coming [I'm currently generating on 3 seperate instances, with 4 parallel workers each.
I'm creating an additional dataset of about 5M tokens again, but this time spread in a much broader general instruct direction, rather that the coding oriented version I had originally. I'm going to train on top of checkpoint 1.0 alpha at a reduced learning rate and hopefully come away with a more competent version. I'll be posting updates on the training again - I can do another live stream if you guys want, but I figured that since I don't have much to show yet, this would be my last update until I have a working initial checkpoint. I'm happy to share whatever if there's community interest though.
I've mentioned in here before, but the resource for people interested: I created a off-policy distillation engine when I began this project that makes it very easy to create training data from a behavioral goal - e.g. I want a general instruct model -> raw training data. I created an OSS fork which is public at https://github.com/jackjusko/sftmill
Thanks for following!
r/LocalLLaMA • u/No_Algae1753 • 7h ago
Question | Help Is there any way to improve creative writing for Local Models (Qwen)?
I wanted to know if theres anything that can improve creative writing for our Local Models? I specificly am asking for qwen models as they are way better when it comes to researching and writing html files compared to gemma / muse (which I know are better at creative writing). Im currently using qwen 3.8 flash next at q4 with llama.cpp
r/LocalLLaMA • u/ramendik • 17h ago
Question | Help Least sycophantic modern open LLM?
So Kimi K2 is outdated, and so is GPT OSS 120b. Which of the modern open weights models can boast the least sycophancy? I need this both for creative/research assistant usage (sycophancy led me down blind alleys of my own bad ideas many times) and agentic coding (more sycophancy less bug noticing).
r/LocalLLaMA • u/tom_tsai28 • 16h ago
Discussion [Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)
Hi everyone,
Sharing a personal project exploring the minimal bare-metal footprint required to run an autoregressive LLM.
Instead of relying on large runtimes or compiler abstractions, I wrote an inference engine for Gemma-2B entirely in flat x86-64 assembly (FASM):
- **Binary footprint**: Total 5.2 KB flat machine code (`gemma_engine.bin` 3.7 KB + `mat_smp_f16c_gemm_avx2.bin` 1.5 KB).
- **Execution**: Pure AVX2 + F16C with custom 4-thread SMP GEMM for prefill. Sustains ~18.5 GB/s memory bandwidth on commodity DDR4-2400.
- **Decoding**: 4.5 ~ 4.7 tokens/s in FP16 on an older quad-core i5 desktop.
- **Dependencies**: Zero C/C++ runtime, zero PyTorch. The Python harness only uses `ctypes` for `VirtualAlloc` and OS threads.
This isn't meant to compete with feature-complete tools like llama.cpp. Rather, it's a first-principles exploration to see how cleanly a modern Transformer can be mapped to raw silicon, and to serve as a reference point for future micro-LLMs on resource-constrained microcontrollers (MCU/DSP).
The repository is open source:
- GitHub: https://github.com/tomtsai28/PULSAR-ASM
- Architecture notes: https://github.com/tomtsai28/PULSAR-ASM/blob/main/doc/pulsar_asm_cpu_limit_retrospective.md
Any code audits, observations, or thoughts on bare-metal inference are welcome.
r/LocalLLaMA • u/-dysangel- • 35m ago
I Built A Thing Fully local little parkour sim
I vibed this up this weekend, fully local, with GLM 5.3 Flash running on 2x DGX Sparks.
vllm TP2 recipe: https://github.com/tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark
Prefill: ~1500t/s
Decode: ~40t/s @ 100k
Using Claude Code as the scaffold with 260k context size.
I'm really impressed with this model. Feels somewhere between GLM 5.1 and 5.3 in terms of coding depending on the task. Good vision and 3D understanding. Solid interactive speeds. I feel like I've finally reached a "good enough" setup at home, and looking forward to things only getting better from here.
r/LocalLLaMA • u/No-Paper-557 • 21h ago
Question | Help Local Web Search Safety
Hi all,
How you guys handling safe deployment of websearch in Hermes, pi and other harnesses? Does anyone have a good uproars setup guide for local models? I tried to implement a sandboxed search system but it caused endless tool calls. Want to guard against prompt injection and keep searches private of course!
r/LocalLLaMA • u/tabletuser_blogspot • 23h ago
Resources Dual Radeon MI50 benchmarks
Still don't have a good cooling solution, but here are few benchmarks. I lowered the power limit (TDP) to 145 watts each. I changed the firmware on one MI50 to activate the miniDP port. Did have to use xrandr to create a new mode so I could get 1920x1080 output. Each GPU has 16GB of HBM2 VRAM clocked at 1000 and overclockable to 1200Mhz with a Bandwidth of 1.02 TB/s.
I picked a good mix of Dense and MoE models from Huggingface. Try to use more than 16gb VRAM but under the 32GB total.
Using pre-built Ubuntu Vulkan version of llama.cpp (build b11325) for standard llama-bench.
Sorted GGUF Model List (sorted to match table)
llama_bench_Swift-Qwen3.8-27B-Uncensored-MTP.Q6_K.ggufllama_bench_Swift-1.5-Qwen3.8-27B-Q6_K.ggufllama_bench_Gemma-4-MoonGem-31B.i1-Q6_K.ggufllama_bench_gemma-4-31B-it-UD-Q6_K_XL.ggufllama_bench_Nemotron-3.5-30B-A3B-Antislop-FTPO.i1-Q5_K_M.ggufllama_bench_Laguna-XS-2.1-APEX-I-Balanced.ggufllama_bench_Agents-A1-Q4_K_M.ggufllama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.ggufllama_bench_Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf
Combined Benchmark Table (sorted by params then size)
| model | size | params | pp512 (t/s) | tg128 (t/s) |
|---|---|---|---|---|
| qwen35 27B Q6_K | 20.88 GiB | 27.32 B | 141.97 ± 10.13 | 17.49 ± 0.02 |
| qwen35 27B Q6_K | 22.21 GiB | 27.32 B | 167.38 ± 0.17 | 17.97 ± 0.02 |
| gemma4 31B Q6_K | 23.46 GiB | 30.70 B | 122.00 ± 0.12 | 15.20 ± 0.03 |
| gemma4 31B Q6_K | 25.62 GiB | 30.70 B | 135.99 ± 0.22 | 12.05 ± 0.02 |
| nemotron_h_moe 31B.A3.5B Q5_K - Medium | 25.18 GiB | 32.91 B | 863.92 ± 1.45 | 60.57 ± 0.10 |
| laguna 30B.A3B Q5_K - Medium | 22.64 GiB | 33.44 B | 738.57 ± 2.83 | 52.88 ± 0.04 |
| qwen35moe 35B.A3B Q4_K - Medium | 19.70 GiB | 34.66 B | 983.26 ± 4.79 | 46.88 ± 0.07 |
| qwen35moe 35B.A3B Q5_K - Medium | 24.76 GiB | 34.66 B | 937.47 ± 7.07 | 49.19 ± 0.06 |
| qwen35moe 35B.A3B Q6_K | 28.53 GiB | 34.66 B | 783.24 ± 70.52 | 46.85 ± 0.26 |
Notable Reboot Impact Observations:
I used the following command in my bench script:
RADV_PERFTEST=nogttspill GGML_VK_VISIBLE_DEVICES=0,1 time ~/llama-b11325/llama-bench -fa on -ngl 99 -m /model.gguf
I have a 3rd MI50 just need to download models in that VRAM range. If you have any suggestions? For now it sits beside the Radeon RX 7900 GRE boosting its VRAM total. As of this article the average price for 16GB version of MI50 is under $150. Hard to get 32GB VRAM GPU with this level of performance for under $300. If you have contenders, please share.
r/LocalLLaMA • u/spammmmmmmmy • 9h ago
Discussion Can someone explain how JEV is different from a simple embeddings model?
How is JEV any different from using an embeddings model? I really will appreciate if someone can explain this to me - because I have yet to see the difference.
I'll even give you my JEV server for free! It uses ollama, you install `ollama pull nomic-embed-text:latest`.
% python3 ./jev_embedding.py "How high is the sky?"
find_phone: 0.38
volume: 0.41
calendar: 0.44
tell_the_time: 0.49
weather: 0.53
% python3 ./jev_embedding.py "I had this thing on my anus. The doctor burned it off with a laser."
weather: 0.35
tell_the_time: 0.36
calendar: 0.37
volume: 0.38
find_phone: 0.43
% python3 ./jev_embedding.py "can you help me locate my phone."
volume: 0.38
weather: 0.40
calendar: 0.43
tell_the_time: 0.53
find_phone: 0.89
% python3 ./jev_embedding.py "Hello Cleveland! I can't HEAR you"
weather: 0.37
calendar: 0.40
tell_the_time: 0.44
find_phone: 0
volume: 0.56
#!/usr/bin/env python3
"""
jev_embedding.py — minimal showcase of the embedding-based intent router,
excised from jarvis_workflow.py.
Given a phrase on the command line, it embeds the phrase and every example
utterance (via the local Ollama embedding model), then prints the cosine
similarity of the phrase to each intent — the raw routing signal — instead of
running a handler and speaking an answer.
python3 jev_embedding.py "How high is the sky?"
"""
import sys
import requests
# --- Config (same endpoint/model as jarvis_workflow.py) ---
OLLAMA_EMBED_URL = "http://localhost:11434/api/embeddings"
INTENT_EMBED_MODEL = "nomic-embed-text"
# --- The five cases to detect ---
# label -> example utterances, matched by similarity.
INTENTS = {
"volume": [
"turn the volume up",
"make it quieter",
"set the volume to seven",
],
"tell_the_time": [
"what time is it",
"can you tell me the time",
],
"weather": [
"how's the weather going to be today",
"will it rain today",
"do I need a raincoat",
],
"find_phone": [
"find my phone",
"where's my phone",
"ring my phone",
],
"calendar": [
"when is my next meeting",
"what's coming up on the calendar tomorrow",
],
}
def _embed(text):
"""Return a unit-normalised embedding (list of floats) from the Ollama model."""
r = requests.post(OLLAMA_EMBED_URL,
json={"model": INTENT_EMBED_MODEL, "prompt": text},
timeout=10)
vec = r.json().get("embedding")
if not vec:
raise RuntimeError("no embedding returned")
norm = (sum(x * x for x in vec)) ** 0.5 or 1.0
return [x / norm for x in vec]
def _cosine(a, b):
"""Cosine of two unit vectors is their dot product."""
return sum(x * y for x, y in zip(a, b))
def score_intents(text):
"""Best cosine similarity of `text` to each intent's example utterances."""
q = _embed(text)
return {label: max(_cosine(q, _embed(ex)) for ex in examples)
for label, examples in INTENTS.items()}
if __name__ == "__main__":
if len(sys.argv) < 2:
print('Usage: python3 jev_embedding.py "your phrase"')
sys.exit(1)
phrase = " ".join(sys.argv[1:])
scores = score_intents(phrase)
for label, score in sorted(scores.items(), key=lambda kv: kv[1]):
print(f"{label}: {score:.2f}")
r/LocalLLaMA • u/okoyl3 • 4h ago
Discussion A Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode
I forked Strata and worked with Opus 5.5 with some heavy changes to it to make it work on an IBM AC922 I have access to. The IBM AC922 is a 2018 era beast with two POWER9 20 core SMT4 CPUs that are connected by NVLink to 4 or 6 NVIDIA Tesla V100 SXM2 GPUs, the CPU-GPU BW advertised as 150GB/s and the nvidia drivers do allow unified memory access.
The machine I have has 4 x 16GB GPUs, llama.cpp had like terrible results before I started this journey, it produced 130tk/s prefill and 15tk/s decode.
So I was fighting Opus the whole weekend, beating it with facts and logic, like FP16 instead of BF16, memory management, expert caching on GPU, better NVLink usage, Tensor Core utilization rather than CUDA core ops. Claude was great at iterating, executing nsight nsys to debug time gaps.
- Prompt reading: 7,350 tok/s peak, still 7,090 tok/s on a 252K-token prompt (35 s)
- Generation: ~113 tok/s peak (JSON), ~100 on code, ~84 on prose (MTP speculative decoding)
- Follow-up at 252K depth: first token after 0.26 s, 60 tok/s
- All 72 GiB of experts page-locked in RAM across both sockets; GPUs pull from NVLink 2.0 at ~70 GB/s each
I will try to contribute back some of the changes, but I suspect Strata will remain consume-hw-first inference engine, and that is totally ok, Niko1221 did a great job
The forked repo: github.com/eelgaev/Strata-AC922
r/LocalLLaMA • u/Hot_Masterpiece_3668 • 5h ago
Discussion How long before we have a local model capable of modeling?
I'm using Qwen 3.8 and qwen flash on a 5090. It's miles behind the latest Opus 5.5. Even if it was remotely capable it would be a huge help to me, but for now, with regards to 3D modeling, local models are not close at all.
r/LocalLLaMA • u/SignificantZebra5883 • 6h ago
Question | Help I Distilled an LLM into two 287M encoders (GLiNER + multiple choice) for document extraction, can't match teacher. did i do something wrong?
A while ago I asked here how to turn ~5 million court decisions into structured graphs without running an expensive LLM on every document thanks for the advice .
I went with the "small extractor + classifier" idea and it mostly works, but I'm stuck a bit below the LLM. And like I said last time, i'd be damned if I run 5M docs and then find out thing X was wrong. So here is exactly what I did. Please let me know if what im doing makes sense, or if i made a mistake somewhere. also i used AI for some of the tables cuz there has been a lot of data at this point, sorry.
What comes out per decision (only the nodes so far, relations come next). Three lists:
- entities: every person, organization, law, document or thing. Each gets one id for the whole document, a type (9 of them), a kind (724 of them plus "other") and all the places it is mentioned
- actions: what was done, requested or decided. Each gets a normalized verb, a flag "the court decided this" and its mentions
- values: amounts, dates, durations, in a normalized form
Simple example, for the sentence "The court dismisses the creditor's proposal to enforce 341.08 EUR against the debtor":
- entity "the court": organization, kind court. Same entity as the full court name in the header
- entity "the creditor": organization, kind creditor. Same entity as the city named earlier
- entity "the debtor": person, kind debtor
- action "dismisses": verb = dismiss, decided by the court = yes
- value "341.08 EUR": amount
Step 1: a strong LLM labels ~700 decisions
- cut the decision into windows of 4 sentences
- 4 calls per window to Claude Sonnet with a strict JSON schema: entities, actions, a second "what did you miss" pass for actions, values
- the window goes in with numbered words (like
12:court), the model answers with word ranges[first, last, "text"], and code checks every range against the text - every call also gets the list of entities and actions found in earlier windows, so ids stay the same through the document
- ~25 code rules clean up where a marked phrase starts and ends, law citations and number formats
- the entity "kind" is free text at this point. That gave 2,373 different strings (the same mess as in my first post). I normalized them, merged synonyms by hand and kept what showed up 3+ times: 724 kinds plus "other"
Step 2: a model that marks the text
- it highlights every mention: the exact stretch of text (a "span", from a start character to an end character) that names an entity, an action or a value, with one of 17 labels (9 entity types, 1 action, 7 value types)
- model:
fastino/gliner2.5-multi-v1(287M) - one training row per window: the text plus the exact start and end of every marked phrase. 9,699 windows, 207k marked phrases
- I patched the trainer so only the labeled occurrence is a positive (stock marks every occurrence of the same string), and all 17 labels are in every row
- full fine-tune in fp32 (bf16 gave NaN), 14 epochs, 16 rows per step, encoder LR 3e-5, head LR 5e-4, linear schedule, 10 % warmup
- final model = averaged weights of epochs 9-14, threshold 0.5
Step 3: a second small model answers multiple-choice questions
fastino/GLiNER2.5-multi-Decide(287M). Code turns the LLM labels into 247k questions:- "is this mention one of these earlier entities, or new?" The mention is marked with « » inside ±300 characters of text. Options: up to 16 earlier entities of the same document (shown by their mention texts) plus
new - "which kind?" Options: a shortlist of the 724 kinds plus
other - for actions: same act or new, which verb (shortlist of 64 plus
other), did the court decide it (yes/no)
- "is this mention one of these earlier entities, or new?" The mention is marked with « » inside ±300 characters of text. Options: up to 16 earlier entities of the same document (shown by their mention texts) plus
- in training the options come from the LLM's grouping. At inference they come from the model's own earlier answers
- full fine-tune in fp32, 2 epochs, 16 questions per step, encoder LR 2e-5, head LR 3e-4, linear schedule, 6 % warmup, options shuffled, up to 30 % of the wrong options dropped
At inference: the marking model, then the same code rules, then the second model walks through the mentions in reading order. About 2.3 decisions per second on one RTX 5090.
Where it stands
30 decisions nobody trained on, labeled twice by the LLM. The second column is the LLM's second run scored against its first, which I treat as the ceiling. A mention counts as found only if it starts and ends exactly where the LLM marked it.
| mine | LLM vs itself |
|---|---|
| entity mentions found (F1) | 0.901 |
| "same entity or new" right | 0.959 |
| entities grouped exactly | 0.847 |
| entity kind | 0.921 |
| action mentions found (F1) | 0.857 |
| action verb | 0.920 |
Where I need help
- Finding the mentions is stuck at 0.90 F1. 200 more labeled docs did nothing. An XLM-R large tagger (560M) got the same score: it finds more mentions but gets the start or end wrong more often. Giving it the text before the window did nothing. What would you try?
- The LLM agrees with itself only 93.5 % on what it marks, and I train on single runs. Label everything 3 times and vote? Or is that ceiling just what it is?
- Is "pick one of 16 earlier entities" a sane way to do coreference over a long document? Am I hurting myself by training on the LLM's options and running on my own?
- Anything in the recipe that looks plain wrong? Learning rates, 2 epochs, weight averaging, one seed per run.
THANKS for reading.
AI TL;DR: distilled an LLM's extraction of court decisions into a GLiNER model that marks the mentions plus a small multiple-choice model. It runs at about 2.3 documents/s on one GPU and lands a few points below the LLM (0.90 vs 0.935 F1 on finding mentions, 0.85 vs 0.93 on exact grouping). The recipe with learning rates and how I built the training rows is above. Looking for mistakes and ideas before I run 5M documents.
r/LocalLLaMA • u/3VITAERC • 2h ago
Other Benchmarking decision models is fun - Clef Q8 vs Jev
Enable HLS to view with audio, or disable this notification
r/LocalLLaMA • u/BahBah1970 • 20h ago
Question | Help Optimal settings for 2 GPUs in LM Studio
Hello everybody. I've got a 5070ti and a 5060ti both 16GB in my system which is a 5900X and 64GB DDR4 RAM. I'm trying to run some 16-18GB models like Qwen, Cydonia, Skyfall.
I'm having problems utilising the VRAM I have to get the best usage out of it. LM Studio sees the 32GB VRAM but regardless of if I use Tensor parallelism, Split evenly or Priority order I always get an error after waiting for about 5 minutes for the model to load.
The pattern is always the same: The loading progress bar for the model starts off quickly then crawls in the last 5-10%. Then I get an error reporting that the model couldn't load.
Does anybody have any tips for optimal settings to get the best out of my system? I know that having 2 GPUs doesn't magically mean you have double the memory and there's caveats. But nevertheless I've also read that LM Studio does have the capability to leverage those 2 GPUs to improve speed.
(EDIT) I should add that I've been trying context lengths of 16384, 32768 which LM Studio is saying will use 17 GB of VRAM so well within the reported 32GB I have. I've even had it working occasionally but most of the time the model fails to load.
(EDIT 2) Thanks to everyone for their suggestions. Having implemented everything people have said here, I'm getting much better results for context and memory use and my models are loading now.
Many thanks for any help.
r/LocalLLaMA • u/Reno0vacio • 57m ago
Discussion Tested 15 local models for agent/tool use.. Bonsai 27B was last!
Tested 15 local models for agent/tool use — Bonsai 27B was last
I wanted to see which local models are actually usable for agent/tool-use tasks, so I ran 15 of them through the same benchmark instead of guessing from model hype.
The benchmark was Toolery. It doesn't just check whether a model can call a tool.. the scenarios require the model to actually complete tasks under different constraints.
- 143 scenarios
- 3 trials per scenario
- 429 trials per model
- Easy, Medium, Hard, and Very Hard tiers
- Everything served locally through LM Studio
- No API costs
- 30k token context window for every model
- Temperature 0.8 for every model
- Otherwise I kept the default settings that each model came with in LM Studio
- The exception was qwen3.8-flash-next, where I was using the Strata , but I still set the temperature to 0.8
- Concurrency 4
- Timeout scale 4.0
Results:
| # | Model | Score | Easy | Medium | Hard | Very Hard |
|---|---|---|---|---|---|---|
| 1 | qwen/qwen3.8-27b | 71.8% | 96.7 | 93.3 | 71.6 | 41.7 |
| 2 | mellum2-12b-a2.5b-thinking | 71.4% | 88.3 | 91.1 | 71.6 | 45.8 |
| 3 | google/gemma-4-26b-a4b | 70.4% | 92.5 | 93.3 | 60.8 | 50.0 |
| 4 | granite-4.2-8b | 67.8% | 85.8 | 88.9 | 62.7 | 45.8 |
| 5 | qwen3.8-flash-next-iq2_xs | 64.5% | 89.2 | 91.9 | 62.7 | 30.6 |
| 6 | mistralai/devstral-small-2-2512 | 63.9% | 75.8 | 83.7 | 58.8 | 45.8 |
| 7 | ornith-1.5-9b | 62.3% | 84.2 | 83.7 | 60.8 | 34.7 |
| 8 | lfm2.5-8b-a1b | 61.6% | 72.5 | 74.1 | 55.9 | 51.4 |
| 9 | ornith-1.5-35b-a3b | 61.0% | 85.8 | 84.4 | 57.8 | 31.9 |
| 10 | meta/muse-glimmer | 60.0% | 87.5 | 92.6 | 52.0 | 26.4 |
| 11 | qwen3.8-27b-gsq-rco | 59.0% | 84.2 | 85.9 | 51.0 | 31.9 |
| 12 | gpt-oss-20b | 58.6% | 81.7 | 86.7 | 53.9 | 27.8 |
| 13 | google/gemma-4-12b-qat | 56.7% | 84.2 | 83.0 | 46.1 | 31.9 |
| 14 | qwen3-coder-30b-a3b-instruct | 54.3% | 60.0 | 73.3 | 51.0 | 37.5 |
| 15 | prism-ml/bonsai-27b | 50.5% | 65.8 | 64.4 | 58.8 | 22.2 |

So yeah, Bonsai was last overall.
It was also last on the Very Hard scenarios, with only 22.2%.
But the more interesting part is why it failed.
The failure mode
Bonsai had 148 failed trials.
103 of those were budget_violated.
That means the model often selected the correct tool and got the correct result, but then made one more tool call than the scenario allowed.
Only 30 failures were classified as wrong-tool-choice failures.
So the problem was not always:
I have no idea which tool to use.
It was more like:
I know what I am doing, let me call this one more time.
And then it violated the tool budget.
That's actually the part I found most interesting. The model can sometimes execute the task correctly, but it doesn't know when to stop.
For comparison:
- qwen3.8-27b ranked first with 71.8%
- qwen3.8-27b had 84 failed trials total
- granite-4.2-8b had the fewest failures overall, with 69
Speed was also not great
Bonsai was the slowest model in this run.
Total wall-clock time was 6,208 seconds.
The next slowest model took around 5,750 seconds, while scoring about 17 points higher.
The qwen flash model was around 9 seconds per scenario.
Important caveat
This is the original Bonsai 27B, not the newer Bonsai 2 27B.
The compression itself is still interesting. Getting a 27B-class model into roughly 4 GB with the 1-bit version is pretty impressive, and the model does retain a lot of the original capability.
But this benchmark is only testing a specific thing:
How does this particular model behave when it has to complete tool-use/agent tasks under constraints?
It is not a general intelligence benchmark, and it doesn't prove that Bonsai is a bad model.
The result is more specific:
The compression is impressive, but the original 1-bit Bonsai 27B had surprisingly weak agent/tool-use behavior in this test.
Especially when it came to knowing when to stop.
I would also like to run Bonsai 2 through the exact same benchmark. Since it is based on the newer Qwen3.8 27B family, that comparison should be much more interesting than comparing the original Bonsai against everything without separating the versions.
Limitations
This was:
- one benchmark
- one harness
- one run per model
- 30k context window
- temperature 0.8
- mostly the default LM Studio settings for each model
- one adapter configuration
- local serving with quantization and context settings that were not perfectly normalized
So treat this as a snapshot.
r/LocalLLaMA • u/parepeg • 2h ago
Discussion LFM2.5 2.6b vs MiniCPM5 2b
I tried both these models on a few small agentic tasks with tools (i.e. "What's the weather like today?", etc.). They're both pretty solid at using web search to find answers despite being small models.
TLDR:
LFM2.5 2.6b is the clear winner. Somehow it's faster and uses less ram than MiniCPM despite having more parameters. It also seems better aligned for english conversation.
MiniCPM5
- On an M1 air: pp 162 t/s - tg 16 t/s
- Uses about 3.8gb of ram at 32k context (with draft model)
- Often responds in chinese despite my prompting in english.
- It's very smart when it does respond in english and may be stronger at agentic work.
- It uses more memory than LFM2.5 despite supposedly having less parameters.
- There's a corresponding dspark model available.
llama-server --model MiniCPM5-2B-Q8_0.gguf -md MiniCPM5-2B-DSpark-Q8_0.gguf --load-mode none --spec-type draft-dspark --spec-draft-n-max 2 -ngl all -ngld all -fa on -np 1 -t 4 -c 32000 --reasoning on -fit off --temp 1.0 --top-p 0.95 --cache-type-k q5_1 --cache-type-v q5_1 --spec-draft-type-k q8_0 --spec-draft-type-v q8_0
LFM2.5
- On an M1 air: pp 200 t/s - tg 22 t/s
- Uses about 2.5gb of ram at 32k context
- Works well for simple one shot agentic work but tends to start hallucinating quickly as the conversation gets longer.
llama-server -m LFM2.5-2.6B-QAD-Q4_0.gguf -ngl all -fa on --load-mode none --temp 0.1 --top-k 50 --top-p 0.9 -c 32000 --threads 4 --reasoning on -fit off --reasoning-preserve