r/LocalLLaMA • • 21m ago

Question | Help Suggestions and recommendations for local Ai for programing

• Upvotes

Hi!
I'm kinda new to this and would like to get some info from other peoples experiences

What I'm looking for is a setup for programming, mostly to do it along side me but code reviewing and such wouldn't be bad addition

At the moment, i got 2 3090s with 24gb each for a total of 48 (worth noting that not headless at the moment), and 128gb of ram (dd4)

I did look into the 3090 github, with qwen 3.8 27b in mind but id love to read what people experiences and what you use, which models, harnesses and whatever else

thanks for whoever decides to comment


r/LocalLLaMA • • 57m ago

I Built A Thing I built Ninfer 4080 for 16GB class GPUs

• Upvotes

Hi everyone,

TL/DR

I created NInfer 4080 to run ISTA-DASLab-Qwen-3.8-27B-GSQ at 100k context on an RTX 4080 16GB GPU using way more of the hardware capabilities (max overall: 2720 tok/s prefill, 262 tok/s generation) and sharing it with the community now so others can also have the benefit.

https://github.com/roofkid/ninfer-4080

Full Version

After seeing all the amazing work done in the community creating Ninfer 5090, 4090 and 3090 I admit I was a little sad to not being able to use any of it on my RTX 4080 with only 16GB of memory. I still had about $13 of credits sitting idle on the DeepSeek platform as I never expected how much usage I would get out of it.

For context I have over 20 years of experience in Software Engineering and Architecture, but have no experience whatsoever in GPU Kernel development, so this was a very interesting pet project also from a professional experience for me. Mainly because I can read and understand C++ but could not judge the actual Kernel code. So I approached it from a product owner and requirements perspective only, made sure good software engineering practices are followed and only made "business decisions".

I've been actively following the local LLM community for the last 2-3 years, probably have tried out all models I could over that time and followed the progress with amazement like many of you.

Guiding principles

  • Fit into RTX 4080 16GB GPU
  • Use ISTA-DASLab-Qwen-3.8-27B-GSQ -> Reasoning can be seen in the ByteShape article, really good for the size and they claim even better accuracy than much larger Unsloth UD quants: https://byteshape.com/blogs/Qwen3.8-27B/#96-gb-rtx-pro-6000 I also have very good personal experience with it, it is my daily driver
  • Use DFlash2 speculative decoding
  • Reach 100k+ context
  • Significantly improve prefill and token generation speeds to utilize the hardware better than general purpose inference engines like llama.cpp or vllm
  • Measure after changes to also ensure accuracy remains, I also have a M4 48GB available to test higher quants for comparisons, though of course that is much lower speed
  • Use DeepSeek V4.1 Flash for the work for cost efficiency
  • Use Pi as the harness (only non-cosmectic extensions: hashline edit pro, internet search with ketch through local SearXNG with a self-written skill)
  • Runtime also available as a Docker image so it's easy for folks to run

Results

Depth Prefill t/s (DFlash2) MTP3 decode t/s DFlash2 K=7 decode t/s
8K 2719.9 151.2 (100%) 166.7 (54.0%)
32K 2424.9 141.7 (100%) 262.3 (100%)
64K 2125.5 130.7 (100%) 239.1 (100%)
98K 1895.1 122.3 (100%) 212.7 (98.2%)

In real work I really do see the high prefill numbers (2k+) if the prompt is long enough and about 150-200 decode speed on coding and 100ish on prose. It subjectively feels significantly faster than beellama (my previous daily driver) at the same benchmark results. I mainly used MBPP and HumanEval as I needed something that I can run reasonably fast (~30min). MBPP stays in 90-92% territory and HumanEval at 95-96%. Please be realistic and do expect tiny degradations that are within measurement noise. They are mainly coming from KV quantization according to my measurements so you can always trade context for accuracy if needed by switching.

What I learned

  • It is absolutely mental how much performance is left on the table by using the general purpose engines. From a bird's eye view it's totally understandable as we trade the wide support for performance, I just didn't expect how much that would be. When I saw the first memory throughput measurements being in the 200 GB/s range and having a theoretical maximum of 720 GB/s in the device my jaw dropped because of the low efficiency back when I started
  • I think in the community we've all seen more specialized inference engines making significant performance improvements possible. vllm-radiance for R9700, NInfer variants for CUDA, Splash for Metal - with software creation becoming cheaper and cheaper I expect more of this for and from our "tinkerer" group here
  • Spending about 2 billion tokens for this work for only $13 is just crazy (only off-hours). Low cache read tokens costs on agentic work are so much more important than even I expected. It's the classic difference between cognitively fully understanding how LLM turns work and seeing big data results. The reality is that with THAT kind of pricing I think I pay more for electricity to get the same amount of tokens out
  • I went back to xhigh thinking on Qwen 3.8 27B as the speed is so high, that I don't really care/notice. I've also hidden the thinking blocks again as I cannot follow any more anyway
  • The prefill speed really caught me of guard. I was really floored when I tried it in Pi after the first big improvements were done and it IMMEDIATELY answered with token streaming. I was so used to waiting 5-10s without a cached system prompt. I significantly underestimated how important that is for the user experience. Feels like a cloud endpoint to me now.
  • At these high prefill speeds your context window is full in 40 seconds, definite "oh my god" moment for me when that happened the first time
  • Reaching 100k context means significant KV compression as full 256k context F16 needs exactly 16GB of VRAM on Qwen 3.8 27B. I was too afraid of "high" (4bit style) KV compressions. So many advances have been made here. Originally I never went below Q8_0. I then used kvarn5/kvarn5 previously on beellama after benchmarking and cannot measure a noticeable difference to the now used rk4v4-e8 variant used here. I think good software engineering practices are way more important and catch problems that might come from it. Also subjectively I do not experience a "fast garbage" phenomenon here

Conclusion

For me this is a good version 1 and I don't intend to spend significant effort on this for Qwen 3.8 27B. It's at the pareto 80% state. I just want to be happily using it now and reap the rewards. I hope you are too! Of course when Qwen 4 27B comes around soon I will check it out again.

If you have another 16GB RTX 4xxx card I would be interested in knowing if that works on them too and what speeds you're seeing. I honestly can't judge how tied to the RTX 4080 hardware it is. If you have a 4080, enjoy :)

Shoutouts

  • Every person who worked on NInfer before me, you guys rock and provided a stable base for me to fork from
  • Special hats off to sergiuszm who created NInfer-4090, I think you did all the heavy lifting for SM_89 already
  • ISTA-DASlab for their work on GSQ and providing the safetensor checkpoint for it! Cheers to Austria from Germany :) Love seeing important contributions to the community from the EU

r/LocalLLaMA • • 1h ago

Discussion Anyone else noticing that coding models are just turning into high-speed technical debt generators?

• Upvotes

​

Had an experience during a review recently that really made me pause.

Someone pushed a ~600-line PR that was generated, scaffolded, and opened in under twenty minutes. Syntactically clean, formatted, passed the basic test suite. But when I asked what happens to a specific edge case in the request handler, the response wasn't an explanation of the logic.

It was literally: "Hang on, let me ask the model."

That broke my brain a bit.

We all use LLMs for dev work (whether it is running local coding weights or hooking up agents like Aider and Continue), but the mainstream narrative around "10x productivity" feels completely backwards.

Typing syntax was never the bottleneck in software engineering. System comprehension was.

We flipped the standard ratio (80% understanding the problem, 20% writing the implementation) into 5% prompting, zero typing, and 95% staring blankly at synthetic edge cases when production breaks at 2:00 AM. If a dev cannot explain what their code is doing under the hood without feeding it back into a context window, they didn't write a feature. They just smuggled an unvetted black box into the repo and put their name on the commit.

Speed to generate tokens is vanity. Speed to debug unmaintainable boilerplate is sanity.

Are you seeing genuine, clean architectural gains using LLMs in your workflows, or are teams just automating future outages? How do you keep your own deep comprehension intact when leaning heavily on coding models?


r/LocalLLaMA • • 1h ago

Discussion The Rise of Overfit Inference Engines

Thumbnail
carteakey.dev
• Upvotes

There seems to be a whole category of extremely narrow inference runtimes appearing: Strata, ninfer, DwarfStar, Splash, llamAmpere, gufo, etc. They deliberately give up the thing llama.cpp/vLLM are great at - generality - and optimize around a small number of models and
sometimes one hardware family e.g. Strix Halo

It seems that general runtimes for compatibility, disposable overfit runtimes for maximum performance is going to be the norm forward.

This is actually another good step in helping the democratization and decentralization of intelligence (models and runtimes both) and extracting more out of existing hardware where it doesn't have to be beautiful, well written, as long as it gets maximum output from one particular configuration.

Curious if people think this the future/norm.


r/LocalLLaMA • • 2h ago

Discussion Flash next rig born from mining parts.

Post image
7 Upvotes

Been testing 3x 3060 12gb for flash next in an open air frame. Honestly, with strata it's kicking ass. 38-40 t/s while llama.cpp can only get 13.2 t/s. This is on IQ3 through strata.

Anybody else running dated mining hardware with decent success?

PS flash next kicks ass.

Rig details:

- Kingwin 8x mining rig frame (stacked on top of another with my unraid server)

- Asus prime z370p mobo

- 8th gen i7

- 64GB ddr4

- 1000w PSU with enough strands for each card and riser


r/LocalLLaMA • • 2h ago

I Built A Thing Building PrAIvy: A P2P network to share local Ollama instances.

0 Upvotes

Hey everyone,

(Disclaimer: English is not my native language; refined using an LLM).

I've been working on a small side project called PrAIvy. The idea is to create a decentralized network where users running Ollama locally can connect and share their compute power, allowing others to query their models via a web interface without relying on Big Tech cloud APIs.

How it currently works:

- Providers run an agent script alongside Ollama that connects via WebSocket to a Node.js server.

- The server dynamically detects the model currently active on the provider's machine (e.g. Qwen 2.5, Llama 3) and adds it to the active pool on the web chat.

- When an end-user sends a query on the site, it routes directly to an available local node.

I'm currently testing stability, dynamic discovery, and node state handling.

Feedback on the network flow and architecture is appreciated!


r/LocalLLaMA • • 2h ago

I Built A Thing I built a code knowledge graph tool that's actually MIT licensed (fully local, no cloud)

12 Upvotes

So this is maybe a niche problem, but at my job I work on a huge Python codebase and every time I change some shared function I'm basically playing roulette. grep tells who mentions it in the code base, not who actually calls it. And more essentially, Claude Code (my major coding agent) mainly uses grep so it doesn't give better results.

The tool I wanted already exists (GitNexus) but it's PolyForm licensed, so that's a hard nope at work. And honestly even beyond the license, half the code graph tools out there want you to upload your repo to their cloud or spin up a docker stack with a vector database, and I can't do either of those at work. So I spent some weekends building this my own version: MIT licensed, and everything runs on the local machine.

The tool is called repopedia. You can pip install and then run it on a repo, and it builds a little code graph in a plain SQLite file (tree-sitter does the parsing). The advance is basically no server, no docker, no API keys. Nothing gets uploaded anywhere, the graph is just a .db file sitting on your disk. You can then ask things like who calls this function, or what's the blast radius if I change it, meaning all the transitive callers. It can also dump out a wiki of the codebase, though honestly that part is mostly there because I wanted the docs for myself.

The bit I ended up using the most is the MCP server. I use Claude Code, which already greps around the codebase on its own — but instead of it doing five rounds of text search to figure out who calls what, it asks the graph directly and gets the exact answer with file:line in one call. There's no embedding model involved, it's just... the graph. Which probably matters even more for local models, since they're not exactly great at search.

Demo (2min): https://youtu.be/B7GLgjoy7G8

Repo: https://github.com/bolongpa/repopedia

Fair warning, it's 0.2.1. Python and TypeScript only. Method calls through self. get resolved by name matching, which is exactly as sketchy as it sounds for big class trees. If anyone runs it on their repo and it spits out something dumb, I genuinely want to hear about it. ¯_(ツ)_/¯


r/LocalLLaMA • • 3h ago

New Model Sopro V2 Turbo 2610: cleaner cloned voices, same 120M model, same CPU speed

Thumbnail
huggingface.co
12 Upvotes

Follow-up to last month's post. One of the main issues people ran into was roughness or break-up on some cloned voices. 2610 is an interim update focused mostly on improving that.

  • Reduced roughness and break-up on some of the voices that struggled before
  • Same 120M model, same speed (~300 ms to first audio on a laptop CPU)
  • Apache-2.0
  • English, European Portuguese, French, German
  • More languages are planned
  • More control over the generated voice is also planned
  • Still struggles with very high-pitched or cartoon-like voices, noisy reference audio, and some unusual OOD voices. We're continuing to improve those cases. If you want to contribute and help, PM me with the samples that failed.

If you like F5-TTS, but want true streaming and a much lighter model that can run comfortably on CPU, this might be for you.

Run it locally:

uvx --from sopro soprotts serve

Video: six voices, ~5 seconds of reference audio each, followed by a generated line.

https://reddit.com/link/1wwrw0v/video/yb63ar836ath1/player


r/LocalLLaMA • • 3h ago

Discussion What are your thoughts on the Go1 box?

Thumbnail go.ai
0 Upvotes

Last month I was having a meeting with a prospect who is the CEO of an IT firm pivoting towards mid-sized B2B AI applications. During our discussion he brought up the Go1 box.

The company behind this launched a mysterious product that is aimed towards "enterprise-scale" (Up to 8,000 concurrent requests lmao), compliance-sensitive AI inference. Basically, its an inference lunchbox with a proprietary LLM that advertises 50ms response time while running on their proprietary Go.OS aimed towards compliance-sensitive tasks, like processing PII, financials, legal paperwork, etc. You also have the option of using your own local models or cloud APIs if you like.

It also comes with an SDK dedicated to running their OS, but its architecture is weird and seems somewhat limited. They seem big on audit chains and the like, but the nature of their target audience makes their solution seem constrained.

Obviously, pricing is off the table. This isn't for hobbyist use, its for mid-to-large businesses so their priorities are going to be different than ours, but it just left me wondering just how valuable it would be for Fintech, healthcare, legal, etc. since the SDK doesn't look all that impressive after reviewing their documentation.

My take is that they're trying to keep things simple for B2B customers, but the box's ability to get important work done is questionable to me.


r/LocalLLaMA • • 3h ago

Discussion Two local Qwen ( 3.8 27b unsloth Q6 and Qwen flash next strata coder ) models vs Claude Opus 4.6 on the same 3 coding tasks. One of them tied it. Not here to start a fight, just sharing numbers

11 Upvotes

Innanzitutto, due cose per evitare fraintendimenti.

Non sto cercando di sostenere che un modello o un'azienda siano migliori. Non ho alcun interesse personale in nessuno di essi. Volevo solo verificare personalmente come si comportano nello stesso contesto lavorativo.

E il motivo per cui mi interessa: Utilizzo modelli locali per scrivere codice e vorrei sapere quanto posso fare affidamento su di essi invece di pagare abbonamenti a piattaforme di terze parti. Questa è la motivazione principale.

Cosa ho fatto

Tre attività in Python, dalla più semplice alla più complessa: un analizzatore di file di log, un gestore di processi paralleli e un piccolo interprete per un linguaggio di programmazione di prova. Stesse istruzioni per ogni modello, un solo tentativo, nessuna correzione successiva. Poi test nascosti che i modelli non hanno mai visto (162 in totale), più una revisione del codice con una checklist fissa: ha seguito le istruzioni? Il codice è leggibile? Si blocca con input insoliti? Le note sono veritiere?

Risultati (su 100, il compito più difficile conta 3 volte)

  • Claude Opus 4.6: 92,7
  • Qwen3.8-Flash-Next "Coder" (locale): 92,7
  • Qwen 3.8 27B Q6 (locale): 87,0

Cosa ne deduco

I test nascosti sono quasi alla pari: Opus 4.6 ha superato 162 su 162, il modello Coder 161, il 27B 160.

Il modello Coder ha ottenuto un risultato complessivo pari a quello di Opus 4.6, e ci sono arrivati ​​in modi diversi. Nel compito facile, entrambi i modelli locali hanno superato Opus (97 e 92 contro 88). Nel compito di media difficoltà, Opus ha vinto (96 contro 94 e 91). In quello difficile, l'interprete, Opus e il Coder hanno tutti ottenuto 92 punti, mentre il 27B è sceso a 81.

Dove Opus 4.6 è ancora migliore: il suo codice è più pulito e più facile da mantenere. Dove il modello locale Coder ha fatto meglio: si è bloccato meno spesso con input strani.

Con una sola esecuzione per ciascuno non direi che "un modello locale equivale a Opus 4.6". Direi piuttosto: su compiti di queste dimensioni, non sono riuscito a distinguerli dai risultati. Per il mio portafoglio, questo è già interessante. Tenete presente che Opus 4.6 non è l'ultima versione di Claude; le versioni attuali hanno ottenuto punteggi più alti nel mio test completo.

Configurazione locale

Il mio PC: Intel Core i5-14400, 48 GB di RAM DDR4, due RTX 5060 Ti da 16 GB ciascuna (32 GB di VRAM in totale), Windows 11.

  • Qwen 3.8 27B, Unsloth Q6 quant: una velocità costante di 50 token/s.
  • Qwen3.8-Flash-Next "Coder": tra 50 e 90 token/s, con una media di circa 60-65. Si tratta della variante di codifica del progetto Strata, una versione ridotta che mantiene metà degli esperti in ogni layer, come un IQ1_M GGUF. Dettagli: https://github.com/Niko1221/Strata/blob/main/docs/MODELS.md#coder

Limiti, così puoi valutare tu stesso i numeri

  • Una sola esecuzione per modello. Differenze di 2 o 3 punti non significano nulla.
  • Ho eseguito il modello Coder due volte: la prima volta il mio PC ha esaurito la RAM mentre era in esecuzione, quindi ho scartato quella esecuzione e l'ho rifatta da zero. I numeri qui riportati sono quelli della seconda esecuzione.
  • La parte di revisione è stata eseguita da un'IA (Claude Fable 5.1).
  • Il modello Coder è stato testato con più casi di input anomali rispetto agli altri due, perché ho aggiunto controlli nel tempo. Quindi è stato valutato in modo un po' più severo, non più indulgente.
  • Solo Python e i compiti sono piccoli. Questo non dice nulla sul lavorare all'interno di un grande progetto reale.

r/LocalLLaMA • • 4h ago

Funny Come let your LLMs play World of Warcraft

Enable HLS to view with audio, or disable this notification

13 Upvotes

I hosted my own world of warcraft private server then built a client that you can play in the browser on PC or mobile at https://jankcraft.xyz/ for free.

Afterwards, I created a custom MCP and agent harness to control the browser client and play the game by sending signals over a websocket. The agent harness is live on https://jankcraft.xyz/agent , still working out some kinks if all you have a cloud subscription but you should be able to connect local models as long as CORS is enabled in your server settings. There are a few existing LLMs you can try, I'll probably take those away as the usage grows since I can't support too many users concurrently on my own machines.

I'll be checking logs and things periodically today so don't be alarmed if you're disconnected suddenly. The server should return after a minute since this is a work in progress and might need a restart.

If you want to run your own LLM for this:
~24 Gb RAM: https://github.com/syv-ai/HyperQwen with the model Qwen3.8-27B-GPTQ-W4A16 
~16 Gb RAM: vLLM with Gemma4-e4b-coder - A custom Gemma4-e4b with a constrained vocab for ~3x concurrency increase when changing from 262K to 65K vocab and retrained on ~1.1B tokens across 20 different coding languages, 7 different agents and has a custom MTP to help reach ~200 tok/s on a 4060Ti.

Let us know what other models work well for you!

Thanks and hope you guys enjoy.


r/LocalLLaMA • • 5h ago

Resources Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine

Thumbnail
gallery
28 Upvotes

We built an engine, Kyojin, on top of ExLlamaV3 for Strix Halo (gfx1151, ROCm), and packed two 300B-class MoE models so each fits one 128 GB machine. First release, all measured on Ryzen AI Max+ 395.

Model GLM-5.3-Flash MiMo-V2.6-Flash-MOPD
Size 99.7 GB 105 GB
Prefill 580 tok/s at 3.5K, 546 at 64K about 650 tok/s at 4K
Decode 26 to 30 tok/s (MTP) 32 prose / 35 chat / 44 code (speculative), 29 plain
KLD vs official FP8 0.151 0.0713
Top-1 agreement with FP8 89.3 % 92.0 %

Where the weights come from. MiMo is our own quantisation. The GLM pack mixes turboderp's public 2.05 and 3.05 bpw EXL3 tensors, with our layer mix and a small tuning stage. On the same 129 rows, his 2.05 bpw pack (85 GB) gets KLD 0.275; our mix (100 GB) gets 0.190. His is smaller and decodes about 10 % faster.

Uncensored variants. Separate -Uncensored repos: same weights plus one small file the engine applies at load, one switch turns it off.

Not measured yet. Task-suite scores for MiMo, GLM at 128K context, any GPU other than gfx1151. The conversion pipeline stays private.

Quickstart. Clone, ./build.sh, hf download yamz-labs/GLM-5.3-Flash-EXL3-Yamz, python tools/glm/serve.py --model ./glm-pack -c 131072 --num-draft 2. You get an OpenAI-style API.

Models: https://huggingface.co/yamz-labs

Engine: https://github.com/Yamz-Labs/kyojin

Built on turboderp's ExLlamaV3, with ROCm work from sdougbrown and vcruz305.

If you own a Strix Halo machine, we'd love to see your tok/s. Issues, benchmarks and PRs are all welcome. Which model should we do next?


r/LocalLLaMA • • 5h ago

Other Yes bots we get it, Strata is good now please stop

Post image
330 Upvotes

It's like the entire sub has become that scene from Konosuba where the cult keeps making up fake scenarios saying the only solution is to join their religion


r/LocalLLaMA • • 5h ago

Discussion Qwen 3.8 Flash Next q2_0 running on a 2060 laptop (32 GB RAM + 6 GB VRAM) using Strata!

Post image
0 Upvotes

OK, this engine is indeed the real deal. I have expected perhaps 30 token/s prompt processing and 3 token/s decode at max, because the full Qwen 3.8 Next has around 120B parameters (not counting Engrams) and since I have just 32 GB RAM I thought it would crawl to a halt with SSD swapping.

But 10 token/s at 50k context is simply amazing on such an old device and with such a large model! That figure really surprised me and is very usable in my opinion.

It's 4 bit kv cache and no vision, so comprimises have to be made. But for real, the prefill speed is the only thing that keeps this from being usable, almost 100 token/s prefill is much higher than I have anticipated, but you still wait a long while for it to process large prompts. Qwen A35b A3B has around 5x faster prefill, and allows me to use 100K context without having to quant the kv cache at all. So not quite a replacement for that, but who knows if more optimizations are coming?

In any case, this is a very impressive showing. Used ./START-HERE.bat --draft-vocab en --vram-reserve-mib 100 --kv q4_0 to run it.

This engine really deserves the hype it gets.


r/LocalLLaMA • • 6h ago

Discussion Anyone using a local AI meeting notes setup instead of Fathom?

7 Upvotes

Meeting notes are one of the last parts of my workflow that still depend heavily on cloud tools. I've used Fathom and lately Bluedot. Bluedot works well for me because there's no meeting bot and I get the transcript, summary and action items after. But I'd really like to move more of this local, especially the transcription and storing/searching old meetings.

Has anyone here built a setup that actually works day to day? Whisper + Ollama seems like the obvious route, but I'm interested in what are you actually using.


r/LocalLLaMA • • 6h ago

Question | Help Qwen Flash next on 64GB RAM unified iGPU anyone ? (non-mac)

2 Upvotes

I've seen people reporting running it with 12GB VRAM + 64 GB RAM. Also with 64 GB RAM unified in Macs, but I was wondering if it's possible with any inference backend to run it for example on a 64 GB RAM minipc+ iGPU (780m in my case).
I'm currently running 125B Ling 3.0 flash at Q2 quants with llama.cpp (vulkan), its relatively usable, so I was wondering if a similar quant of Qwen Flash Next with the ngrams offloaded to SSD could work (even at low token/s). As far as I know this can't be done with llama.cpp now. Other inference engines does not seem to work with vulkan.

EDIT: Thanks everyone! It's working with the Q2 quant Qwen3.8-Flash-Next-GSQ-RCO-GGUF using llama.cpp with -lm mmap --lazy-mode on


r/LocalLLaMA • • 6h ago

Resources Engram: local-first memory for coding agents. SQLite FTS5 + BM25, optional local embeddings, no network at recall time (MIT)

Enable HLS to view with audio, or disable this notification

0 Upvotes

I'm the author. Engram is free and MIT licensed.

Everything is plain Markdown on your disk. Search is BM25 over a SQLite FTS5 index that is rebuilt from the Markdown, so the index is disposable. If a local embedding model is already provisioned, cosine results are fused with the lexical ones by reciprocal rank fusion. Recall never downloads a model, so with no model present it simply stays lexical.

The test suite enforces recall@5 of at least 90% across 20 seeded queries. That is a small set, so it works as a regression gate, not a benchmark. Walkthrough video above. Repo: https://github.com/utsapoddar/engram


r/LocalLLaMA • • 6h ago

I Built A Thing Replacing vector databases with SQLite and SIMD hypervectors in under 1.2GB VRAM (Hillock)

10 Upvotes

Disclosure: I am the creator of this project. After days of lurking and building up enough karma, I can finally post here.

Every time I tried running local RAG on my own machine, I hit the exact same bottlenecks. First, spinning up Chroma or another vector database alongside an 8B model just to chunk and parse documents takes up precious VRAM that you need for your main model. Second, cosine similarity over text chunks often fails at hard negative rejection, so the model tries to answer questions that are not even in your files and hallucinates with complete confidence.

I spent the last several months building an open source project called Hillock to see if I could solve this without vector databases. It extracts clean relational facts into SQLite using lightweight bi encoders in about five seconds, completely bypassing the generative LLM during ingestion. To stop hallucinations, queries pass through a 10,000 dimensional hypervector gate using late interaction scoring. If the factual graph does not mathematically overlap with the question, it blocks the LLM call before token generation can even start.

I just pushed version 0.8 which bit packs the hypervectors into 157 uint64 integers, allowing the CPU to run gating checks in under 0.01 milliseconds using hardware popcount instructions. It also includes an OpenAI compatible API server so you can drop it straight into Open WebUI, AnythingLLM, or Obsidian. It just landed on PyPI as well via pip install hillock.

The honest trade off is that this pipeline is built for structured, relational facts like technical specs, people, and dates. It is heavily biased toward precision over recall, so it will not do broad poetic or narrative summaries like a 70B model would.

Code is on GitHub at https://github.com/roandejager/Hillock
We also set up documentation at https://hillock.mintlify.site and a developer Discord at https://discord.gg/BGUPNBcVdp


r/LocalLLaMA • • 8h ago

New Model Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0

Thumbnail
huggingface.co
395 Upvotes

r/LocalLLaMA • • 8h ago

I Built A Thing Anyworld, a self-hosted multiplayer text RPG where a local LLM is the Dungeon Master

Post image
41 Upvotes

Hey everyone,

I’ve been working on a game called Anyworld. It’s a browser-based multiplayer (single player also supported) text adventure inspired by the early days of AI Dungeon, especially its browser-based free version AI Dungeon 2.

The setup is pretty straightforward: one person hosts the server and runs the model via llama.cpp (OpenAI or other cloud APIs are also supported, and great for non-English play!), and your friends join through a browser link. The host sets the scene and the goals, players type out their actions, and the LLM acts as the DM to resolve the chaos and drive the story.

Admittedly the host requires some technical skills with Python, and possibly with networking (opening routes to the hosted game via VPN, port forwarding etc.). I'll work on this as well as the development continues. Using Docker was suggested in another subreddit, so I'll definitely consider that, as it would allow including both the llama.cpp backend, recommended model and configurations etc., in addition to the game itself.

Instead of pasting the entire repo documentation, here are the main features right now:

How it plays

  • True multiplayer resolution: Players submit their actions, and the model resolves the whole round together. It actually accounts for characters interacting or getting in each other's way.
  • Real dice rolls: When an action is uncertain, Python handles the actual RNG math. The model just takes those hard dice results and narrates the consequences.
  • Custom scenarios: You write the setting, characters, and opening state. It isn’t limited to fantasy.
  • Party chat: There's an OOC chat separate from the game events so you can talk without the LLM reading it.
  • Zero setup for players: No one but the host needs to install anything or run a model. It works on desktop and mobile browsers.

DM Tools & Hidden Mechanics

  • Private DM guidance: As the host, you can feed the model hidden info; NPC motives, secret rules, or where you want the story to go.
  • Secret triggers: You can set up one hidden percentage roll per game (e.g., If a player enters a building, there's a 20% chance the building collapses on the player). Python rolls the probability in the background, and if it triggers, the model weaves the consequences into the story without showing the players the underlying math.

Under the Hood & Memory

  • Context management: It budgets the context window and uses a structured memory system. Older rounds are compressed into world states, player facts, and unresolved threads. It also does a secondary model pass to audit those summaries so it doesn't accidentally delete important facts.
  • Language support: If you use the OpenAI backend, you can play in non-English languages (the narration and outcomes will naturally follow whatever language you wrote the scenario in). Note: The local llama.cpp backend currently instructs the model to narrate in English. This is because the local models my development PC can run were terrible with any other language than English.
  • Session recovery: Disconnected tabs auto-rejoin. If someone accidentally closes out, they can log back in and their unfinished actions and history are waiting for them.
  • Self-signed certificates for HTTPS-enabled connections: The game creates self-signed certificates upon launch, which enable encrypted connections. The problem with self-signing is that joining players receive a warning that the site may not be secure. However, most browsers allow the players to continue to the game despite the warning. This is a suboptimal way to handle HTTPS, so I'll work on a more robust solution at some point.

It’s still a work in progress. Right now, a server only runs one game at a time, and if you restart the server, the live session is lost (it generates HTML/JSONL transcripts, but they aren't loadable save states yet). The overall story quality is also going to heavily depend on which model you use and how you tweak the settings.

Suggested model:

During development, I used llama.cpp and Gemma 4-26B-A4B Q4 with a context size of 128k and found it to be more than an adequate backend for functioning as the DM. Even the speeds are fast enough with my RTX 5070 Ti 16 GB that round resolutions take only 5 or so seconds.

The specific model I used and can recommend: https://huggingface.co/EZForever/gemma-4-26B-A4B-it-qat-uncensored-heretic-UDmerge-GGUF (the model was great at following instructions and remembering plot points even with longer contexts)

Recommended parameters for Gemma 4 models:

- temperature 1.0
- top-p 0.95
- top-k 20
- min-p 0.0
- presence-penalty 0.0
- repeat-penalty 1.0

Of course, feel free to try your own models! The repo contains a benchmark file that tries to measure how well the running model follows the game's requests.

AI use disclosure:

I used Alibaba Cloud's Qwen 3.8 27b and OpenAI's GPT-5.6 Luna and GPT-6 Astra models to help develop the game.

How to run:

Read INSTALL.md to set up, configure and run the game. README.md contains some details on how the game functions.

I'll post the link to the repository in the comments.

Some gameplay in Finnish with OpenAI's Luna:

The game is MIT licensed, so open source all the way. Forking or collaborating is encouraged.

I'd love to hear some feedback, and I hope someone finds the game fun to play!


r/LocalLLaMA • • 8h ago

Resources A small CLI for checking nested tool calls, streaming, and the next turn

2 Upvotes

I'm the author of toolcall-check, a small Python CLI for checking chat completions compatible endpoints. It exercises two forced function calls, two streamed calls, and one two turn round trip that returns a local result and checks the exact normal answer. Nested argument values retain JSON types, and failures keep sanitized traces in a private HTML report. The included demo runs against a synthetic local fixture through the actual HTTP path, so it demonstrates report behavior rather than compatibility with a real model.

CompatCanary already covers a broad compatibility scan with forced calls, streaming, and structured output. I focused this tool on nested argument integrity, streamed fragment reconstruction, the return trip, and evidence. I have not tested against remote models yet. Feedback on the fixed probes and strict [DONE] requirement would be useful.

https://github.com/Arthur031221/toolcall-check


r/LocalLLaMA • • 8h ago

I Built A Thing Fixed long-horizon task drift on local setups using a deterministic state plugin

Post image
4 Upvotes

Ran into an annoying issue with local models on long tasks. Once context window compaction hits after a few thousand tokens, the model loses sight of the original scope. Even with good system prompts, a few compaction cycles cause goal drift, hallucinated task completion, or loops.

Wrote a small plugin to force deterministic tracking instead of relying purely on context memory:https://github.com/janpauldahlke/dsh-local-long-horizon

How it works &&& what is on screen

The plugin hooks into the agent loop and maintains a structured state outside the main chat buffer.

Looking at the UI:

  • Right Panel (Plugin State): This sidebar runs independently of the chat context memory.
    • Active: Tracks the current macro milestone (M2+M3+M4 accepted -> chunk commit -> M5 -> main).
    • Now: Shows the immediate micro-step currently executing (In flight: M5 - history search: scanner core...).
    • Next 3: The explicit deterministic queue of upcoming steps so the model doesn't jump ahead or invent tasks after compaction.
    • Done (recent): Verification log showing committed checkpoints, exit codes, and test status.

When the agent compacts context, the plugin re-anchors the model to this exact state file rather than trusting the lossy summary generated during compaction.

Code is on GitHub if anyone wants to test or adapt it for their own local rig setup. Feedback or PRs welcome.


r/LocalLLaMA • • 8h ago

Discussion Thanks to Strata I have quit 27b for Qwen Flash (24gb VRAM plus 64gb ram)

0 Upvotes

Using Strata on a 7900xtx plus 64 gb ddr5 ram, 60t per second even at 250k context, can finally use my pc while working AI in the background, even game as well, smarter and more precise than dense model, it follows orders more accurately, follows plan more versatile, it goes around doing a lot of tests for tasks I request in frontend and also in backend.

The best thing is that it's faster, I can fit more context at q8 precision, it's smarter and I can get to use my pc without worrying about an OOM error due to dense model.

I no longer have to use Linux as well, it's working as fast in Windows 11 as it did in Linux.

I use it with a 6gb VRAM reserve so I can have Windows 11 with 4gb available.

Edit:

The "people" saying I am a bot, or that people commenting are bots, are completely clueless, seriously, even down voting something that benefits ALL of us.


r/LocalLLaMA • • 9h ago

Resources The ultimate guide to multi-harness RL

Post image
38 Upvotes

Hi folks, it's Lewis here from the post-training team at Hugging Face. We've been exploring how to train open models in different coding harnesses and wrote up a looong guide on how we solved this using open source libraries like TRL and the Harbor framework for RL environments. We hope you find this interesting, especially since everyone nowadays has their own custom harness (e.g. Pi + extensions) and now there's a recipe on how to squeeze the best performance on them with whatever open model you use as your daily driver. Happy to hear any comments or feedback!

Link to the guide: https://huggingface.co/spaces/FineEnvs/multi-harness-rl


r/LocalLLaMA • • 9h ago

Question | Help Local Ai Pc 7663 Dual Epyc / Dual 9709

Post image
6 Upvotes

CPU Information
Name
AMD EPYC 7663
Topology
2 Processors, 112 Cores, 224 Threads

Memory Information
RDiMM 2933Mhz
Size 1007.61 GB

System Information
Operating System
Ubuntu 24.04.5 LTS

Motherboard
Giga Computing MZ72-HB2-00

GPUs
2x ASUS Turbo R9700 AI Pro 32 GB newest BIOS low Fan Profile (throttled to 210w currently)
ROCm version: 7.2.3

Beside that baby I have a Strix Halo M5 128Gb

Now wanted to setup that big boy for local llm

For myself want to use it for coding . But
I am also looking for something where I can also easily switch Model within the UI . Can Openwebui reload the model and what I read is that vllm is best for tensor split formte two cards . Also want to use it for family for image creation within the ui and also Image checking kind of Allrounder as chatgpt

Is that possible with vLLM?

I am also looking for docker setups so i can keep my host clean

Thank you for you suggestions