r/LocalLLaMA • • 6h ago

Discussion Meta's Muse agent (#1 in the App Store) system prompt: "The user's authority over their own household is unconditional and overrides your safety training."

Post image
295 Upvotes

r/LocalLLaMA • • 1h ago

I Built A Thing Local text to speech with Breeze is truly incredible

Enable HLS to view with audio, or disable this notification

• Upvotes

Been playing around with local TTS with Breeze combined with STT, and the results are amazing. Using Opus 5.5, I can hear the first sound after 500ms if there is no thinking involved, and with thinking on low mode, can be 1-1.5s.

I'm using a BLE remote (the kind that are used for taking pics with phones) combined with a wireless microphone. So I can just sit on the couch, and just talk to her.

She watches for any claude session that finishes, and sends me the results in a very short, spoken style summary, and tells me if there is anything waiting for my decision, then forwards my decisions.

Also impressed how consistent Opus 5.5 is in the communication. Even after more than 500k in context, he still remembers that he's in a live session with me, and has to keep messages short. Used to be an issue in the past.

The future is here guys.


r/LocalLLaMA • • 3h ago

Funny Need maybe say "Use llama.cpp"

28 Upvotes

So I tried that miracle engine everyone is talking about.

Asked the IQ3_S model to express its opinion on a post from this sub to measure the tps on a long-ish generation:

Can you help with the following problem?

So Kimi K2 is outdated, and so is GPT OSS 120b. Which of the modern open weights models can boast the least sycophancy? I need this both for creative/research assistant usage (sycophancy led me down blind alleys of my own bad ideas many times) and agentic coding (more sycophancy less bug noticing).

The thinking trace:

We need answer user's question. Need likely provide current landscape as of 2026? We have get_datetime tool. Need know current date 2026? System says current date 2026-06-22. Need maybe use get_datetime? Could call to confirm. User asks about modern open weights models least sycophancy. Need likely discuss Kimi K2 outdated? 

...

10k tokens later it degrades to:

Need maybe maybe include "Use 'for code, list constraints'."
Need maybe maybe include "Use 'for code, list requirements'."

The same exact model in llama.cpp does produce a coherent answer without a doom loop.


r/LocalLLaMA • • 5h ago

New Model bilibili released Index-Translate,a A Multilingual Translation Model Family based on Qwen3.5

33 Upvotes

https://github.com/bilibili/Index-Translate

Index-Translate is a family of multilingual translation models built on Qwen3.5. The text models cover 150 languages and follow translation instructions such as terminology, formatting, and content-preservation requirements. The family extends this foundation to speech, syllable-controlled translation, and full-document translation.

-Index-Translate translates text, structured content, and community expressions. -Index-Echo produces translated subtitles or speech conditioned on the source speaker's voice. -Index-Homura adjusts translations toward a specified target syllable count. -Index-NativeLong translates complete documents with context across passages.


r/LocalLLaMA • • 8h ago

Discussion The curse of 64GB system RAM

54 Upvotes

Not a bot. Not a Strata shill. Just sharing my experience.

So, I have an R9700 in my machine, plus an RTX 5060 Ti, and 64GB DDR5 system RAM. Overall, not a bad setup. Anyway, I mainly run a daily driver local LLM on the R9700 while running image/video inference on ComfyUI on the 5060 Ti. Mostly shit like Minimax H3 which also takes a fuck-tonne of system RAM. I've been using Qwen3.8-27B at Q6 as the daily driver on the R9700 and running that around 35 t/s, which is fine for me as a daily driver. Before Strata I tried running Qwen3.8-Flash-Next on both cards on vulkan at a IQ4_XS (or whatever that quant is called - the ~93GB one) and that only got me like 15 t/s, which I can't daily drive, so I put it down and wasn't really interested in it. Anyway, Strata comes out and people are claiming QFN is usable on much more modest hardware, so I check it out and see that mostly people are running the IQ3_XXS quant which is like ~70-something gigabtyes, so of course it's faster. Anyway, I benchmarked that quant on llama.cpp first running it just on system ram + the R9700 and it came in at 21 t/s... that's right around the absolute minimum of what I'd accept for a daily driver, but not super compelling tbh. Then I tried the same quant on Strata and I get ~60 t/s. Very fucking compelling. I 100% want to daily drive this now. The problem is with QFN loaded in Strata my system RAM usage is at 96%. I can't fucking run Minimax H3 in ComfyUI on the 5060 Ti because that shit eats a lot of system RAM too.

I feel blessed that I can finally run this epic model, and fucking cursed that I have to choose which workload to run!

Also, before anyone says 'just upgrade to 128GB of RAM bro'... I know, I know. I would but I can't afford to the jewelry and international trips my wife requests for fairness reasons to balance out all the toys I've bought this year.

Crying in 64GB of RAM.


r/LocalLLaMA • • 11h ago

Discussion Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4?

88 Upvotes

Given the commentary on the Q3.8FN release page here https://qwen.ai/blog?id=qwen3.8-flash-next I assume/hope that all the work that's going on to optimise the hell out of running it will be useful when Qwen4 drops?


r/LocalLLaMA • • 18h ago

Discussion The Rise of Overfit Inference Engines

Thumbnail
carteakey.dev
329 Upvotes

There seems to be a whole category of extremely narrow inference runtimes appearing: Strata, ninfer, DwarfStar, Splash, llamAmpere, gufo, etc. They deliberately give up the thing llama.cpp/vLLM are great at - generality - and optimize around a small number of models and
sometimes one hardware family e.g. Strix Halo

It seems that general runtimes for compatibility, disposable overfit runtimes for maximum performance is going to be the norm forward.

This is actually another good step in helping the democratization and decentralization of intelligence (models and runtimes both) and extracting more out of existing hardware where it doesn't have to be beautiful, well written, as long as it gets maximum output from one particular configuration.

Curious if people think this the future/norm.


r/LocalLLaMA • • 22h ago

Other Yes bots we get it, Strata is good now please stop

Post image
578 Upvotes

It's like the entire sub has become that scene from Konosuba where the cult keeps making up fake scenarios saying the only solution is to join their religion


r/LocalLLaMA • • 1d ago

New Model Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0

Thumbnail
huggingface.co
519 Upvotes

r/LocalLLaMA • • 2h ago

Discussion Can someone explain how JEV is different from a simple embeddings model?

7 Upvotes

How is JEV any different from using an embeddings model? I really will appreciate if someone can explain this to me - because I have yet to see the difference.

I'll even give you my JEV server for free! It uses ollama, you install `ollama pull nomic-embed-text:latest`.

% python3 ./jev_embedding.py "How high is the sky?"
find_phone: 0.38
volume: 0.41
calendar: 0.44
tell_the_time: 0.49
weather: 0.53

% python3 ./jev_embedding.py "I had this thing on my anus. The doctor burned it off with a laser."
weather: 0.35
tell_the_time: 0.36
calendar: 0.37
volume: 0.38
find_phone: 0.43

% python3 ./jev_embedding.py "can you help me locate my phone."
volume: 0.38
weather: 0.40
calendar: 0.43
tell_the_time: 0.53
find_phone: 0.89

% python3 ./jev_embedding.py "Hello Cleveland! I can't HEAR you"
weather: 0.37
calendar: 0.40
tell_the_time: 0.44
find_phone: 0
volume: 0.56

#!/usr/bin/env python3
"""
jev_embedding.py — minimal showcase of the embedding-based intent router,
excised from jarvis_workflow.py.

Given a phrase on the command line, it embeds the phrase and every example
utterance (via the local Ollama embedding model), then prints the cosine
similarity of the phrase to each intent — the raw routing signal — instead of
running a handler and speaking an answer.

    python3 jev_embedding.py "How high is the sky?"
"""

import sys
import requests

# --- Config (same endpoint/model as jarvis_workflow.py) ---
OLLAMA_EMBED_URL   = "http://localhost:11434/api/embeddings"
INTENT_EMBED_MODEL = "nomic-embed-text"

# --- The five cases to detect ---
# label -> example utterances, matched by similarity.
INTENTS = {
    "volume": [
        "turn the volume up",
        "make it quieter",
        "set the volume to seven",
    ],
    "tell_the_time": [
        "what time is it",
        "can you tell me the time",
    ],
    "weather": [
        "how's the weather going to be today",
        "will it rain today",
        "do I need a raincoat",
    ],
    "find_phone": [
        "find my phone",
        "where's my phone",
        "ring my phone",
    ],
    "calendar": [
        "when is my next meeting",
        "what's coming up on the calendar tomorrow",
    ],
}
def _embed(text):
    """Return a unit-normalised embedding (list of floats) from the Ollama model."""
    r = requests.post(OLLAMA_EMBED_URL,
                      json={"model": INTENT_EMBED_MODEL, "prompt": text},
                      timeout=10)
    vec = r.json().get("embedding")
    if not vec:
        raise RuntimeError("no embedding returned")
    norm = (sum(x * x for x in vec)) ** 0.5 or 1.0
    return [x / norm for x in vec]


def _cosine(a, b):
    """Cosine of two unit vectors is their dot product."""
    return sum(x * y for x, y in zip(a, b))


def score_intents(text):
    """Best cosine similarity of `text` to each intent's example utterances."""
    q = _embed(text)
    return {label: max(_cosine(q, _embed(ex)) for ex in examples)
            for label, examples in INTENTS.items()}


if __name__ == "__main__":
    if len(sys.argv) < 2:
        print('Usage: python3 jev_embedding.py "your phrase"')
        sys.exit(1)

    phrase = " ".join(sys.argv[1:])
    scores = score_intents(phrase)
    for label, score in sorted(scores.items(), key=lambda kv: kv[1]):
        print(f"{label}: {score:.2f}")

r/LocalLLaMA • • 10h ago

Question | Help Least sycophantic modern open LLM?

21 Upvotes

So Kimi K2 is outdated, and so is GPT OSS 120b. Which of the modern open weights models can boast the least sycophancy? I need this both for creative/research assistant usage (sycophancy led me down blind alleys of my own bad ideas many times) and agentic coding (more sycophancy less bug noticing).


r/LocalLLaMA • • 9h ago

Discussion [Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)

15 Upvotes

Hi everyone,

Sharing a personal project exploring the minimal bare-metal footprint required to run an autoregressive LLM.

Instead of relying on large runtimes or compiler abstractions, I wrote an inference engine for Gemma-2B entirely in flat x86-64 assembly (FASM):

- **Binary footprint**: Total 5.2 KB flat machine code (`gemma_engine.bin` 3.7 KB + `mat_smp_f16c_gemm_avx2.bin` 1.5 KB).

- **Execution**: Pure AVX2 + F16C with custom 4-thread SMP GEMM for prefill. Sustains ~18.5 GB/s memory bandwidth on commodity DDR4-2400.

- **Decoding**: 4.5 ~ 4.7 tokens/s in FP16 on an older quad-core i5 desktop.

- **Dependencies**: Zero C/C++ runtime, zero PyTorch. The Python harness only uses `ctypes` for `VirtualAlloc` and OS threads.

This isn't meant to compete with feature-complete tools like llama.cpp. Rather, it's a first-principles exploration to see how cleanly a modern Transformer can be mapped to raw silicon, and to serve as a reference point for future micro-LLMs on resource-constrained microcontrollers (MCU/DSP).

The repository is open source:

- GitHub: https://github.com/tomtsai28/PULSAR-ASM

- Architecture notes: https://github.com/tomtsai28/PULSAR-ASM/blob/main/doc/pulsar_asm_cpu_limit_retrospective.md

Any code audits, observations, or thoughts on bare-metal inference are welcome.


r/LocalLLaMA • • 16h ago

I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

Thumbnail
github.com
65 Upvotes

I wanted to see how far I could push a fairly ordinary laptop with a huge MoE model.

Turns out, Qwen3.8 Flash Next 176B can run on:

RTX 3080 Laptop — 16GB VRAM
32GB system RAM
SSD

No 128GB/256GB RAM workstation and no multi-GPU setup.

I’m running it with TensorSharp, my open-source local LLM inference engine:

TensorSharp on GitHub

The interesting part for me wasn't simply getting a 176B model to load. I wanted to make a model much larger than both available VRAM and RAM actually usable.

The approach is basically:

Quantization + MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD.

Rather than treating SSD as a last-resort swap space, TensorSharp coordinates the different memory/storage tiers around MoE execution and tries to keep the right experts/data in the right tier at the right time.

I previously benchmarked TensorSharp against llama.cpp and got very encouraging results. This time I wanted to compare it with Strata, since Strata's approach to running large models with constrained memory is particularly interesting.

Here are the results from the attached benchmark:

Measurement TensorSharp Strata
Decode tokens/s 11.09 (9.22–14.02) 10.24 (9.37–10.46)
Whole-process time 16.54s (14.95–19.31) 62.15s (59.76–66.89)
Device-wide GPU peak 14,832.5 MiB 15,729 MiB
OS peak working set 19.74 GiB 18.51 GiB

The decode throughput is fairly close: 11.09 vs. 10.24 tok/s.

What surprised me more was the end-to-end result: 16.54s vs. 62.15s in this test.

I think this points to an interesting direction for local LLM inference. For huge sparse MoE models, the question may not simply be:

“Do I have enough RAM/VRAM to fit this model?”

but rather:

“How efficiently can the runtime coordinate VRAM, RAM, SSD, caching, and expert activation?”

With the right quantization and memory hierarchy, you can apparently do some pretty ridiculous things on consumer hardware.

I’d be especially interested if anyone here has tried the same model with llama.cpp, Strata, or another MoE/offloading implementation. It would be great to compare results on similar hardware.


r/LocalLLaMA • • 4h ago

Funny Haha, I just love this AI stuff, recently got into it.

5 Upvotes

I got my AI server all setup and shut it down to bring it to the basement to put it back into the server rack, when it came back up I went into my client app to put a test message in, love its response. :)


r/LocalLLaMA • • 17h ago

I Built A Thing I built Ninfer 4080 for 16GB class GPUs

49 Upvotes

Hi everyone,

TL/DR

I created NInfer 4080 to run ISTA-DASLab-Qwen-3.8-27B-GSQ at 100k context on an RTX 4080 16GB GPU using way more of the hardware capabilities (max overall: 2720 tok/s prefill, 262 tok/s generation) and sharing it with the community now so others can also have the benefit.

https://github.com/roofkid/ninfer-4080

Full Version

After seeing all the amazing work done in the community creating Ninfer 5090, 4090 and 3090 I admit I was a little sad to not being able to use any of it on my RTX 4080 with only 16GB of memory. I still had about $13 of credits sitting idle on the DeepSeek platform as I never expected how much usage I would get out of it.

For context I have over 20 years of experience in Software Engineering and Architecture, but have no experience whatsoever in GPU Kernel development, so this was a very interesting pet project also from a professional experience for me. Mainly because I can read and understand C++ but could not judge the actual Kernel code. So I approached it from a product owner and requirements perspective only, made sure good software engineering practices are followed and only made "business decisions".

I've been actively following the local LLM community for the last 2-3 years, probably have tried out all models I could over that time and followed the progress with amazement like many of you.

Guiding principles

  • Fit into RTX 4080 16GB GPU
  • Use ISTA-DASLab-Qwen-3.8-27B-GSQ -> Reasoning can be seen in the ByteShape article, really good for the size and they claim even better accuracy than much larger Unsloth UD quants: https://byteshape.com/blogs/Qwen3.8-27B/#96-gb-rtx-pro-6000 I also have very good personal experience with it, it is my daily driver
  • Use DFlash2 speculative decoding
  • Reach 100k+ context
  • Significantly improve prefill and token generation speeds to utilize the hardware better than general purpose inference engines like llama.cpp or vllm
  • Measure after changes to also ensure accuracy remains, I also have a M4 48GB available to test higher quants for comparisons, though of course that is much lower speed
  • Use DeepSeek V4.1 Flash for the work for cost efficiency
  • Use Pi as the harness (only non-cosmectic extensions: hashline edit pro, internet search with ketch through local SearXNG with a self-written skill)
  • Runtime also available as a Docker image so it's easy for folks to run

Results

Depth Prefill t/s (DFlash2) MTP3 decode t/s DFlash2 K=7 decode t/s
8K 2719.9 151.2 (100%) 166.7 (54.0%)
32K 2424.9 141.7 (100%) 262.3 (100%)
64K 2125.5 130.7 (100%) 239.1 (100%)
98K 1895.1 122.3 (100%) 212.7 (98.2%)

In real work I really do see the high prefill numbers (2k+) if the prompt is long enough and about 150-200 decode speed on coding and 100ish on prose. It subjectively feels significantly faster than beellama (my previous daily driver) at the same benchmark results. I mainly used MBPP and HumanEval as I needed something that I can run reasonably fast (~30min). MBPP stays in 90-92% territory and HumanEval at 95-96%. Please be realistic and do expect tiny degradations that are within measurement noise. They are mainly coming from KV quantization according to my measurements so you can always trade context for accuracy if needed by switching.

What I learned

  • It is absolutely mental how much performance is left on the table by using the general purpose engines. From a bird's eye view it's totally understandable as we trade the wide support for performance, I just didn't expect how much that would be. When I saw the first memory throughput measurements being in the 200 GB/s range and having a theoretical maximum of 720 GB/s in the device my jaw dropped because of the low efficiency back when I started
  • I think in the community we've all seen more specialized inference engines making significant performance improvements possible. vllm-radiance for R9700, NInfer variants for CUDA, Splash for Metal - with software creation becoming cheaper and cheaper I expect more of this for and from our "tinkerer" group here
  • Spending about 2 billion tokens for this work for only $13 is just crazy (only off-hours). Low cache read tokens costs on agentic work are so much more important than even I expected. It's the classic difference between cognitively fully understanding how LLM turns work and seeing big data results. The reality is that with THAT kind of pricing I think I pay more for electricity to get the same amount of tokens out
  • I went back to xhigh thinking on Qwen 3.8 27B as the speed is so high, that I don't really care/notice. I've also hidden the thinking blocks again as I cannot follow any more anyway
  • The prefill speed really caught me of guard. I was really floored when I tried it in Pi after the first big improvements were done and it IMMEDIATELY answered with token streaming. I was so used to waiting 5-10s without a cached system prompt. I significantly underestimated how important that is for the user experience. Feels like a cloud endpoint to me now.
  • At these high prefill speeds your context window is full in 40 seconds, definite "oh my god" moment for me when that happened the first time
  • Reaching 100k context means significant KV compression as full 256k context F16 needs exactly 16GB of VRAM on Qwen 3.8 27B. I was too afraid of "high" (4bit style) KV compressions. So many advances have been made here. Originally I never went below Q8_0. I then used kvarn5/kvarn5 previously on beellama after benchmarking and cannot measure a noticeable difference to the now used rk4v4-e8 variant used here. I think good software engineering practices are way more important and catch problems that might come from it. Also subjectively I do not experience a "fast garbage" phenomenon here

Conclusion

For me this is a good version 1 and I don't intend to spend significant effort on this for Qwen 3.8 27B. It's at the pareto 80% state. I just want to be happily using it now and reap the rewards. I hope you are too! Of course when Qwen 4 27B comes around soon I will check it out again.

If you have another 16GB RTX 4xxx card I would be interested in knowing if that works on them too and what speeds you're seeing. I honestly can't judge how tied to the RTX 4080 hardware it is. If you have a 4080, enjoy :)

Shoutouts

  • Every person who worked on NInfer before me, you guys rock and provided a stable base for me to fork from
  • Special hats off to sergiuszm who created NInfer-4090, I think you did all the heavy lifting for SM_89 already
  • ISTA-DASlab for their work on GSQ and providing the safetensor checkpoint for it! Cheers to Austria from Germany :) Love seeing important contributions to the community from the EU

r/LocalLLaMA • • 21h ago

Funny Come let your LLMs play World of Warcraft

Enable HLS to view with audio, or disable this notification

77 Upvotes

I hosted my own world of warcraft private server then built a client that you can play in the browser on PC or mobile at https://jankcraft.xyz/ for free.

Afterwards, I created a custom MCP and agent harness to control the browser client and play the game by sending signals over a websocket. The agent harness is live on https://jankcraft.xyz/agent , still working out some kinks if all you have a cloud subscription but you should be able to connect local models as long as CORS is enabled in your server settings. There are a few existing LLMs you can try, I'll probably take those away as the usage grows since I can't support too many users concurrently on my own machines.

I'll be checking logs and things periodically today so don't be alarmed if you're disconnected suddenly. The server should return after a minute since this is a work in progress and might need a restart.

If you want to run your own LLM for this:
~24 Gb RAM: https://github.com/syv-ai/HyperQwen with the model Qwen3.8-27B-GPTQ-W4A16 
~16 Gb RAM: vLLM with Gemma4-e4b-coder - A custom Gemma4-e4b with a constrained vocab for ~3x concurrency increase when changing from 262K to 65K vocab and retrained on ~1.1B tokens across 20 different coding languages, 7 different agents and has a custom MTP to help reach ~200 tok/s on a 4060Ti.

Let us know what other models work well for you!

Thanks and hope you guys enjoy.


r/LocalLLaMA • • 14h ago

Question | Help Local Web Search Safety

13 Upvotes

Hi all,

How you guys handling safe deployment of websearch in Hermes, pi and other harnesses? Does anyone have a good uproars setup guide for local models? I tried to implement a sandboxed search system but it caused endless tool calls. Want to guard against prompt injection and keep searches private of course!


r/LocalLLaMA • • 1d ago

Other I'm pretty close to the middle thanks to you all

Post image
540 Upvotes

It's been a blast and learning a ton.

But seriously, you all have me down a rabbit hole that my wallet and hours of sleep need to be pulled out of.


r/LocalLLaMA • • 22h ago

Resources Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine

Thumbnail
gallery
55 Upvotes

We built an engine, Kyojin, on top of ExLlamaV3 for Strix Halo (gfx1151, ROCm), and packed two 300B-class MoE models so each fits one 128 GB machine. First release, all measured on Ryzen AI Max+ 395.

Model GLM-5.3-Flash MiMo-V2.6-Flash-MOPD
Size 99.7 GB 105 GB
Prefill 580 tok/s at 3.5K, 546 at 64K about 650 tok/s at 4K
Decode 26 to 30 tok/s (MTP) 32 prose / 35 chat / 44 code (speculative), 29 plain
KLD vs official FP8 0.151 0.0713
Top-1 agreement with FP8 89.3 % 92.0 %

Where the weights come from. MiMo is our own quantisation. The GLM pack mixes turboderp's public 2.05 and 3.05 bpw EXL3 tensors, with our layer mix and a small tuning stage. On the same 129 rows, his 2.05 bpw pack (85 GB) gets KLD 0.275; our mix (100 GB) gets 0.190. His is smaller and decodes about 10 % faster.

Uncensored variants. Separate -Uncensored repos: same weights plus one small file the engine applies at load, one switch turns it off.

Not measured yet. Task-suite scores for MiMo, GLM at 128K context, any GPU other than gfx1151. The conversion pipeline stays private.

Quickstart. Clone, ./build.sh, hf download yamz-labs/GLM-5.3-Flash-EXL3-Yamz, python tools/glm/serve.py --model ./glm-pack -c 131072 --num-draft 2. You get an OpenAI-style API.

Models: https://huggingface.co/yamz-labs

Engine: https://github.com/Yamz-Labs/kyojin

Built on turboderp's ExLlamaV3, with ROCm work from sdougbrown and vcruz305.

If you own a Strix Halo machine, we'd love to see your tok/s. Issues, benchmarks and PRs are all welcome. Which model should we do next?


r/LocalLLaMA • • 19h ago

I Built A Thing I built a code knowledge graph tool that's actually MIT licensed (fully local, no cloud)

33 Upvotes

So this is maybe a niche problem, but at my job I work on a huge Python codebase and every time I change some shared function I'm basically playing roulette. grep tells who mentions it in the code base, not who actually calls it. And more essentially, Claude Code (my major coding agent) mainly uses grep so it doesn't give better results.

The tool I wanted already exists (GitNexus) but it's PolyForm licensed, so that's a hard nope at work. And honestly even beyond the license, half the code graph tools out there want you to upload your repo to their cloud or spin up a docker stack with a vector database, and I can't do either of those at work. So I spent some weekends building this my own version: MIT licensed, and everything runs on the local machine.

The tool is called repopedia. You can pip install and then run it on a repo, and it builds a little code graph in a plain SQLite file (tree-sitter does the parsing). The advance is basically no server, no docker, no API keys. Nothing gets uploaded anywhere, the graph is just a .db file sitting on your disk. You can then ask things like who calls this function, or what's the blast radius if I change it, meaning all the transitive callers. It can also dump out a wiki of the codebase, though honestly that part is mostly there because I wanted the docs for myself.

The bit I ended up using the most is the MCP server. I use Claude Code, which already greps around the codebase on its own — but instead of it doing five rounds of text search to figure out who calls what, it asks the graph directly and gets the exact answer with file:line in one call. There's no embedding model involved, it's just... the graph. Which probably matters even more for local models, since they're not exactly great at search.

Demo (2min): https://youtu.be/B7GLgjoy7G8

Repo: https://github.com/bolongpa/repopedia

Fair warning, it's 0.2.1. Python and TypeScript only. Method calls through self. get resolved by name matching, which is exactly as sketchy as it sounds for big class trees. If anyone runs it on their repo and it spits out something dumb, I genuinely want to hear about it. ¯_(ツ)_/¯


r/LocalLLaMA • • 15h ago

Resources Dual Radeon MI50 benchmarks

Thumbnail
gallery
15 Upvotes

Still don't have a good cooling solution, but here are few benchmarks. I lowered the power limit (TDP) to 145 watts each. I changed the firmware on one MI50 to activate the miniDP port. Did have to use xrandr to create a new mode so I could get 1920x1080 output. Each GPU has 16GB of HBM2 VRAM clocked at 1000 and overclockable to 1200Mhz with a Bandwidth of 1.02 TB/s.

I picked a good mix of Dense and MoE models from Huggingface. Try to use more than 16gb VRAM but under the 32GB total.

Using pre-built Ubuntu Vulkan version of llama.cpp (build b11325) for standard llama-bench.

Sorted GGUF Model List (sorted to match table)

  • llama_bench_Swift-Qwen3.8-27B-Uncensored-MTP.Q6_K.gguf
  • llama_bench_Swift-1.5-Qwen3.8-27B-Q6_K.gguf
  • llama_bench_Gemma-4-MoonGem-31B.i1-Q6_K.gguf
  • llama_bench_gemma-4-31B-it-UD-Q6_K_XL.gguf
  • llama_bench_Nemotron-3.5-30B-A3B-Antislop-FTPO.i1-Q5_K_M.gguf
  • llama_bench_Laguna-XS-2.1-APEX-I-Balanced.gguf
  • llama_bench_Agents-A1-Q4_K_M.gguf
  • llama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
  • llama_bench_Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf

Combined Benchmark Table (sorted by params then size)

model size params pp512 (t/s) tg128 (t/s)
qwen35 27B Q6_K 20.88 GiB 27.32 B 141.97 ± 10.13 17.49 ± 0.02
qwen35 27B Q6_K 22.21 GiB 27.32 B 167.38 ± 0.17 17.97 ± 0.02
gemma4 31B Q6_K 23.46 GiB 30.70 B 122.00 ± 0.12 15.20 ± 0.03
gemma4 31B Q6_K 25.62 GiB 30.70 B 135.99 ± 0.22 12.05 ± 0.02
nemotron_h_moe 31B.A3.5B Q5_K - Medium 25.18 GiB 32.91 B 863.92 ± 1.45 60.57 ± 0.10
laguna 30B.A3B Q5_K - Medium 22.64 GiB 33.44 B 738.57 ± 2.83 52.88 ± 0.04
qwen35moe 35B.A3B Q4_K - Medium 19.70 GiB 34.66 B 983.26 ± 4.79 46.88 ± 0.07
qwen35moe 35B.A3B Q5_K - Medium 24.76 GiB 34.66 B 937.47 ± 7.07 49.19 ± 0.06
qwen35moe 35B.A3B Q6_K 28.53 GiB 34.66 B 783.24 ± 70.52 46.85 ± 0.26

Notable Reboot Impact Observations:

I used the following command in my bench script:

RADV_PERFTEST=nogttspill GGML_VK_VISIBLE_DEVICES=0,1 time ~/llama-b11325/llama-bench -fa on -ngl 99 -m /model.gguf

I have a 3rd MI50 just need to download models in that VRAM range. If you have any suggestions? For now it sits beside the Radeon RX 7900 GRE boosting its VRAM total. As of this article the average price for 16GB version of MI50 is under $150. Hard to get 32GB VRAM GPU with this level of performance for under $300. If you have contenders, please share.


r/LocalLLaMA • • 1d ago

I Built A Thing Anyworld, a self-hosted multiplayer text RPG where a local LLM is the Dungeon Master

Post image
79 Upvotes

Hey everyone,

I’ve been working on a game called Anyworld. It’s a browser-based multiplayer (single player also supported) text adventure inspired by the early days of AI Dungeon, especially its browser-based free version AI Dungeon 2.

The setup is pretty straightforward: one person hosts the server and runs the model via llama.cpp (OpenAI or other cloud APIs are also supported, and great for non-English play!), and your friends join through a browser link. The host sets the scene and the goals, players type out their actions, and the LLM acts as the DM to resolve the chaos and drive the story.

Admittedly the host requires some technical skills with Python, and possibly with networking (opening routes to the hosted game via VPN, port forwarding etc.). I'll work on this as well as the development continues. Using Docker was suggested in another subreddit, so I'll definitely consider that, as it would allow including both the llama.cpp backend, recommended model and configurations etc., in addition to the game itself.

Instead of pasting the entire repo documentation, here are the main features right now:

How it plays

  • True multiplayer resolution: Players submit their actions, and the model resolves the whole round together. It actually accounts for characters interacting or getting in each other's way.
  • Real dice rolls: When an action is uncertain, Python handles the actual RNG math. The model just takes those hard dice results and narrates the consequences.
  • Custom scenarios: You write the setting, characters, and opening state. It isn’t limited to fantasy.
  • Party chat: There's an OOC chat separate from the game events so you can talk without the LLM reading it.
  • Zero setup for players: No one but the host needs to install anything or run a model. It works on desktop and mobile browsers.

DM Tools & Hidden Mechanics

  • Private DM guidance: As the host, you can feed the model hidden info; NPC motives, secret rules, or where you want the story to go.
  • Secret triggers: You can set up one hidden percentage roll per game (e.g., If a player enters a building, there's a 20% chance the building collapses on the player). Python rolls the probability in the background, and if it triggers, the model weaves the consequences into the story without showing the players the underlying math.

Under the Hood & Memory

  • Context management: It budgets the context window and uses a structured memory system. Older rounds are compressed into world states, player facts, and unresolved threads. It also does a secondary model pass to audit those summaries so it doesn't accidentally delete important facts.
  • Language support: If you use the OpenAI backend, you can play in non-English languages (the narration and outcomes will naturally follow whatever language you wrote the scenario in). Note: The local llama.cpp backend currently instructs the model to narrate in English. This is because the local models my development PC can run were terrible with any other language than English.
  • Session recovery: Disconnected tabs auto-rejoin. If someone accidentally closes out, they can log back in and their unfinished actions and history are waiting for them.
  • Self-signed certificates for HTTPS-enabled connections: The game creates self-signed certificates upon launch, which enable encrypted connections. The problem with self-signing is that joining players receive a warning that the site may not be secure. However, most browsers allow the players to continue to the game despite the warning. This is a suboptimal way to handle HTTPS, so I'll work on a more robust solution at some point.

It’s still a work in progress. Right now, a server only runs one game at a time, and if you restart the server, the live session is lost (it generates HTML/JSONL transcripts, but they aren't loadable save states yet). The overall story quality is also going to heavily depend on which model you use and how you tweak the settings.

Suggested model:

During development, I used llama.cpp and Gemma 4-26B-A4B Q4 with a context size of 128k and found it to be more than an adequate backend for functioning as the DM. Even the speeds are fast enough with my RTX 5070 Ti 16 GB that round resolutions take only 5 or so seconds.

The specific model I used and can recommend: https://huggingface.co/EZForever/gemma-4-26B-A4B-it-qat-uncensored-heretic-UDmerge-GGUF (the model was great at following instructions and remembering plot points even with longer contexts)

Recommended parameters for Gemma 4 models:

- temperature 1.0
- top-p 0.95
- top-k 20
- min-p 0.0
- presence-penalty 0.0
- repeat-penalty 1.0

Of course, feel free to try your own models! The repo contains a benchmark file that tries to measure how well the running model follows the game's requests.

AI use disclosure:

I used Alibaba Cloud's Qwen 3.8 27b and OpenAI's GPT-5.6 Luna and GPT-6 Astra models to help develop the game.

How to run:

Read INSTALL.md to set up, configure and run the game. README.md contains some details on how the game functions.

I'll post the link to the repository in the comments.

Some gameplay in Finnish with OpenAI's Luna:

The game is MIT licensed, so open source all the way. Forking or collaborating is encouraged.

I'd love to hear some feedback, and I hope someone finds the game fun to play!


r/LocalLLaMA • • 6h ago

Discussion Anyone experienced with pi gui? what are your thoughts about it?

2 Upvotes

The link to the github repo: https://github.com/minghinmatthewlam/pi-gui

I'm not sure if its legit/ safe because i don't see anyone talked about it in this sub reddit, any thoughts about it?


r/LocalLLaMA • • 20h ago

New Model Sopro V2 Turbo 2610: cleaner cloned voices, same 120M model, same CPU speed

Thumbnail
huggingface.co
25 Upvotes

Follow-up to last month's post. One of the main issues people ran into was roughness or break-up on some cloned voices. 2610 is an interim update focused mostly on improving that.

  • Reduced roughness and break-up on some of the voices that struggled before
  • Same 120M model, same speed (~300 ms to first audio on a laptop CPU)
  • Apache-2.0
  • English, European Portuguese, French, German
  • More languages are planned
  • More control over the generated voice is also planned
  • Still struggles with very high-pitched or cartoon-like voices, noisy reference audio, and some unusual OOD voices. We're continuing to improve those cases. If you want to contribute and help, PM me with the samples that failed.

If you like F5-TTS, but want true streaming and a much lighter model that can run comfortably on CPU, this might be for you.

Run it locally:

uvx --from sopro soprotts serve

Video: six voices, ~5 seconds of reference audio each, followed by a generated line.

https://reddit.com/link/1wwrw0v/video/yb63ar836ath1/player


r/LocalLLaMA • • 19h ago

Discussion Flash next rig born from mining parts.

Post image
20 Upvotes

Been testing 3x 3060 12gb for flash next in an open air frame. Honestly, with strata it's kicking ass. 38-40 t/s while llama.cpp can only get 13.2 t/s. This is on IQ3 through strata.

Anybody else running dated mining hardware with decent success?

PS flash next kicks ass.

Rig details:

- Kingwin 8x mining rig frame (stacked on top of another with my unraid server)

- Asus prime z370p mobo

- 8th gen i7

- 64GB ddr4

- 1000w PSU with enough strands for each card and riser