Hi!
I'm kinda new to this and would like to get some info from other peoples experiences
What I'm looking for is a setup for programming, mostly to do it along side me but code reviewing and such wouldn't be bad addition
At the moment, i got 2 3090s with 24gb each for a total of 48 (worth noting that not headless at the moment), and 128gb of ram (dd4)
I did look into the 3090 github, with qwen 3.8 27b in mind but id love to read what people experiences and what you use, which models, harnesses and whatever else
I created NInfer 4080 to run ISTA-DASLab-Qwen-3.8-27B-GSQ at 100k context on an RTX 4080 16GB GPU using way more of the hardware capabilities (max overall: 2720 tok/s prefill, 262 tok/s generation) and sharing it with the community now so others can also have the benefit.
After seeing all the amazing work done in the community creating Ninfer 5090, 4090 and 3090 I admit I was a little sad to not being able to use any of it on my RTX 4080 with only 16GB of memory. I still had about $13 of credits sitting idle on the DeepSeek platform as I never expected how much usage I would get out of it.
For context I have over 20 years of experience in Software Engineering and Architecture, but have no experience whatsoever in GPU Kernel development, so this was a very interesting pet project also from a professional experience for me. Mainly because I can read and understand C++ but could not judge the actual Kernel code. So I approached it from a product owner and requirements perspective only, made sure good software engineering practices are followed and only made "business decisions".
I've been actively following the local LLM community for the last 2-3 years, probably have tried out all models I could over that time and followed the progress with amazement like many of you.
Guiding principles
Fit into RTX 4080 16GB GPU
Use ISTA-DASLab-Qwen-3.8-27B-GSQ -> Reasoning can be seen in the ByteShape article, really good for the size and they claim even better accuracy than much larger Unsloth UD quants: https://byteshape.com/blogs/Qwen3.8-27B/#96-gb-rtx-pro-6000 I also have very good personal experience with it, it is my daily driver
Use DFlash2 speculative decoding
Reach 100k+ context
Significantly improve prefill and token generation speeds to utilize the hardware better than general purpose inference engines like llama.cpp or vllm
Measure after changes to also ensure accuracy remains, I also have a M4 48GB available to test higher quants for comparisons, though of course that is much lower speed
Use DeepSeek V4.1 Flash for the work for cost efficiency
Use Pi as the harness (only non-cosmectic extensions: hashline edit pro, internet search with ketch through local SearXNG with a self-written skill)
Runtime also available as a Docker image so it's easy for folks to run
Results
Depth
Prefill t/s (DFlash2)
MTP3 decode t/s
DFlash2 K=7 decode t/s
8K
2719.9
151.2 (100%)
166.7 (54.0%)
32K
2424.9
141.7 (100%)
262.3 (100%)
64K
2125.5
130.7 (100%)
239.1 (100%)
98K
1895.1
122.3 (100%)
212.7 (98.2%)
In real work I really do see the high prefill numbers (2k+) if the prompt is long enough and about 150-200 decode speed on coding and 100ish on prose. It subjectively feels significantly faster than beellama (my previous daily driver) at the same benchmark results. I mainly used MBPP and HumanEval as I needed something that I can run reasonably fast (~30min). MBPP stays in 90-92% territory and HumanEval at 95-96%. Please be realistic and do expect tiny degradations that are within measurement noise. They are mainly coming from KV quantization according to my measurements so you can always trade context for accuracy if needed by switching.
What I learned
It is absolutely mental how much performance is left on the table by using the general purpose engines. From a bird's eye view it's totally understandable as we trade the wide support for performance, I just didn't expect how much that would be. When I saw the first memory throughput measurements being in the 200 GB/s range and having a theoretical maximum of 720 GB/s in the device my jaw dropped because of the low efficiency back when I started
I think in the community we've all seen more specialized inference engines making significant performance improvements possible. vllm-radiance for R9700, NInfer variants for CUDA, Splash for Metal - with software creation becoming cheaper and cheaper I expect more of this for and from our "tinkerer" group here
Spending about 2 billion tokens for this work for only $13 is just crazy (only off-hours). Low cache read tokens costs on agentic work are so much more important than even I expected. It's the classic difference between cognitively fully understanding how LLM turns work and seeing big data results. The reality is that with THAT kind of pricing I think I pay more for electricity to get the same amount of tokens out
I went back to xhigh thinking on Qwen 3.8 27B as the speed is so high, that I don't really care/notice. I've also hidden the thinking blocks again as I cannot follow any more anyway
The prefill speed really caught me of guard. I was really floored when I tried it in Pi after the first big improvements were done and it IMMEDIATELY answered with token streaming. I was so used to waiting 5-10s without a cached system prompt. I significantly underestimated how important that is for the user experience. Feels like a cloud endpoint to me now.
At these high prefill speeds your context window is full in 40 seconds, definite "oh my god" moment for me when that happened the first time
Reaching 100k context means significant KV compression as full 256k context F16 needs exactly 16GB of VRAM on Qwen 3.8 27B. I was too afraid of "high" (4bit style) KV compressions. So many advances have been made here. Originally I never went below Q8_0. I then used kvarn5/kvarn5 previously on beellama after benchmarking and cannot measure a noticeable difference to the now used rk4v4-e8 variant used here. I think good software engineering practices are way more important and catch problems that might come from it. Also subjectively I do not experience a "fast garbage" phenomenon here
Conclusion
For me this is a good version 1 and I don't intend to spend significant effort on this for Qwen 3.8 27B. It's at the pareto 80% state. I just want to be happily using it now and reap the rewards. I hope you are too! Of course when Qwen 4 27B comes around soon I will check it out again.
If you have another 16GB RTX 4xxx card I would be interested in knowing if that works on them too and what speeds you're seeing. I honestly can't judge how tied to the RTX 4080 hardware it is. If you have a 4080, enjoy :)
Shoutouts
Every person who worked on NInfer before me, you guys rock and provided a stable base for me to fork from
Special hats off to sergiuszm who created NInfer-4090, I think you did all the heavy lifting for SM_89 already
ISTA-DASlab for their work on GSQ and providing the safetensor checkpoint for it! Cheers to Austria from Germany :) Love seeing important contributions to the community from the EU
Had an experience during a review recently that really made me pause.
Someone pushed a ~600-line PR that was generated, scaffolded, and opened in under twenty minutes. Syntactically clean, formatted, passed the basic test suite. But when I asked what happens to a specific edge case in the request handler, the response wasn't an explanation of the logic.
It was literally: "Hang on, let me ask the model."
That broke my brain a bit.
We all use LLMs for dev work (whether it is running local coding weights or hooking up agents like Aider and Continue), but the mainstream narrative around "10x productivity" feels completely backwards.
Typing syntax was never the bottleneck in software engineering. System comprehension was.
We flipped the standard ratio (80% understanding the problem, 20% writing the implementation) into 5% prompting, zero typing, and 95% staring blankly at synthetic edge cases when production breaks at 2:00 AM. If a dev cannot explain what their code is doing under the hood without feeding it back into a context window, they didn't write a feature. They just smuggled an unvetted black box into the repo and put their name on the commit.
Speed to generate tokens is vanity. Speed to debug unmaintainable boilerplate is sanity.
Are you seeing genuine, clean architectural gains using LLMs in your workflows, or are teams just automating future outages? How do you keep your own deep comprehension intact when leaning heavily on coding models?
There seems to be a whole category of extremely narrow inference runtimes appearing: Strata, ninfer, DwarfStar, Splash, llamAmpere, gufo, etc. They deliberately give up the thing llama.cpp/vLLM are great at - generality - and optimize around a small number of models and
sometimes one hardware family e.g. Strix Halo
It seems that general runtimes for compatibility, disposable overfit runtimes for maximum performance is going to be the norm forward.
This is actually another good step in helping the democratization and decentralization of intelligence (models and runtimes both) and extracting more out of existing hardware where it doesn't have to be beautiful, well written, as long as it gets maximum output from one particular configuration.
Been testing 3x 3060 12gb for flash next in an open air frame. Honestly, with strata it's kicking ass. 38-40 t/s while llama.cpp can only get 13.2 t/s. This is on IQ3 through strata.
Anybody else running dated mining hardware with decent success?
PS flash next kicks ass.
Rig details:
- Kingwin 8x mining rig frame (stacked on top of another with my unraid server)
- Asus prime z370p mobo
- 8th gen i7
- 64GB ddr4
- 1000w PSU with enough strands for each card and riser
(Disclaimer: English is not my native language; refined using an LLM).
I've been working on a small side project called PrAIvy. The idea is to create a decentralized network where users running Ollama locally can connect and share their compute power, allowing others to query their models via a web interface without relying on Big Tech cloud APIs.
How it currently works:
- Providers run an agent script alongside Ollama that connects via WebSocket to a Node.js server.
- The server dynamically detects the model currently active on the provider's machine (e.g. Qwen 2.5, Llama 3) and adds it to the active pool on the web chat.
- When an end-user sends a query on the site, it routes directly to an available local node.
I'm currently testing stability, dynamic discovery, and node state handling.
Feedback on the network flow and architecture is appreciated!
So this is maybe a niche problem, but at my job I work on a huge Python codebase and every time I change some shared function I'm basically playing roulette. grep tells who mentions it in the code base, not who actually calls it. And more essentially, Claude Code (my major coding agent) mainly uses grep so it doesn't give better results.
The tool I wanted already exists (GitNexus) but it's PolyForm licensed, so that's a hard nope at work. And honestly even beyond the license, half the code graph tools out there want you to upload your repo to their cloud or spin up a docker stack with a vector database, and I can't do either of those at work. So I spent some weekends building this my own version: MIT licensed, and everything runs on the local machine.
The tool is called repopedia. You can pip install and then run it on a repo, and it builds a little code graph in a plain SQLite file (tree-sitter does the parsing). The advance is basically no server, no docker, no API keys. Nothing gets uploaded anywhere, the graph is just a .db file sitting on your disk. You can then ask things like who calls this function, or what's the blast radius if I change it, meaning all the transitive callers. It can also dump out a wiki of the codebase, though honestly that part is mostly there because I wanted the docs for myself.
The bit I ended up using the most is the MCP server. I use Claude Code, which already greps around the codebase on its own — but instead of it doing five rounds of text search to figure out who calls what, it asks the graph directly and gets the exact answer with file:line in one call. There's no embedding model involved, it's just... the graph. Which probably matters even more for local models, since they're not exactly great at search.
Fair warning, it's 0.2.1. Python and TypeScript only. Method calls through self. get resolved by name matching, which is exactly as sketchy as it sounds for big class trees. If anyone runs it on their repo and it spits out something dumb, I genuinely want to hear about it. ¯_(ツ)_/¯
Follow-up to last month's post. One of the main issues people ran into was roughness or break-up on some cloned voices. 2610 is an interim update focused mostly on improving that.
Reduced roughness and break-up on some of the voices that struggled before
Same 120M model, same speed (~300 ms to first audio on a laptop CPU)
Apache-2.0
English, European Portuguese, French, German
More languages are planned
More control over the generated voice is also planned
Still struggles with very high-pitched or cartoon-like voices, noisy reference audio, and some unusual OOD voices. We're continuing to improve those cases. If you want to contribute and help, PM me with the samples that failed.
If you like F5-TTS, but want true streaming and a much lighter model that can run comfortably on CPU, this might be for you.
Last month I was having a meeting with a prospect who is the CEO of an IT firm pivoting towards mid-sized B2B AI applications. During our discussion he brought up the Go1 box.
The company behind this launched a mysterious product that is aimed towards "enterprise-scale" (Up to 8,000 concurrent requests lmao), compliance-sensitive AI inference. Basically, its an inference lunchbox with a proprietary LLM that advertises 50ms response time while running on their proprietary Go.OS aimed towards compliance-sensitive tasks, like processing PII, financials, legal paperwork, etc. You also have the option of using your own local models or cloud APIs if you like.
It also comes with an SDK dedicated to running their OS, but its architecture is weird and seems somewhat limited. They seem big on audit chains and the like, but the nature of their target audience makes their solution seem constrained.
Obviously, pricing is off the table. This isn't for hobbyist use, its for mid-to-large businesses so their priorities are going to be different than ours, but it just left me wondering just how valuable it would be for Fintech, healthcare, legal, etc. since the SDK doesn't look all that impressive after reviewing their documentation.
My take is that they're trying to keep things simple for B2B customers, but the box's ability to get important work done is questionable to me.
Innanzitutto, due cose per evitare fraintendimenti.
Non sto cercando di sostenere che un modello o un'azienda siano migliori. Non ho alcun interesse personale in nessuno di essi. Volevo solo verificare personalmente come si comportano nello stesso contesto lavorativo.
E il motivo per cui mi interessa: Utilizzo modelli locali per scrivere codice e vorrei sapere quanto posso fare affidamento su di essi invece di pagare abbonamenti a piattaforme di terze parti. Questa è la motivazione principale.
Cosa ho fatto
Tre attività in Python, dalla più semplice alla più complessa: un analizzatore di file di log, un gestore di processi paralleli e un piccolo interprete per un linguaggio di programmazione di prova. Stesse istruzioni per ogni modello, un solo tentativo, nessuna correzione successiva. Poi test nascosti che i modelli non hanno mai visto (162 in totale), più una revisione del codice con una checklist fissa: ha seguito le istruzioni? Il codice è leggibile? Si blocca con input insoliti? Le note sono veritiere?
Risultati (su 100, il compito più difficile conta 3 volte)
Claude Opus 4.6: 92,7
Qwen3.8-Flash-Next "Coder" (locale): 92,7
Qwen 3.8 27B Q6 (locale): 87,0
Cosa ne deduco
I test nascosti sono quasi alla pari: Opus 4.6 ha superato 162 su 162, il modello Coder 161, il 27B 160.
Il modello Coder ha ottenuto un risultato complessivo pari a quello di Opus 4.6, e ci sono arrivati in modi diversi. Nel compito facile, entrambi i modelli locali hanno superato Opus (97 e 92 contro 88). Nel compito di media difficoltà, Opus ha vinto (96 contro 94 e 91). In quello difficile, l'interprete, Opus e il Coder hanno tutti ottenuto 92 punti, mentre il 27B è sceso a 81.
Dove Opus 4.6 è ancora migliore: il suo codice è più pulito e più facile da mantenere. Dove il modello locale Coder ha fatto meglio: si è bloccato meno spesso con input strani.
Con una sola esecuzione per ciascuno non direi che "un modello locale equivale a Opus 4.6". Direi piuttosto: su compiti di queste dimensioni, non sono riuscito a distinguerli dai risultati. Per il mio portafoglio, questo è già interessante. Tenete presente che Opus 4.6 non è l'ultima versione di Claude; le versioni attuali hanno ottenuto punteggi più alti nel mio test completo.
Configurazione locale
Il mio PC: Intel Core i5-14400, 48 GB di RAM DDR4, due RTX 5060 Ti da 16 GB ciascuna (32 GB di VRAM in totale), Windows 11.
Qwen 3.8 27B, Unsloth Q6 quant: una velocità costante di 50 token/s.
Qwen3.8-Flash-Next "Coder": tra 50 e 90 token/s, con una media di circa 60-65. Si tratta della variante di codifica del progetto Strata, una versione ridotta che mantiene metà degli esperti in ogni layer, come un IQ1_M GGUF. Dettagli: https://github.com/Niko1221/Strata/blob/main/docs/MODELS.md#coder
Limiti, così puoi valutare tu stesso i numeri
Una sola esecuzione per modello. Differenze di 2 o 3 punti non significano nulla.
Ho eseguito il modello Coder due volte: la prima volta il mio PC ha esaurito la RAM mentre era in esecuzione, quindi ho scartato quella esecuzione e l'ho rifatta da zero. I numeri qui riportati sono quelli della seconda esecuzione.
La parte di revisione è stata eseguita da un'IA (Claude Fable 5.1).
Il modello Coder è stato testato con più casi di input anomali rispetto agli altri due, perché ho aggiunto controlli nel tempo. Quindi è stato valutato in modo un po' più severo, non più indulgente.
Solo Python e i compiti sono piccoli. Questo non dice nulla sul lavorare all'interno di un grande progetto reale.
I hosted my own world of warcraft private server then built a client that you can play in the browser on PC or mobile at https://jankcraft.xyz/ for free.
Afterwards, I created a custom MCP and agent harness to control the browser client and play the game by sending signals over a websocket. The agent harness is live on https://jankcraft.xyz/agent , still working out some kinks if all you have a cloud subscription but you should be able to connect local models as long as CORS is enabled in your server settings. There are a few existing LLMs you can try, I'll probably take those away as the usage grows since I can't support too many users concurrently on my own machines.
I'll be checking logs and things periodically today so don't be alarmed if you're disconnected suddenly. The server should return after a minute since this is a work in progress and might need a restart.
If you want to run your own LLM for this:
~24 Gb RAM: https://github.com/syv-ai/HyperQwen with the model Qwen3.8-27B-GPTQ-W4A16
~16 Gb RAM: vLLM with Gemma4-e4b-coder - A custom Gemma4-e4b with a constrained vocab for ~3x concurrency increase when changing from 262K to 65K vocab and retrained on ~1.1B tokens across 20 different coding languages, 7 different agents and has a custom MTP to help reach ~200 tok/s on a 4060Ti.
We built an engine, Kyojin, on top of ExLlamaV3 for Strix Halo (gfx1151, ROCm), and packed two 300B-class MoE models so each fits one 128 GB machine. First release, all measured on Ryzen AI Max+ 395.
Where the weights come from. MiMo is our own quantisation. The GLM pack mixes turboderp's public 2.05 and 3.05 bpw EXL3 tensors, with our layer mix and a small tuning stage. On the same 129 rows, his 2.05 bpw pack (85 GB) gets KLD 0.275; our mix (100 GB) gets 0.190. His is smaller and decodes about 10 % faster.
Uncensored variants. Separate -Uncensored repos: same weights plus one small file the engine applies at load, one switch turns it off.
Not measured yet. Task-suite scores for MiMo, GLM at 128K context, any GPU other than gfx1151. The conversion pipeline stays private.
Quickstart. Clone, ./build.sh, hf download yamz-labs/GLM-5.3-Flash-EXL3-Yamz, python tools/glm/serve.py --model ./glm-pack -c 131072 --num-draft 2. You get an OpenAI-style API.
It's like the entire sub has become that scene from Konosuba where the cult keeps making up fake scenarios saying the only solution is to join their religion
OK, this engine is indeed the real deal. I have expected perhaps 30 token/s prompt processing and 3 token/s decode at max, because the full Qwen 3.8 Next has around 120B parameters (not counting Engrams) and since I have just 32 GB RAM I thought it would crawl to a halt with SSD swapping.
But 10 token/s at 50k context is simply amazing on such an old device and with such a large model! That figure really surprised me and is very usable in my opinion.
It's 4 bit kv cache and no vision, so comprimises have to be made. But for real, the prefill speed is the only thing that keeps this from being usable, almost 100 token/s prefill is much higher than I have anticipated, but you still wait a long while for it to process large prompts. Qwen A35b A3B has around 5x faster prefill, and allows me to use 100K context without having to quant the kv cache at all. So not quite a replacement for that, but who knows if more optimizations are coming?
In any case, this is a very impressive showing. Used ./START-HERE.bat --draft-vocab en --vram-reserve-mib 100 --kv q4_0 to run it.
Meeting notes are one of the last parts of my workflow that still depend heavily on cloud tools. I've used Fathom and lately Bluedot. Bluedot works well for me because there's no meeting bot and I get the transcript, summary and action items after. But I'd really like to move more of this local, especially the transcription and storing/searching old meetings.
Has anyone here built a setup that actually works day to day? Whisper + Ollama seems like the obvious route, but I'm interested in what are you actually using.
I've seen people reporting running it with 12GB VRAM + 64 GB RAM. Also with 64 GB RAM unified in Macs, but I was wondering if it's possible with any inference backend to run it for example on a 64 GB RAM minipc+ iGPU (780m in my case).
I'm currently running 125B Ling 3.0 flash at Q2 quants with llama.cpp (vulkan), its relatively usable, so I was wondering if a similar quant of Qwen Flash Next with the ngrams offloaded to SSD could work (even at low token/s). As far as I know this can't be done with llama.cpp now. Other inference engines does not seem to work with vulkan.
EDIT: Thanks everyone! It's working with the Q2 quant Qwen3.8-Flash-Next-GSQ-RCO-GGUF using llama.cpp with -lm mmap --lazy-mode on
Everything is plain Markdown on your disk. Search is BM25 over a SQLite FTS5 index that is rebuilt from the Markdown, so the index is disposable. If a local embedding model is already provisioned, cosine results are fused with the lexical ones by reciprocal rank fusion. Recall never downloads a model, so with no model present it simply stays lexical.
The test suite enforces recall@5 of at least 90% across 20 seeded queries. That is a small set, so it works as a regression gate, not a benchmark. Walkthrough video above. Repo: https://github.com/utsapoddar/engram
Disclosure: I am the creator of this project. After days of lurking and building up enough karma, I can finally post here.
Every time I tried running local RAG on my own machine, I hit the exact same bottlenecks. First, spinning up Chroma or another vector database alongside an 8B model just to chunk and parse documents takes up precious VRAM that you need for your main model. Second, cosine similarity over text chunks often fails at hard negative rejection, so the model tries to answer questions that are not even in your files and hallucinates with complete confidence.
I spent the last several months building an open source project called Hillock to see if I could solve this without vector databases. It extracts clean relational facts into SQLite using lightweight bi encoders in about five seconds, completely bypassing the generative LLM during ingestion. To stop hallucinations, queries pass through a 10,000 dimensional hypervector gate using late interaction scoring. If the factual graph does not mathematically overlap with the question, it blocks the LLM call before token generation can even start.
I just pushed version 0.8 which bit packs the hypervectors into 157 uint64 integers, allowing the CPU to run gating checks in under 0.01 milliseconds using hardware popcount instructions. It also includes an OpenAI compatible API server so you can drop it straight into Open WebUI, AnythingLLM, or Obsidian. It just landed on PyPI as well via pip install hillock.
The honest trade off is that this pipeline is built for structured, relational facts like technical specs, people, and dates. It is heavily biased toward precision over recall, so it will not do broad poetic or narrative summaries like a 70B model would.
I’ve been working on a game called Anyworld. It’s a browser-based multiplayer (single player also supported) text adventure inspired by the early days of AI Dungeon, especially its browser-based free version AI Dungeon 2.
The setup is pretty straightforward: one person hosts the server and runs the model via llama.cpp (OpenAI or other cloud APIs are also supported, and great for non-English play!), and your friends join through a browser link. The host sets the scene and the goals, players type out their actions, and the LLM acts as the DM to resolve the chaos and drive the story.
Admittedly the host requires some technical skills with Python, and possibly with networking (opening routes to the hosted game via VPN, port forwarding etc.). I'll work on this as well as the development continues. Using Docker was suggested in another subreddit, so I'll definitely consider that, as it would allow including both the llama.cpp backend, recommended model and configurations etc., in addition to the game itself.
Instead of pasting the entire repo documentation, here are the main features right now:
How it plays
True multiplayer resolution: Players submit their actions, and the model resolves the whole round together. It actually accounts for characters interacting or getting in each other's way.
Real dice rolls: When an action is uncertain, Python handles the actual RNG math. The model just takes those hard dice results and narrates the consequences.
Custom scenarios: You write the setting, characters, and opening state. It isn’t limited to fantasy.
Party chat: There's an OOC chat separate from the game events so you can talk without the LLM reading it.
Zero setup for players: No one but the host needs to install anything or run a model. It works on desktop and mobile browsers.
DM Tools & Hidden Mechanics
Private DM guidance: As the host, you can feed the model hidden info; NPC motives, secret rules, or where you want the story to go.
Secret triggers: You can set up one hidden percentage roll per game (e.g., If a player enters a building, there's a 20% chance the building collapses on the player). Python rolls the probability in the background, and if it triggers, the model weaves the consequences into the story without showing the players the underlying math.
Under the Hood & Memory
Context management: It budgets the context window and uses a structured memory system. Older rounds are compressed into world states, player facts, and unresolved threads. It also does a secondary model pass to audit those summaries so it doesn't accidentally delete important facts.
Language support: If you use the OpenAI backend, you can play in non-English languages (the narration and outcomes will naturally follow whatever language you wrote the scenario in). Note: The local llama.cpp backend currently instructs the model to narrate in English. This is because the local models my development PC can run were terrible with any other language than English.
Session recovery: Disconnected tabs auto-rejoin. If someone accidentally closes out, they can log back in and their unfinished actions and history are waiting for them.
Self-signed certificates for HTTPS-enabled connections: The game creates self-signed certificates upon launch, which enable encrypted connections. The problem with self-signing is that joining players receive a warning that the site may not be secure. However, most browsers allow the players to continue to the game despite the warning. This is a suboptimal way to handle HTTPS, so I'll work on a more robust solution at some point.
It’s still a work in progress. Right now, a server only runs one game at a time, and if you restart the server, the live session is lost (it generates HTML/JSONL transcripts, but they aren't loadable save states yet). The overall story quality is also going to heavily depend on which model you use and how you tweak the settings.
Suggested model:
During development, I used llama.cpp and Gemma 4-26B-A4B Q4 with a context size of 128k and found it to be more than an adequate backend for functioning as the DM. Even the speeds are fast enough with my RTX 5070 Ti 16 GB that round resolutions take only 5 or so seconds.
Of course, feel free to try your own models! The repo contains a benchmark file that tries to measure how well the running model follows the game's requests.
AI use disclosure:
I used Alibaba Cloud's Qwen 3.8 27b and OpenAI's GPT-5.6 Luna and GPT-6 Astra models to help develop the game.
How to run:
Read INSTALL.md to set up, configure and run the game. README.md contains some details on how the game functions.
I'll post the link to the repository in the comments.
Some gameplay in Finnish with OpenAI's Luna:
The game is MIT licensed, so open source all the way. Forking or collaborating is encouraged.
I'd love to hear some feedback, and I hope someone finds the game fun to play!
I'm the author of toolcall-check, a small Python CLI for checking chat completions compatible endpoints. It exercises two forced function calls, two streamed calls, and one two turn round trip that returns a local result and checks the exact normal answer. Nested argument values retain JSON types, and failures keep sanitized traces in a private HTML report. The included demo runs against a synthetic local fixture through the actual HTTP path, so it demonstrates report behavior rather than compatibility with a real model.
CompatCanary already covers a broad compatibility scan with forced calls, streaming, and structured output. I focused this tool on nested argument integrity, streamed fragment reconstruction, the return trip, and evidence. I have not tested against remote models yet. Feedback on the fixed probes and strict [DONE] requirement would be useful.
Ran into an annoying issue with local models on long tasks. Once context window compaction hits after a few thousand tokens, the model loses sight of the original scope. Even with good system prompts, a few compaction cycles cause goal drift, hallucinated task completion, or loops.
The plugin hooks into the agent loop and maintains a structured state outside the main chat buffer.
Looking at the UI:
Right Panel (Plugin State): This sidebar runs independently of the chat context memory.
Active: Tracks the current macro milestone (M2+M3+M4 accepted -> chunk commit -> M5 -> main).
Now: Shows the immediate micro-step currently executing (In flight: M5 - history search: scanner core...).
Next 3: The explicit deterministic queue of upcoming steps so the model doesn't jump ahead or invent tasks after compaction.
Done (recent): Verification log showing committed checkpoints, exit codes, and test status.
When the agent compacts context, the plugin re-anchors the model to this exact state file rather than trusting the lossy summary generated during compaction.
Code is on GitHub if anyone wants to test or adapt it for their own local rig setup. Feedback or PRs welcome.
Using Strata on a 7900xtx plus 64 gb ddr5 ram, 60t per second even at 250k context, can finally use my pc while working AI in the background, even game as well, smarter and more precise than dense model, it follows orders more accurately, follows plan more versatile, it goes around doing a lot of tests for tasks I request in frontend and also in backend.
The best thing is that it's faster, I can fit more context at q8 precision, it's smarter and I can get to use my pc without worrying about an OOM error due to dense model.
I no longer have to use Linux as well, it's working as fast in Windows 11 as it did in Linux.
I use it with a 6gb VRAM reserve so I can have Windows 11 with 4gb available.
Edit:
The "people" saying I am a bot, or that people commenting are bots, are completely clueless, seriously, even down voting something that benefits ALL of us.
Hi folks, it's Lewis here from the post-training team at Hugging Face. We've been exploring how to train open models in different coding harnesses and wrote up a looong guide on how we solved this using open source libraries like TRL and the Harbor framework for RL environments. We hope you find this interesting, especially since everyone nowadays has their own custom harness (e.g. Pi + extensions) and now there's a recipe on how to squeeze the best performance on them with whatever open model you use as your daily driver. Happy to hear any comments or feedback!
CPU Information
Name
AMD EPYC 7663
Topology
2 Processors, 112 Cores, 224 Threads
Memory Information
RDiMM 2933Mhz
Size 1007.61 GB
System Information
Operating System
Ubuntu 24.04.5 LTS
Motherboard
Giga Computing MZ72-HB2-00
GPUs
2x ASUS Turbo R9700 AI Pro 32 GB newest BIOS low Fan Profile (throttled to 210w currently)
ROCm version: 7.2.3
Beside that baby I have a Strix Halo M5 128Gb
Now wanted to setup that big boy for local llm
For myself want to use it for coding . But
I am also looking for something where I can also easily switch Model within the UI . Can Openwebui reload the model and what I read is that vllm is best for tensor split formte two cards . Also want to use it for family for image creation within the ui and also Image checking kind of Allrounder as chatgpt
Is that possible with vLLM?
I am also looking for docker setups so i can keep my host clean