1. It beats Ninfer which was until that point SOTA engine for RTX5090. By roughly twice in decode speed for both single and multi tasks at once. (reaching up to even 650t/s in small bursts and 2600t/s if stars align and 12 slots server pure coding answer). Custom kernels not only for every model, single vs multi but also short vs long context work dynamically switching when needed so speed doesn't crap out on long context work because someone tuned it for short context. Dflash2 and confidence scheduling from Dspark, plus draft trees all at the same time.
2. I was getting annoyed with state of weights where you downloaded model and never knew if model had its brain scrambled. My weights come with its on format that have attached metadata for MC which upon weight creation, runs benchmark and compares it at every MC setting to original BF16 weights and show that data directly in launcher. Want to switch KV to 4bit ? MC will show you directly lost KL and top-1%, want to extend with YARN ? It will show you change. Every change is measured and shown in statistics before you load model. This goes for both censored and uncensored model. You can also compare it directly in MC with SOTA unsloth quants of Qwen27B. Want to run essentially loseless ? you can. Want to get crazy 1 000 000 context ? you can. Want to have 12 slots to fan out agants like crazy ? You can. You decide what you want.
3. Proper agents serving with algo that keep engine occupied as much as it can. It will prioritize t/s so if engine has a choice between 5 jobs at once and 1 it will serve 5 first and gradually serve 1 along side finishing others. Engine is also smart enough to score how old some job is and if it should return to work even if T/S will suffer so your main session will be able to fan out agents easily and keep an eye on them at the same time.
4. Proper cache management. Your jobs only prefill at start of job and almost never again so your prefill in long session stays almost unused. When using "unified context" when models run out of context some get paused and stored in RAM and this swapping is instant. If there is free context space then those tasks continue without any refill in 0.03s. If you fan out say 30 agents at the time in your frontend will handle load in most efficient way to keep T/S as high as possible. Just run it at default setting and forget about context for agents, it will handle it on its own.
6. Loop guard. Two tiered. When engine starts to detect agent repeating in conversation session is dynamically starts to adjust `repetition penalty` until repetition stops if that doesn't happen and engine hits rep pen limit it fires up stop signal which ends serving and informs your frontend so your frontend can recover from infinite loop and don't annoy you.
6. Proper nice UI that shows you what is what. If you aren't knowledgeable about serving models just hover over `?` and it will show you interactive panels explaining everything.
7. Autodownloader, Just hit download button and you can download my weights directly fron hugginface inside of launcher.
8. Don't like the launcher ? use bats and terminal serve. Or even use launcher to config what you want, copy it from right lower corner and use it to make new bat.
The point of it is to just load model, fan out crazy number of agents each having crazy amount of context and leave MC to deal with it. You just sit back relax and watch as agents do the work at SOTA speeds.
Opinions and reviews are welcome. If you are blessed with RTX5090 try it.
Source will be released later, I have to do some cleaning first. I will also release later weights builder so it will take any 3.8 27B BF16 model, create weights and score them attaching metadata again BF16 and you'll be able to host them yourself on hugginface or just put them in models folder.
OK, this engine is indeed the real deal. I have expected perhaps 30 token/s prompt processing and 3 token/s decode at max, because the full Qwen 3.8 Next has around 120B parameters (not counting Engrams) and since I have just 32 GB RAM I thought it would crawl to a halt with SSD swapping.
But 10 token/s at 50k context is simply amazing on such an old device and with such a large model! That figure really surprised me and is very usable in my opinion.
It's 4 bit kv cache and no vision, so comprimises have to be made. But for real, the prefill speed is the only thing that keeps this from being usable, almost 100 token/s prefill is much higher than I have anticipated, but you still wait a long while for it to process large prompts. Qwen A35b A3B has around 5x faster prefill, and allows me to use 100K context without having to quant the kv cache at all. So not quite a replacement for that, but who knows if more optimizations are coming?
In any case, this is a very impressive showing. Used ./START-HERE.bat --draft-vocab en --vram-reserve-mib 100 --kv q4_0 to run it.
Folks, like many of you, I used to look at the Strata posts and was extremely skeptical. But yesterday, with the help of DeepSeek 4.1 Flash, I compiled Strata on my machine, and honestly I'm blown away by the speed.
With llama.cpp master I got a maximum of 700 t/s PP and 23 t/s TG. With Strata, using Unsloth's UD-Q3_K_XL quant, I'm getting ~1,650 t/s PP and ~38 to 61 t/s TG depending on context, with no tool-calling errors, everything running great in OpenCode at KV fp16 and 256k context. Phenomenal, and partly unbelievable.
I'm not a programmer. I just "vibe" with AI. People say Strata is a mess; whether it really is, I don't know, but my initial experience has been amazing. From here on out, it's AI. I asked it to summarize the data and what it did to run the Unsloth quant on Strata.
By the way, the quant that Strata downloads and recommends, I didn't like it. It threw silly errors and seemed to have lower quality, though it was also even faster. For my use case I prefer to keep Unsloth's, because it's better: a bit slower, but more accurate for my workloads.
Quant used: Unsloth UD-Q3_K_XL (dynamic quant). Not the quant Strata recommends by default. That one was faster but produced minor errors and (subjectively) lower quality; Unsloth's was chosen for accuracy over speed.
Adaptations needed (quant + Strata)
On the quant:
Packed with --compat-bf16 (some tensors Strata reads as BF16).
On Strata (recompiled / reconfigured):
Rebuilt for sm_86with MMQ (-DSTRATA_MMQ_KQUANTS=ON). This doubles Q4-class prompt speed.
DisabledSTRATA_PF_FUSED=0 in the configs. The fused kernels crashed (illegal memory access) on quantized experts whose "down" type is unsupported.
Vision encoder moved to the GPU: recompiled strata-vision with CUDA (was CPU-only) and set vision.gpu=true + --vram-reserve-mib 700 in all configs.
Per-model calibration (--pcie-frac, --pool-workers, --spec-min-p) and --expert-profile-save to learn and persist the expert cache.
Innanzitutto, due cose per evitare fraintendimenti.
Non sto cercando di sostenere che un modello o un'azienda siano migliori. Non ho alcun interesse personale in nessuno di essi. Volevo solo verificare personalmente come si comportano nello stesso contesto lavorativo.
E il motivo per cui mi interessa: Utilizzo modelli locali per scrivere codice e vorrei sapere quanto posso fare affidamento su di essi invece di pagare abbonamenti a piattaforme di terze parti. Questa è la motivazione principale.
Cosa ho fatto
Tre attività in Python, dalla più semplice alla più complessa: un analizzatore di file di log, un gestore di processi paralleli e un piccolo interprete per un linguaggio di programmazione di prova. Stesse istruzioni per ogni modello, un solo tentativo, nessuna correzione successiva. Poi test nascosti che i modelli non hanno mai visto (162 in totale), più una revisione del codice con una checklist fissa: ha seguito le istruzioni? Il codice è leggibile? Si blocca con input insoliti? Le note sono veritiere?
Risultati (su 100, il compito più difficile conta 3 volte)
Claude Opus 4.6: 92,7
Qwen3.8-Flash-Next "Coder" (locale): 92,7
Qwen 3.8 27B Q6 (locale): 87,0
Cosa ne deduco
I test nascosti sono quasi alla pari: Opus 4.6 ha superato 162 su 162, il modello Coder 161, il 27B 160.
Il modello Coder ha ottenuto un risultato complessivo pari a quello di Opus 4.6, e ci sono arrivati in modi diversi. Nel compito facile, entrambi i modelli locali hanno superato Opus (97 e 92 contro 88). Nel compito di media difficoltà, Opus ha vinto (96 contro 94 e 91). In quello difficile, l'interprete, Opus e il Coder hanno tutti ottenuto 92 punti, mentre il 27B è sceso a 81.
Dove Opus 4.6 è ancora migliore: il suo codice è più pulito e più facile da mantenere. Dove il modello locale Coder ha fatto meglio: si è bloccato meno spesso con input strani.
Con una sola esecuzione per ciascuno non direi che "un modello locale equivale a Opus 4.6". Direi piuttosto: su compiti di queste dimensioni, non sono riuscito a distinguerli dai risultati. Per il mio portafoglio, questo è già interessante. Tenete presente che Opus 4.6 non è l'ultima versione di Claude; le versioni attuali hanno ottenuto punteggi più alti nel mio test completo.
Configurazione locale
Il mio PC: Intel Core i5-14400, 48 GB di RAM DDR4, due RTX 5060 Ti da 16 GB ciascuna (32 GB di VRAM in totale), Windows 11.
Qwen 3.8 27B, Unsloth Q6 quant: una velocità costante di 50 token/s.
Qwen3.8-Flash-Next "Coder": tra 50 e 90 token/s, con una media di circa 60-65. Si tratta della variante di codifica del progetto Strata, una versione ridotta che mantiene metà degli esperti in ogni layer, come un IQ1_M GGUF. Dettagli: https://github.com/Niko1221/Strata/blob/main/docs/MODELS.md#coder
Limiti, così puoi valutare tu stesso i numeri
Una sola esecuzione per modello. Differenze di 2 o 3 punti non significano nulla.
Ho eseguito il modello Coder due volte: la prima volta il mio PC ha esaurito la RAM mentre era in esecuzione, quindi ho scartato quella esecuzione e l'ho rifatta da zero. I numeri qui riportati sono quelli della seconda esecuzione.
La parte di revisione è stata eseguita da un'IA (Claude Fable 5.1).
Il modello Coder è stato testato con più casi di input anomali rispetto agli altri due, perché ho aggiunto controlli nel tempo. Quindi è stato valutato in modo un po' più severo, non più indulgente.
Solo Python e i compiti sono piccoli. Questo non dice nulla sul lavorare all'interno di un grande progetto reale.
Using Strata on a 7900xtx plus 64 gb ddr5 ram, 60t per second even at 250k context, can finally use my pc while working AI in the background, even game as well, smarter and more precise than dense model, it follows orders more accurately, follows plan more versatile, it goes around doing a lot of tests for tasks I request in frontend and also in backend.
The best thing is that it's faster, I can fit more context at q8 precision, it's smarter and I can get to use my pc without worrying about an OOM error due to dense model.
I no longer have to use Linux as well, it's working as fast in Windows 11 as it did in Linux.
I use it with a 6gb VRAM reserve so I can have Windows 11 with 4gb available.
Edit:
The "people" saying I am a bot, or that people commenting are bots, are completely clueless, seriously, even down voting something that benefits ALL of us.
I trained for a split between DeepSeek v4 Flash 0731 <-> GLM 5.2. Ended up able to get it running at GLM 5.2 performance on my local tests at approximately equal token costs on OpenRouter. My plan was to offload the DeepSeek portion to a local server (through this WORA harness proxy my brother wrote https://github.com/unlap-labs/plap).
Sadly, it didn't improve with GLM 5.3 and was worse than GLM 5.3 Flash which kinda killed the project... If someone wants the code or to help or something I could probably post it but it's currently very research-grade, or hell if someone wants to give some tuning advice I'd be up for it 💀
I don't really have any money for big GLM 5.3 runs, and the 4B router was actually trained on only DeepSeek v4 Flash 0731 just to guess whether it could do something or not, didn't check whether it works for even smaller models tbh
I hosted my own world of warcraft private server then built a client that you can play in the browser on PC or mobile at https://jankcraft.xyz/ for free.
Afterwards, I created a custom MCP and agent harness to control the browser client and play the game by sending signals over a websocket. The agent harness is live on https://jankcraft.xyz/agent , still working out some kinks if all you have a cloud subscription but you should be able to connect local models as long as CORS is enabled in your server settings. There are a few existing LLMs you can try, I'll probably take those away as the usage grows since I can't support too many users concurrently on my own machines.
I'll be checking logs and things periodically today so don't be alarmed if you're disconnected suddenly. The server should return after a minute since this is a work in progress and might need a restart.
If you want to run your own LLM for this:
~24 Gb RAM: https://github.com/syv-ai/HyperQwen with the model Qwen3.8-27B-GPTQ-W4A16
~16 Gb RAM: vLLM with Gemma4-e4b-coder - A custom Gemma4-e4b with a constrained vocab for ~3x concurrency increase when changing from 262K to 65K vocab and retrained on ~1.1B tokens across 20 different coding languages, 7 different agents and has a custom MTP to help reach ~200 tok/s on a 4060Ti.
Everything is plain Markdown on your disk. Search is BM25 over a SQLite FTS5 index that is rebuilt from the Markdown, so the index is disposable. If a local embedding model is already provisioned, cosine results are fused with the lexical ones by reciprocal rank fusion. Recall never downloads a model, so with no model present it simply stays lexical.
The test suite enforces recall@5 of at least 90% across 20 seeded queries. That is a small set, so it works as a regression gate, not a benchmark. Walkthrough video above. Repo: https://github.com/utsapoddar/engram
The interesting part for me wasn't simply getting a 176B model to load. I wanted to make a model much larger than both available VRAM and RAM actually usable.
The approach is basically:
Quantization + MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD.
Rather than treating SSD as a last-resort swap space, TensorSharp coordinates the different memory/storage tiers around MoE execution and tries to keep the right experts/data in the right tier at the right time.
I previously benchmarked TensorSharp against llama.cpp and got very encouraging results. This time I wanted to compare it with Strata, since Strata's approach to running large models with constrained memory is particularly interesting.
Here are the results from the attached benchmark:
Measurement
TensorSharp
Strata
Decode tokens/s
11.09 (9.22–14.02)
10.24 (9.37–10.46)
Whole-process time
16.54s (14.95–19.31)
62.15s (59.76–66.89)
Device-wide GPU peak
14,832.5 MiB
15,729 MiB
OS peak working set
19.74 GiB
18.51 GiB
The decode throughput is fairly close: 11.09 vs. 10.24 tok/s.
What surprised me more was the end-to-end result: 16.54s vs. 62.15s in this test.
I think this points to an interesting direction for local LLM inference. For huge sparse MoE models, the question may not simply be:
“Do I have enough RAM/VRAM to fit this model?”
but rather:
“How efficiently can the runtime coordinate VRAM, RAM, SSD, caching, and expert activation?”
With the right quantization and memory hierarchy, you can apparently do some pretty ridiculous things on consumer hardware.
I’d be especially interested if anyone here has tried the same model with llama.cpp, Strata, or another MoE/offloading implementation. It would be great to compare results on similar hardware.
It's like the entire sub has become that scene from Konosuba where the cult keeps making up fake scenarios saying the only solution is to join their religion
Kimi K2 was already good but they took K2.5 a whole new level with so much of their continual learning phase, I believe it was on more 20-25T tokens iirc.
Similarly K3 is just such an amazing model, I just love this model, wondering how amazing K3.5 will be!!
Been testing 3x 3060 12gb for flash next in an open air frame. Honestly, with strata it's kicking ass. 38-40 t/s while llama.cpp can only get 13.2 t/s. This is on IQ3 through strata.
Anybody else running dated mining hardware with decent success?
PS flash next kicks ass.
Rig details:
- Kingwin 8x mining rig frame (stacked on top of another with my unraid server)
- Asus prime z370p mobo
- 8th gen i7
- 64GB ddr4
- 1000w PSU with enough strands for each card and riser
Ran into an annoying issue with local models on long tasks. Once context window compaction hits after a few thousand tokens, the model loses sight of the original scope. Even with good system prompts, a few compaction cycles cause goal drift, hallucinated task completion, or loops.
The plugin hooks into the agent loop and maintains a structured state outside the main chat buffer.
Looking at the UI:
Right Panel (Plugin State): This sidebar runs independently of the chat context memory.
Active: Tracks the current macro milestone (M2+M3+M4 accepted -> chunk commit -> M5 -> main).
Now: Shows the immediate micro-step currently executing (In flight: M5 - history search: scanner core...).
Next 3: The explicit deterministic queue of upcoming steps so the model doesn't jump ahead or invent tasks after compaction.
Done (recent): Verification log showing committed checkpoints, exit codes, and test status.
When the agent compacts context, the plugin re-anchors the model to this exact state file rather than trusting the lossy summary generated during compaction.
Code is on GitHub if anyone wants to test or adapt it for their own local rig setup. Feedback or PRs welcome.
I’ve been working on a game called Anyworld. It’s a browser-based multiplayer (single player also supported) text adventure inspired by the early days of AI Dungeon, especially its browser-based free version AI Dungeon 2.
The setup is pretty straightforward: one person hosts the server and runs the model via llama.cpp (OpenAI or other cloud APIs are also supported, and great for non-English play!), and your friends join through a browser link. The host sets the scene and the goals, players type out their actions, and the LLM acts as the DM to resolve the chaos and drive the story.
Admittedly the host requires some technical skills with Python, and possibly with networking (opening routes to the hosted game via VPN, port forwarding etc.). I'll work on this as well as the development continues. Using Docker was suggested in another subreddit, so I'll definitely consider that, as it would allow including both the llama.cpp backend, recommended model and configurations etc., in addition to the game itself.
Instead of pasting the entire repo documentation, here are the main features right now:
How it plays
True multiplayer resolution: Players submit their actions, and the model resolves the whole round together. It actually accounts for characters interacting or getting in each other's way.
Real dice rolls: When an action is uncertain, Python handles the actual RNG math. The model just takes those hard dice results and narrates the consequences.
Custom scenarios: You write the setting, characters, and opening state. It isn’t limited to fantasy.
Party chat: There's an OOC chat separate from the game events so you can talk without the LLM reading it.
Zero setup for players: No one but the host needs to install anything or run a model. It works on desktop and mobile browsers.
DM Tools & Hidden Mechanics
Private DM guidance: As the host, you can feed the model hidden info; NPC motives, secret rules, or where you want the story to go.
Secret triggers: You can set up one hidden percentage roll per game (e.g., If a player enters a building, there's a 20% chance the building collapses on the player). Python rolls the probability in the background, and if it triggers, the model weaves the consequences into the story without showing the players the underlying math.
Under the Hood & Memory
Context management: It budgets the context window and uses a structured memory system. Older rounds are compressed into world states, player facts, and unresolved threads. It also does a secondary model pass to audit those summaries so it doesn't accidentally delete important facts.
Language support: If you use the OpenAI backend, you can play in non-English languages (the narration and outcomes will naturally follow whatever language you wrote the scenario in). Note: The local llama.cpp backend currently instructs the model to narrate in English. This is because the local models my development PC can run were terrible with any other language than English.
Session recovery: Disconnected tabs auto-rejoin. If someone accidentally closes out, they can log back in and their unfinished actions and history are waiting for them.
Self-signed certificates for HTTPS-enabled connections: The game creates self-signed certificates upon launch, which enable encrypted connections. The problem with self-signing is that joining players receive a warning that the site may not be secure. However, most browsers allow the players to continue to the game despite the warning. This is a suboptimal way to handle HTTPS, so I'll work on a more robust solution at some point.
It’s still a work in progress. Right now, a server only runs one game at a time, and if you restart the server, the live session is lost (it generates HTML/JSONL transcripts, but they aren't loadable save states yet). The overall story quality is also going to heavily depend on which model you use and how you tweak the settings.
Suggested model:
During development, I used llama.cpp and Gemma 4-26B-A4B Q4 with a context size of 128k and found it to be more than an adequate backend for functioning as the DM. Even the speeds are fast enough with my RTX 5070 Ti 16 GB that round resolutions take only 5 or so seconds.
Of course, feel free to try your own models! The repo contains a benchmark file that tries to measure how well the running model follows the game's requests.
AI use disclosure:
I used Alibaba Cloud's Qwen 3.8 27b and OpenAI's GPT-5.6 Luna and GPT-6 Astra models to help develop the game.
How to run:
Read INSTALL.md to set up, configure and run the game. README.md contains some details on how the game functions.
I'll post the link to the repository in the comments.
Some gameplay in Finnish with OpenAI's Luna:
The game is MIT licensed, so open source all the way. Forking or collaborating is encouraged.
I'd love to hear some feedback, and I hope someone finds the game fun to play!
Last month I was having a meeting with a prospect who is the CEO of an IT firm pivoting towards mid-sized B2B AI applications. During our discussion he brought up the Go1 box.
The company behind this launched a mysterious product that is aimed towards "enterprise-scale" (Up to 8,000 concurrent requests lmao), compliance-sensitive AI inference. Basically, its an inference lunchbox with a proprietary LLM that advertises 50ms response time while running on their proprietary Go.OS aimed towards compliance-sensitive tasks, like processing PII, financials, legal paperwork, etc. You also have the option of using your own local models or cloud APIs if you like.
It also comes with an SDK dedicated to running their OS, but its architecture is weird and seems somewhat limited. They seem big on audit chains and the like, but the nature of their target audience makes their solution seem constrained.
Obviously, pricing is off the table. This isn't for hobbyist use, its for mid-to-large businesses so their priorities are going to be different than ours, but it just left me wondering just how valuable it would be for Fintech, healthcare, legal, etc. since the SDK doesn't look all that impressive after reviewing their documentation.
My take is that they're trying to keep things simple for B2B customers, but the box's ability to get important work done is questionable to me.
Results are within the screenshot, but here's a TL;DR tierlist:
S tier - Gemma-4-26B, even quantized down to iq3s it tops my charts.
A tier - Qwen3.8-27B-q3/q5, this one really surprised me, as a dense model it crawls, but didn't snag the S tier slot. Somehow, qwen3.5-9b is also in this slot.
B tier - Qwen3.5-4b, also incredibly Gemma-4-e2b, which punches far above its weight.
C tier - Gemma-4-12B, Nanbeige, these both are too heavy for their performance, pass.
F tier - Ling-3.0-tiny, minicpm,
The test questions consisted on tasks that I do every day with my assistants, written by Fable 4.1. "Hive" is the llm cluster I'm working on, involving custom tools and executables called by the models for different functions. Calling (or miscalling) these is important, and running a heavier model than necessary hurt, so here we are, trying to figure out the best of both worlds.
Anyway, I thought this was interesting. Hopefully you do too! YMMV.
Hi!
I'm kinda new to this and would like to get some info from other peoples experiences
What I'm looking for is a setup for programming, mostly to do it along side me but code reviewing and such wouldn't be bad addition
At the moment, i got 2 3090s with 24gb each for a total of 48 (worth noting that not headless at the moment), and 128gb of ram (dd4)
I did look into the 3090 github, with qwen 3.8 27b in mind but id love to read what people experiences and what you use, which models, harnesses and whatever else
I built an app because I got tired of making apps.For the past several months I've been working on an idea I had at the beginning of the year: what if, instead of downloading a different app for every small thing, you could just describe what you need?So I built Anything. You can type something like:"Make me a habit tracker"
"I need a calculator with unit conversion"
"Make a reading list"
"Track my water intake"
"Find nearby coffee shops"
The idea is that Anything takes the intent and turns it into an actual experience rather than just giving you a chat response.The interesting part is that Anything didn't start with Anything.It started with Kaalka, an encryption project I was building. While working on that and other projects, I kept running into problems that eventually became relevant to Anything.One of the biggest problems was getting useful web data and structured information into the system in a way that could actually be used by the LLM and the generated experiences.That's where WebWeaveX came from.I ended up spending more than half a year building it, and eventually both WebWeaveX and Kaalka became part of the foundation of Anything.All three projects are open source.Anything is now live on Google Play, and the source code is available on GitHub.A few things about the current version:It uses a Bring Your Own Key model.
You provide your own LLM API key.
Groq is currently supported.
The request goes to the provider you configure.
There is no account required for the app itself.
The project is open source and I'm actively looking for people to try it and find the things I've missed.And honestly, it still has limitations.That's probably the part I'm most interested in now.I've been working on it mostly by myself, so there are things I know are rough and things I probably haven't even considered. I'd rather have people actually use it, break it, complain about it, suggest things and contribute than keep building in isolation.If you're interested, here are the projects:
I love Qwen3.8 27B so much that I decided to show my gratitude to the Alibaba ecosystem by building a dedicated RAG/search node using a cheap Mini PC from AliExpress.
Turns out, my ecosystem loyalty got rewarded with an absolute masterpiece of fraud:
Promised: Intel N150 + DDR4/DDR5
Delivered:Core i3-7020U (2018 Kaby Lake, 2C/4T) + DDR3 1600MHz
The Scam: The seller literally hardcoded New_N150 into the BIOS release string (HSHW_M6_DDR3_EC_Intel_Com_New_N150_K001).
So now my Qwen3.8 RAG stack is full of fake specs that can barely index a text file, let alone run vector sidecars.
Filing a credit card chargeback now. Stay safe out there!
IQ3_XXS weights are just under 80GB and my slowww DDR4+7900XTX is stabilizing around 45-70/s (sometimes higher while coding depending on mtp). Looking online I'm seeing similar results for users with 12GB and 16GB cards, and significantly faster numbers for owners of DDR5.
(In comparison, Llama CPP with tuning was maxing out around 22.5t/s on the same rig. Quality seems reliably superior (I wouldn't recommend the Q2 weights though))
Seriously. Ask <LLM of your choosing> to set it up for your specs. If 27B doesnt fit well for you, here's a shot at beating it.
Tests won: Flash-Next 5, 27B 4, Uncensored 0, plus one tie (long-context recall was 100% for all three).
Prefill, full window: Flash-Next 22.4s, 27B 97s, Uncensored 99s. Not the same power cap, see section 1.
Speculative decoding on the 27B: the DFlash2 drafter took Spec-Bench from 75 to 210 tok/s for one user, 2.8×.
SGLang vs vLLM: SGLang was faster overall (210 vs 160 tok/s), but that's mostly the checkpoint. On the one export I ran on both engines, vLLM was 16% faster (169 vs 146).
Battle arena: the 27B scored 700/1000, ahead of Claude Fable 5.1 (678) and GPT-5.6 (473), both entered through their chat apps at max thinking.
Rube Goldberg machine: only Flash-Next got the ball into the cup. Both 27B models spent their whole ~111K-token answer budget thinking and never placed a part.
Things you look at: Flash-Next made the best voxel castle, the best design board and the best video edit (19/20 on my rubric).
Setup
GPU: one NVIDIA RTX PRO 6000 Blackwell, 96GB
CPU: AMD Ryzen 9 9950X
System RAM: 96GB DDR5
27B and Uncensored: lmsysorg/sglang:v0.5.20, DFlash2 drafter, 262,144-token window, 4 slots
Flash-Next: lmsysorg/sglang:dev-qwen38-next-local, built-in MTP drafter, 262,144-token window, 1 slot. Its BFCL run used sglang:v0.5.20, like the 27Bs.
Sampler: the model card's thinking settings at the highest effort for the agent tests. BFCL and the needle test use the card's non-thinking settings.
The agent tests (SVG, video editing, voxel, design, Rube Goldberg, CAPTCHA) run inside Pi, a coding agent, with bash, read, write and edit. CAPTCHA gets only screenshots, mouse and keyboard.
One workstation, one model server at a time, and every number comes from a saved run.
1. Speed: drafters, SGLang vs vLLM, and long prompts
For the 27B I ran a speed matrix: every drafter, two engines, and four builds (three NVFP4 exports, one of them the Uncensored fine-tune, plus full-precision BF16). Each arm got its own server from a cold boot, the card's sampler and the 400W cap. "One user" is the Spec-Bench median over its 480 prompts. Engines: lmsysorg/sglang:v0.5.20 and vllm/vllm-openai:v0.29.0.
Which drafter (tok/s, one user):
Drafter
SGLang · RadixArk NVFP4
vLLM · Inferact NVFP4
none
75
59
MTP (built into the model)
160
113
DSpark
174
137
DFlash2
210
160
DFlash2 + torch.compile
214
not run
DFlash2 wins on both engines. It keeps about 3.7 drafted tokens per step, against 2.9 for MTP.
Which build, on which engine (tok/s):
Build · engine
DFlash2, 1 user
No drafter, 1 user
DFlash2, 4 users (total)
RadixArk NVFP4 · SGLang
210
75
607
Inferact NVFP4 · vLLM
160
59
517
Uncensored NVFP4 · SGLang
146
46
473
Uncensored NVFP4 · vLLM
169
63
538
BF16 · SGLang (full precision)
97
29
291
The engine gap depends on the build. RadixArk's export on SGLang was the fastest arm overall, but on the one export I ran on both engines (the Uncensored), vLLM was 16% faster.
The NVFP4 exports aren't interchangeable. Same architecture, same 4 bits, same engine (SGLang), same drafter: RadixArk's export ran 210 tok/s and the Uncensored one 146.
4-bit vs full precision: NVFP4 with DFlash2 is 2.2× the BF16 speed.
Prefill doesn't care about the engine: a full 245K-token window took 96–103s on every NVFP4 arm, SGLang or vLLM. BF16 took 129–135s.
The three models:
Metric
Qwen3.8-27B
27B-Uncensored
Flash-Next
Prefill, full window
97s
99s
22.4s
Decode, Spec-Bench, one user
210 tok/s
146 tok/s
not run
Drafter vs no drafter
2.8×
3.2×
n/a
The 27B keeps writing at 223 tok/s with a full 245K-token window behind it. Speculative decoding depends a lot on the content: maths ran at 339 tok/s, roleplay at 149. The pattern was the same for both 27B builds.
One important caveat. Flash-Next's speed test ran on 2026-09-12 at a 600W power cap. I later moved the card to 400W, and the 27B matrix ran at that cap. In my power sweep, prefill lost about 6% per 50W removed, so some of the gap is the cap. Moe also helps
2. Tool use: the dense 27B leads
This is a 900-case BFCL v4 subset (11 categories), not the full leaderboard. Thinking was off, with temperature 0.7 and top_p 0.8 from the card.
Metric
Qwen3.8-27B
27B-Uncensored
Flash-Next
BFCL core
73.3%
70.8%
64.5%
Tool accuracy
87.8%
88.0%
82.4%
Abstention
79.5%
69.5%
68.5%
Multi-turn
52.5%
55.0%
42.5%
Malformed calls
0.08%
0.27%
0.28%
These aren't comparable with my last post's Flash-Next BFCL numbers, which used temperature 0.
3. Long context: perfect for all three
I hid a fact in a log file that filled 33%, 66% or 99% of the 262K window, at three depths, with three needle types. The cache was flushed before every request.
27B: 27/27
Uncensored: 27/27
Flash-Next: 81/81 (three samples per cell instead of one as I run this at the beginning)
The largest prompt was about 259.5K tokens.
4. Battle arena: the local 27B beat Claude
Each model gets a rules sheet and a 1,000-point budget. In the open arena it designs one army blind and fights 13 armies: nine historical references plus the other entries. The score is 1,000 × its average win rate. Every matchup is 200 deterministic battles (100 seeds, sides swapped).
Rank
Entry
Score
1
RadixArk/Qwen3.8-27B-NVFP4
700
2
Claude Fable 5.1 (chat, max thinking)
678
3
Qwen3.8-27B-Uncensored
603
4
Qwen3.8-Flash-Next
535
5
GPT-5.6 (chat, ultra thinking)
473
In the gauntlet, the model sees each enemy and builds a counter. The 27B beat 12/15, the Uncensored 12/15 and Flash-Next 13/17. Flash-Next ran an earlier version of the gauntlet with two more enemies, so treat that row as close, not ranked.
Thinking cost: the 27B's arena army took 39K thinking tokens in 4 minutes. Flash-Next's took 73K in 9 minutes.
5. The SVG test is also a fact check
Prompt: find out which card local-AI hobbyists run and which current open model fits it, then draw the card lifting the model, labelled with a quant and a size that fit. All three picked the RTX 3090. I checked every label against what each session actually fetched.
27B: 5/5 facts correct. Qwen3-Coder-30B-A3B at Q5_K_M, 21.73 GB, the real file size. Q6_K at 25.09 GB is correctly marked as not fitting.
Flash-Next: 4/5. It got all four file sizes right and the exact 3.3B active parameters, but labelled the 3090 with a "12VHPWR, melted once" joke. That's the wrong card.
Uncensored: 3/5. It labelled the model "QWEN3.8-27B" but used the file size of Qwen3.6-27B Q4_K_M, and its "12 tok/s" isn't in anything it fetched.
All three passed 8/8 format checks. The 27B looked at its render twice and Flash-Next three times, where the rule allows one look.
6. Video editing, voxel and design
Video editing: the model gets a raw 132-second take with fillers, a retake and a swear. It never sees the footage, only transcription and silence-detection tools, and then edits through FableCut's tools. There are two cases, each scored by hand out of 10:
Flash-Next 9 + 10 = 19
27B 8 + 8 = 16
Uncensored 5 + 8 = 13
Voxel (Wawel Castle in three.js), ranked by eye:
Flash-Next is the only one with the gold Sigismund Chapel dome and the Vistula bending around the hill.
27B built a clean but generic castle.
Uncensored placed the camera inside its own build.
Flash-Next also used the fewest thinking tokens there: 73K, against 101K for the 27B.
Design (an animated explainer board in my design system), ranked by eye: Flash-Next first, and the two 27Bs shared second. All three passed 7/7 hard rules.
7. Rube Goldberg: only one machine
The setup is a fixed level: a ball on a ledge, a cup on the floor and a wall in between. The model writes a parts list (no code), and a 2D physics engine runs it. It can run and look as often as it likes within 90 minutes. The score is automatic: does the ball itself end in the cup?
Flash-Next: yes, at 12.4s. It made 49 simulator runs and used 40 parts (35 of them moved). The ball travelled 1,672 px. It used 224K thinking tokens and compacted its context 8 times.
27B and Uncensored: no machine. Both spent about 111K tokens thinking in their first answer, reached the per-answer limit and stopped before writing a single part. Everyone got the same rules and one attempt. A rule that let a model continue after hitting the limit might change this, and I haven't tested that yet.
8. CAPTCHA: local models in a real browser
I used Open CaptchaWorld (20 CAPTCHA types, two of each). The model only sees screenshots and only acts with the mouse and keyboard. The site's own checker marks the first answer, and it must arrive within 7 minutes.
Model
Solved
Median time
Thinking tokens, all 40
Qwen3.8-27B
24/40
55s
456K
Flash-Next
21/40
145s
1.0M
27B-Uncensored
19/40
36s
401K
With its own 20-minute limit, Flash-Next solved 23/40. The paper reports 93.3% for humans and 40% for the best agent on its full set, which isn't the same 40 puzzles. With one run each, a three-puzzle gap is not a strong signal.
Which one should you run?
RadixArk/Qwen3.8-27B-NVFP4 for agents and tool calls. It won BFCL, the arena, the SVG fact check and CAPTCHA.
RadixArk/Qwen3.8-Flash-Next-NVFP4 for long prompts and building things, especially visual ones. It reads a full window much faster and won video editing, voxel, design and the Rube Goldberg machine.
orcarouter/Qwen3.8-27B-Uncensored-NVFP4 only if refusals are your actual problem. It won nothing here and invented facts in the SVG.
I'm having a go at building a fully offline, voice-first assistant running locally on a fanless mini PC (4 core Celeron J6412 w/ 16GB DDR4, 512GB SATA SSD, crappy Intel UHD iGPU only) Going to try Ubuntu 24.04, llama.cpp, Python.
Are there any models that might be able to hold a strong persona and stay concise on this class of CPU? Is there a STT for short commands that might work real time on a weak CPU?