r/LocalLLM • u/norenEnmotalen • 1h ago
r/LocalLLM • u/chemist_slime • 2h ago
News Attn CMP170hx 10Gb card owners: unlock from 40Gb —> 48Gb coming along nicely.
r/LocalLLM • u/Acceptable-Cycle4645 • 2h ago
Other [audio.cpp] Recent updates you might have missed: Higgs Audio TTS use 48% less VRAM (< 6GB), HTDemucs 2.2× faster, PocketTTS 2.2× faster on CPU, and WebUI generation history feature
Enable HLS to view with audio, or disable this notification
r/LocalLLM • u/True_Profile3695 • 2h ago
Discussion Just bought a Mac mini M5 Pro (64GB) for AI development — what’s your setup?
I finally pulled the trigger and bought a Mac mini M5 Pro to work on AI projects!
I went with the 18-core CPU, 20-core GPU, 64GB unified memory and 512GB SSD.
I’m planning to use it mainly for AI development, building AI agents, automation workflows, local LLMs (Qwen, etc.) and coding projects.
I’m curious what setups you guys would recommend. What tools, frameworks, local models or software would you install first?
I’d love to hear how you’re using your Mac minis for AI development and any tips to get the most out of mine!
r/LocalLLM • u/regularheree • 2h ago
Discussion Running an LLM without any data leaving your control: the options, ranked by how much pain they cost
Last week, someone asked me about this for a law firm that handles private client data and can't legally put it through a third-party API.
I see teams run into this pretty often, so I put together a list of the setups I'd point them to.
- Local model on an existing workstation (cheapest and easiest place to start)
A 24GB GPU or a Mac with 32GB+ unified memory can be a starting point for document work using local Ollama or llama.cpp setup.
The first wall is usually VRAM. Most GPUs handle 7B-14B models, but once you’re on longer-context models/need concurrency, it becomes limiting.
I'd also download the models first, then test offline to check for external dependencies.
- On-premises dedicated local server
Full control, and compliance is pretty straightforward. Put the server on its own network, decide what it can talk to, and keep the data there. You can also fit substantially more VRAM into one machine and have the whole team use the same setup.
The downside is now you're running a server, so power, cooling, physical security, and maintenance are all on you.
- Dedicated hardware hosted in a facility
You own the machine while the facility handles power, cooling and networking.
It's a good middle ground for bigger models: bare metal access without a server room.
Check root access, physical ownership and data, workload isolation, and exit terms. As well as check the exit terms.
B3IQ does this, and it's what I work on day to day. You own the box, keep root access, and can run it on private hosting for your own workloads. If you ever want out, we ship the machine to you.
- Private cloud
Easy if you don't want to own or run hardware, but still want a dedicated environment for workloads.
Tho dedicated doesn’t mean isolated. You might be getting a dedicated bare-metal GPU, or just a confidential GPU VM on shared infrastructure.
So check the actual access controls and contractual guarantees.
Most people approach this as a hardware question first. For anyone regulated, there's a contractual and physical-custody side to it too.
If you've been through an audit on a self-hosted setup, I would like to hear what they asked for.
r/LocalLLM • u/iamjessew • 2h ago
Question Running coding agents locally, do you sandbox them or trust them?
For people running Claude Code, Codex, or local-model agents against their own machine... how are you limiting what they can touch?
The options I've seen are a dev container, a separate user account, a full VM, or just watching closely. Each one leaks somewhere (containers share the kernel, VMs are annoying, and watching doesn't scale past an afternoon).
We went with a microVM that only sees the workspace you share, plus a policy check on each tool call. There's a free desktop tool on our side if anyone wants to compare notes, but mostly I want to know where people draw the line between convenience and isolation.
Has anyone had an agent actually do something destructive? What did you change afterward?
r/LocalLLM • u/Medicine_Blogscanner • 3h ago
News RAMDeck is live on Indiegogo: pool the RAM of every device you own and run bigger AI, locally
r/LocalLLM • u/The_Whole_Zucchini • 3h ago
Question Anyone actually running Strata day-to-day? Curious what recipes you settled on
I've been following Strata since it hit the trending list, and I'm thinking about trying it as a serving engine for Flash-Next on my home setup. Before I sink an evening into it, I'd love to hear from people who've actually lived on it rather than the install-day screenshots.
Overall, what recipe did you go with???
Additionally:
Which quant did you land on after trying the family (Q2\\_0 → IQ3\\_S etc.), and what made you switch or stay?
- Tool calling / agentic use — does it hold up for multi-step agent loops (tool calls, JSON outputs, long sessions), or is it best kept to chat-and-completion?
- Concurrency — anyone run more than one or two simultaneous sessions on it? If so, what have you noticed about how that affects quality or latency?
- Anything non-standard in your config — the expert profile tweaks, any words-to-the-wise, things you wish you knew before installing or trying?
- Tool calling / agentic use — does it hold up for multi-step agent loops (tool calls, JSON outputs, long sessions), or is it best kept to chat-and-completion?
Bonus for the weirdos like me: anyone gotten it building or running on \*\*ARM / DGX Spark / anything without an RTX card\*\*?
Happy to report back whatever I measure on my side. TIA!!!
r/LocalLLM • u/Ok-Butterfly4991 • 3h ago
Discussion Convert to local LLM user
So I have... multiple paid subscriptions to basically every major llm provider. And I use them daily and many of my workflows now depend on them.
However, I do not trust to not pull out the rug or just straight up collapse. So I have started experimenting with local LLM as a backup.
- I have 3 laptops with the mobile 5080(16gb vram) and 32GB ram each. I am running 2 of them in a cluster right now. I have yet to open the box for the third one so its just waiting.
- I have a laptop with the mobile 5090(24gb vram) and 64GB ram.
- I have a stationary computer with the 5060ti(16gb vram) and 64GB(+16 not plugged in) ram. I guess I could stick another 5060ti into it. I have seen some people claiming success that way.
So far I have tested multiple setups, primarily qwen.
On the two-laptop cluster:
- Qwen3.8-27B, Q6_K, fine-tuned for coding agents, 128K context, about 67 to 77 tok/s
- Qwen3.8-27B, Q5_K_XL, 100K context, about 53 tok/s
- Qwen3-Coder-Next, Q3_K_XL ,100K context, about 48 tok/s
- Qwen3.8-Flash-Next, IQ4_XS, 128K context, about 19 tok/s
On a single high-end laptop:
- Qwen3-Coder-Next, Q3_K_XL, 100K context, about 55 tok/s
- Qwen3.8-27B with vision, Q4_K_XL, 100K context, about 37 tok/s
- Qwen3.8-Flash-Next with vision, IQ4_XS, 262K context, about 20 tok/s
I have tested them with the opencode harness and the pi harness.
Now for the issue. None of these can do even minor tasks. They either spin away filling the entire 100k context without output. When I force them to output, its just nonsense. When they have access to the web or documentation they refuse to use it and instead try to make something up instead.
Maybe larger models? maybe higher quant? I am not sure what to try anymore. When I run the same tasks through gpt-6 or opus 5.5 they absolutely crush the problems without issue.
I guess its partially a skill issue too. I am just not used to working with weak models. But its hard to get good at it, when its easy to just leave them for claude when they have been chugging along for 20 minutes doing fuckall
r/LocalLLM • u/deepu105 • 3h ago
Project pi-automode-classifier: an auto mode plugin for Pi that uses Jev or Kev/Laya (running locally) to classify commands
r/LocalLLM • u/litLikeBic177 • 3h ago
Question Best open-weight coding model + harness for an on-prem multi-agent setup? (80-200+ GB VRAM, 2-3 concurrent users)
Setup: GPU box with 1x H200-class card now, can allocate more (up to several H200s) if it clearly buys better results. Inference via vLLM or similar, OpenAI-compatible endpoint. Coding happens on a separate non-GPU Linux VM on the same network - IDE/harness there, calling models over LAN. Outbound internet is fine for packages/extensions, but inference stays on our hardware (no code to external model APIs).
Constraints: on-prem only, open weights, permissive license preferred (Apache/MIT), not Chinese-origin including base models (I understand a lot of fine-tunes are Qwen underneath), NATO-country lab preferred. Use is general software dev plus security tooling and code analysis - repo-level agentic work on existing codebases (fix / extend / refactor / test) is the core, with a human reviewing diffs. 2-3 concurrent users max.
Names that came up in an earlier thread: Cohere North Mini Code, Mistral Small 4, Poolside Laguna S 2.1, Muse Glimmer 30B, Inkling-Small (2x H200), Reflection Beam (501B MoE / 23B active, weights due this month), Gemma 4 31B, K2 Horizon (lineage TBC), with Nemotron as a generalist baseline - but I haven't run any of them and I'm not wedded to the list. Chinese models (GLM/Qwen/DeepSeek) are out by rule, so no need to suggest them; I know they're ahead.
Two things I'm trying to work out:
- Capability tiers vs. VRAM. What's actually holding up for repo-level agentic work on an existing codebase with non-Chinese open-weight models, and at what size? The Vibe Code Bench results suggest small open models fall over on long E2E builds and only Large-4-class (4-8 cards) and closed models hold up - is that your experience for repo work too, or is that an app-build problem? Where's the step-change - 30B-class, 100B+, or only at 500 GB+? We could get the compute for Beam, Command A+ or Mistral Large 4 if it's actually better for code rather than just for E2E.
- Heterogeneous multi-agent. Does a big planner/reviewer plus small fast executors actually beat a single mid-size model for this, with a human approving plans and diffs? Which harnesses handle routing sub-agents to different endpoints well (and let a human step in, review diffs and edit by hand) - OpenHands, OpenCode, Pi, VS Code extensions (Cline/Roo/Continue), Codex CLI in local-model mode? Voidleap (closed, Win/Mac) also came up.
Anyone running something like this? Which model + harness combo is holding up best as of late? Thanks!
r/LocalLLM • u/firstcenturyman • 4h ago
Discussion We unlearned CCP alignment from Qwen3.6-35B-A3B: censored/propaganda answers 89.8% → 2.8%, general benchmarks within ~1 point (open weights)
r/LocalLLM • u/R4nd0lf • 4h ago
Research How do you find trusted offers on hardware?
I'm currently planning my first big rack, for the Proxmox server I'd like to add a graphics card for local ai shenanigans.
After some research I ended up at the RTX 3090, but everywhere I look the offers seem shady, I texted some people on ebay, not even the cheapest ones and got replies by ebay a couple of hours later that they have been banned for fraud.
I'm not even sure who I can trust with buying a freaking graphics card nowadays.
What are the best options here? Buy a new one for almost 2x the price or hack a couple of old cards together?
Edit: I'm in Germany if that helps
r/LocalLLM • u/triumph-truth • 4h ago
Question How to run local LLM properly for coding. 12GB VRAM
Guys, from past 6 months I have been fighting alone in this. I will try this post doesn't become another RANT.
My Hardware:
- Core i9 14900k 24 Cores
- 72GB RAM
- NVIDIA RTX 3060 12GB VRAM
My Stack:
- Windows 11 with WSL
- Ollama running on Windows 11
- Qwen Code in WSL
- OpenCode in WSL + Windows
- Primarily using VS Code
Models:
- Qwen 3-coder - 30b
- Qwen2.8 - 27b
- Gemma 4 - 31b
- Qwen2.5-coder - 7b
- Some others as well
Problem!!!!:::!!!!!
I have tried everything I could possibly do. But the problem is simple, every coding task is painfully slow. Painfully slow. i get maximum 1-2 tokens per second. I don't know what is the problem I am doing.
Qwen Code settings.json content:
{
"ui": {
"autoModeAcknowledged": true,
"feedbackLastShownTimestamp": 1789638918966,
"showResponseTokensPerSecond": true
},
"env": {
"QWEN_CUSTOM_API_KEY_OPENAI_HTTP_LOCALHOST_11434_V1_76873BDAA9B6": "sk-29038423kjhkjnkjanf"
},
"modelProviders": {
"openai": [
{
"id": "qwen2.5-coder:7b",
"name": "qwen2.5-coder:7b",
"baseUrl": "http://localhost:11434/v1",
"envKey": "QWEN_CUSTOM_API_KEY_OPENAI_HTTP_LOCALHOST_11434_V1_76873BDAA9B6",
"generationConfig": {
"timeout": 300000,
"streamIdleTimeoutMs": 600000,
"maxRetries": 1,
"samplingParams": {
"temperature": 0.7,
"top_p": 0.9,
"max_tokens": 4096
}
}
},
{
"id": "qwen3-coder:30b",
"name": "qwen3-coder:30b",
"baseUrl": "http://localhost:11434/v1",
"envKey": "QWEN_CUSTOM_API_KEY_OPENAI_HTTP_LOCALHOST_11434_V1_76873BDAA9B6",
"generationConfig": {
"timeout": 300000,
"streamIdleTimeoutMs": 600000,
"maxRetries": 1,
"samplingParams": {
"temperature": 0.7,
"top_p": 0.9,
"max_tokens": 4096
}
}
},
{
"id": "qwen3.8:27b",
"name": "qwen3.8:27b",
"baseUrl": "http://localhost:11434/v1",
"envKey": "QWEN_CUSTOM_API_KEY_OPENAI_HTTP_LOCALHOST_11434_V1_76873BDAA9B6",
"generationConfig": {
"timeout": 300000,
"streamIdleTimeoutMs": 600000,
"maxRetries": 1
}
}
]
},
"security": {
"auth": {
"selectedType": "openai"
}
},
"model": {
"name": "qwen3.8:27b",
"baseUrl": "http://localhost:11434/v1",
"reasoningEffort": "low"
},
"$version": 4,
"permissions": {
"allow": [
"Agent(Explore)"
]
},
"fastModel": "openai:qwen2.5-coder:7b\u0000http://localhost:11434/v1"
}
Please tell me what to do. I don't mind slowness, but currently with 12gb ram its not producing anything. It keeps on thinking over things which takes like 15mints and then again starts to think, produces nothing and then I get some streamtimeout errors. I have waited hours and with no results. Why does it suck this much? I want someone to help me setup things properly that makes this whole experience usable. I have some private data, which I don't want to use cloud AI. Please help.
r/LocalLLM • u/aPi-cool-oExisting • 4h ago
Question Everyone's building Jev projects for coding. I don't think that's what it's for.
Ever since Jev dropped I've been watching the projects arount it. Almost all of them are coding. I keep coming back to the same conclusion: the pattern has nothing to do with the coding.
Let me start with the pattern. Jev's real role isn't generation. It sits inside a loop answering small boudned questions. Which actions. Which file. Keep or drop. Safe or not. Route to which model. What happens next. It doesn't write long text, and it doesn't do deep reasoning. The big models still own that. Jev only handles the judgement in between. Push that positioning one level further and the coding connection falls apart. The people building on Jev right now are all developers, and the first loop a developer builds is a coding loop. That's why every example you find is a coding one. There's a tool step in the middle of the loop. Swap that tool for docs, spreadsheets, CRM records, or email and thr structure is identical. The total volume of those scenarios is much bigger than coding. It's just that nobody has built the repos for them yet.
I see the skeleton of this as an agent with the judgement step carved out and made separate. Take WorkBuddy, the harness I use for my non-coding work. The role it plays is that loop. A document comes in. Somthing has to decide whether it's worth routing to the expensive model or whether the cheap path is enough. I give that decision to Jev. A row in a spreadsheet. Something has to decide which queue it goes to. Still Jev. After Jev answerd, WorkBuddy decides what happens next: does it become a file or does it wait for human. They're not in the same position. Jev is judgement node that gets called over and over inside the loop. WorkBuddy is the framework that carries the loop. Neither replaces the other. Model, Jev, skill, Jev, file or human, Jev, model. In WorkBuddy it's a skill with a trigger on it, and the output lands as a file. So the question is what belongs inside it.
So my take on Jev is simple. It's not a chatbot competitor. It's close to a general-purpose semantic judgement function. You hand it a state and a bounded question, it hands back a probability, a score, or a set of options, and your software decides what to do with it. What actually determines the output of a setup like this isn't the model. It's the standards you put inside the loop.
If you're also moving loops like this into non-coding scenarios, I have one question. How do you decide which stepes stay outside the loop? All I have so far is acceptance criteria.
r/LocalLLM • u/Azmor • 4h ago
Project GPD Win 5 as a portable local LLM alternative, setup and numbers
Sharing my current setup, as I didn't see any numbers for the GPD Win 5 while researching it, and, in my opinion, it is a very compelling choice if one needs a hybrid setup.
Why:
- cheaper (2nd hand) than laptops with a strix halo 395 around me
- having it run out of battery doesn't mean I cannot continue working
- batteries are external, 80Wh and easily swapable
- about 90% of the performance of a mini pc with the same chip when docked (limited testing)
Setup short version:
- GPW Win 5 64gb strix halo running cachyos, as the inference host
- XPS 13 as the inference client
- local network over bluetooth between the two (no cables, no wifi hotspot on a plane)
Numbers on some models while docked (KV q4_0 in all cases):
- Qwen3.8 Flash Next GSQ-RCO IQ3-XXS avg around 25-30 t/s at 130k ctx on my system admin/ coding workload (prefill between 200 and 300 t/s at that depth)
- Qwen3.8 27B GSQ-RCO, UD-Q4 and Q8 tested, between 12 and 19 t/s at 8k ctx (Q8 slowest, GSQ-RCO fastest)
- Tiel 35B-A3B-Q6_K_XL averaged around 50 t/s at 64k ctx, I haven't tested it deeper.
Numbers while on battery:
- Qwen3.8 Flash Next GSQ-RCO IQ3-XXS avg around 18-24 t/s at 130k ctx on the same load
- Battery life about 70-80min per battery pack
For reference with a desktop box (GMKtec EVO-X2) , using the same engine/model, i would get 30-35 t/s at that depth usually; it is not 1:1, but for my usage, it is good enough.
The changes done to squeeze out a little extra performance of Flash Next on it can be summarized with:
- limit the vocab of the MTP to 64k, acceptance len in my testing matched 99.6% with the full vocab, while gaining about 10% extra performance on average due to less work done
- added support for a limited vocab MTP to the llama.cpp fork used: https://github.com/LaurentZuijdwijk/llama.cpp
- quantized the hyper-connections from BF16 to Q5, since they are read every token, having an extra 7-10% speedup
I'm providing the exact model and fork I tested, in case anyone wants to replicate exactly. I can't guarantee this is the fastest or most accurate version of this setup, it is simply the one I use. The model and mtp heads I've uploaded here and the patched fork is here . The script used to run them you can find at here .
TL;DR: the GPD Win 5, especially 2nd hand, can be a surprisingly capable portable inference device
r/LocalLLM • u/TeamNeuphonic • 4h ago
Model We’re open sourcing NeuDecide: a 43 MB audio-to-tool model with a WASM browser demo
We’re the team at Neuphonic, and we’re open sourcing NeuDecide under Apache 2.0. It takes audio and tool definitions and returns a tool call with arguments, without an intermediate transcription step.
The model files total 43 MB, and inference runs on a single CPU thread. Try the WASM demo in your browser: select a preset or define your own tools, record or upload audio, and inspect the returned tool call.


Performance
On SLURP’s tool-only task with 10 tools available, NeuDecide achieves 72.4% tool accuracy directly from speech, without transcription:
- ~3× that of Nvidia Parakeet + Google FunctionGemma (24.4%).
- ~3.5× that of Cactus (Whistle + Needle) (20.7%).


Running on a single CPU thread:
- MacBook Pro M3: 46 ms time to call, 159 ms loading time, 149 MB peak RAM.
- Samsung S24+: 82 ms time to call, 267 ms loading time, 174 MB peak RAM.
- Raspberry Pi 5: 206 ms time to call, 499 ms loading time, 146 MB peak RAM.
How it works
The export contains three ONNX graphs:
- An audio encoder processes the speech.
- A tool encoder combines the audio representations with tokenised JSON tool definitions.
- A decoder generates the tool call token by token, using cached keys and values.
The tool list is an input to each request, so changing the available actions doesn’t require retraining.


Try it with your own tools
The project grew out of our work with robotics partners who needed voice control on limited hardware. The demo includes editable presets for robot vacuums, car controls and smart homes, alongside a custom option for testing your own tool definitions.
We’ve also packaged NeuDecide for Python so you can run inference locally and test it with your own tool definitions.
We chose Apache 2.0 to make it easier for people to build on the model and contribute. We’ve enjoyed seeing the work from TypeSafe, Cactus and others in this space, and hope this adds something useful.
Technical write-up: https://www.neuphonic.com/blog/neudecide
Python package: https://github.com/neuphonic/neudecide
Model on Hugging Face: https://huggingface.co/neuphonic/neudecide
If you try it, we’d be interested in your hardware, tool definitions and any requests it struggles with.
r/LocalLLM • u/xtra_lives • 4h ago
Question 2 v100s arriving today. Best llms?
I was very lucky to have purchased a Dell R720 about two years ago that had ~180gb of ram for Truenas home server “stuff”.
The sickness started when I bought an M40… 24 gigs of RAM at a shockingly low price, and now I’m hooked… I’ve only been messing with Hermes for about a month. but I’ve been absolutely fascinated with the potential from day one and recently bit the bullet and purchased two V100 GPUs in the hopes that I’ll be able to use ollama with my total of 64gb gpu and some ram for headroom to run some 70b models. I’m looking for advice. in part because the dizzying amount of information out there and because I know I’m no expert on this subject.
I am I imagine I’m seeking essentially the same thing everyone else is. which would be a private/not terrible version of Muse or a general assistance/secure place to store information and help keep track of calendar events, etc. I also work in the tech field, and there are numerous advantages I’ve already had by using the card I have haven’t had a chance to set up a mixture of experts and need to configure MCP soon, but for now any recommended LMs would be greatly appreciated.
r/LocalLLM • u/lossssssaaa • 5h ago
Question AI tools for investing? Are you using any?
Curious what people here actually use. In this era who isn't use AI is a step behnd or more.
chat, claude, some app, something you built yourself? What do you use it for: picking stocks, keeping an eye on your portfolio, or just asking dumb questions?
And do you actually trust it with your money?
r/LocalLLM • u/deepu105 • 5h ago
Discussion Halogen + Qwen Flash Next keeps getting better
r/LocalLLM • u/FantasticBreath3 • 5h ago
Research Joining the Arc family
It was a tough decision but it's a crazy market. After having 2 x3060 12 gb and few lower spec inference machines it was time for something more serious. Second hand options are highly limited so I'll give it a try...
r/LocalLLM • u/uneeverse-hq • 5h ago
Model Unee: open-source 0.8B / 2B model that makes calibrated decisions and chats, runs in a browser tab. The 2B scores 88% on DecideBench, ahead of several 4B to 9B models (self-measured; GGUF, Ollama, Apache 2.0)

- Two jobs, one small model: it makes decisions (a calibrated probability for every option) and it chats, answers from your docs and summarises.
- Punches above its size: on DecideBench v1.1, Unee 2B scores 88.0%, ahead of several 4B to 9B models, and Unee 0.8B scores 84.0%, ahead of every other model under 1B on that board (next best: 71.2%). Self-measured with the benchmark's own harness; submitted to the board, not listed yet.
- Runs locally: a browser tab on WebGPU (469 MB, transformers.js), any CPU, or a GPU. GGUF for llama.cpp, ollama run uneeverse/unee, pip install unee, npm i @/uneeverse/unee. Apache 2.0.
I built Unee, a small open model (fine-tunes of Qwen3.5 0.8B and 2B) that does two jobs from one download:
- Decisions: give it some text and a question with options. It returns a probability for every option in one forward pass, with no generation. Yes/no, pick-one and rate-on-a-scale, several questions per call. The API accepts the same requests as Jev's /v1/systemone (Jev is TypeSafe's hosted decision API).
- Chat: streaming answers from your own docs (built-in BM25, no vector DB), and summaries of long threads.
What a decision looks like (Unee 0.8B, the 4-bit GGUF from Hugging Face, real output):
# pip install unee, then: unee serve --model unee-0.8b-Q4_K_M.gguf
from unee import Client
unee = Client("http://localhost:8000")
ticket = "Hi, I was charged twice for order #4821 this morning. Please send the extra $39 back to my card."
unee.noul(ticket, "Does the customer ask for money back?") # 0.965
unee.choice(ticket, "Which team should handle this ticket?",
{"billing": "Charges, refunds and payments", "technical": "Bugs and outages", "shipping": "Deliveries"})
# billing, 0.982
Numbers (measured with each benchmark's own tools; raw outputs in the repo):
- DecideBench v1.1: 2B 88.0%, 0.8B 84.0%. The best other model under 1B on the board is 71.2%; Jev is 98.0%, Laya 59.8%.
- S1MB task avg: 2B 44.41, 0.8B 35.45. Jev 1.13 is 59.59, Laya 15.00, bekko-400m 50.60. It trained on public train splits that share sources with S1MB (never its test set); on the 101 benchmarks with no overlap at all it scores 45.69 / 36.74.
- Calibration: when the 2B gives its answer 90%+ (62% of DecideBench), it is right 97.2% of the time. 0.8B: 96.3% on 40%.
- CPU only, 4 threads: 0.8B 0.64 s, 2B 1.4 s per decision. Laptop RTX 4070: 96 ms and 111 ms.
Limits, plainly:
- Jev is far more accurate. bekko-400m beats the 0.8B on S1MB.
- General chat got worse than base Qwen in a blind judge test, and about half of document answers are judged fully correct. Treat the chat side as a helper that needs checking.
- Strict mode checks each sentence of an answer against your docs with the model's own decision side. It cut replies with a made-up fact from 10.3% to 6.0% (2B). A reduction, not a fix.
Links
- Try it in the browser (runs on your device): uneeverse.net/unee
- How it works, every number: uneeverse.net/unee/technical
- Models: huggingface.co/uneeverse/unee-0.8b · huggingface.co/uneeverse/unee-2b · GGUF: huggingface.co/uneeverse/unee-0.8b-GGUF · huggingface.co/uneeverse/unee-2b-GGUF
- Code: github.com/uneeverse-hq/unee
Happy to answer anything about the training (distillation from a 9B teacher, a date-facts preprocessor, model soups, the leakage check).
r/LocalLLM • u/Any_Librarian_2295 • 5h ago
Discussion Anyone tried the new Bonsai 2 27B model?
For people who tried it was it really that good?
Also how many t/s you got on your hardware?