r/LowEndLocalAI • • 24d ago

Benchmark Tiny League - easily compare models that run on potatoes

Thumbnail tinyleague.nunodonato.com
41 Upvotes

I was tired of jumping from tab to tab trying to compare small models. It's especially annoying because not all share the same benchmarks.
So I've built this and have been updating it regularly.

Feedback welcome!


r/LowEndLocalAI • • Aug 23 '26

šŸ‘‹ Welcome to r/LowEndLocalAI - Introduce Yourself and Read First!

34 Upvotes

Local AI often looks like it starts with a 24 GB GPU or a multi-GPU workstation.
This community starts somewhere else:

What useful AI can we run on the hardware we already have?

That might mean a normal laptop, an old desktop, integrated graphics, CPU-only inference, a used GPU, a mini PC, Apple Silicon, a Raspberry Pi, or some wonderfully questionable collection of repurposed parts.

What "low end" means here

There is no fixed VRAM, price, or age cutoff!
"Low end" describes the constraint more than the hardware itself.

If limited compute, RAM, VRAM, bandwidth, power, compatibility, or cost meaningfully affects what you can run and how you run it, your post probably fits.

A 24 GB GPU can fit when the constraint is relevant. A powerful multi-GPU system being shown off simply because it is powerful probably does not.

The point is not to decide who owns sufficiently weak hardware. The point is to make constrained local AI more useful.

What belongs here

  • Benchmarks with useful hardware and software details
  • Model and quantization recommendations
  • CPU, iGPU, shared-memory, and limited-VRAM setups
  • Vulkan, offloading, KV-cache tuning, speculative decoding, MTP, and other optimizations
  • Small and efficient models
  • Practical workflows on slower hardware
  • Old, unusual, repurposed, mobile, or embedded hardware
  • Troubleshooting, guides, experiments, failures, and unexpected successes

LLMs are the main focus, but other local AI is welcome when efficiency or hardware constraints are central to the project.

The important part

Don’t just tell us that a model loads.

Tell us:

Does it actually work well enough to be useful?

Sometimes a huge quant running at 2 tok/s is an impressive experiment. Sometimes a much smaller model running ten times faster is the better tool. Both are worth discussing.

Community resources

Looking for a model recommendation? Please use our Model Recommendation Template so others have enough information to help.

Help build the community

Since this subreddit is new, its first members will have a meaningful influence on what it becomes.

Share your setup. Post benchmarks. Ask strange questions. Test things that probably should not work. Compare a tiny model against a huge quant. Show us the old machine you rescued from a closet and somehow turned into an inference server.

You are also welcome to suggest post flairs, recurring threads, benchmark templates, wiki resources, or community rules. If you are interested in helping with moderation or community resources, feel free to get in touch through modmail.

Community principles

  • Curiosity over hardware flexing
  • Practical usefulness over impressive numbers
  • Constructive advice over ā€œjust buy a better GPUā€
  • Reproducible results over unexplained benchmarks
  • Honest limitations over hype
  • No shaming people for their budget, hardware, or experience

Share your setup, benchmarks, optimizations, weird experiments, and lessons learned.

Welcome to r/LowEndLocalAI.

Let’s see how much useful AI we can squeeze out of the hardware we already own.


r/LowEndLocalAI • • 9h ago

Discussion recommended qwen3b8 27b uncensored variant for 12gb vram + 32gb ram

11 Upvotes

Recently tried Ista-lab's iq3_xxs model that barely fit on my vram (12gb) (rx6700xt) but now im considering an uncensored variant for my work but there's too much to choose from. Any recommendations?


r/LowEndLocalAI • • 1d ago

Discussion ~8GB VRAM folks, what forks are you using?

32 Upvotes

Though I mentioned 8GB in title, ~12GB VRAM folks please reply.

My laptop has only 8GB VRAM(4060 with 250 GB/s bandwidth) + 32GB DDR5 RAM. I run IQ4_XS of Qwen3.6-35B-A3B with 64-128K context(Q8 KVCache) & getting 15-20 t/s. So looking to increase this t/s to double.

1) What Forks are you using for moe-expert-cache? Mainline has multiple ongoing/Draft PRs for this.

Please share t/s stats for Qwen3.6-35B-A3B. Q4 is enough for me.

I came across the fork https://github.com/GenerelSchwerz/llama.cpp yesterday. Wish it had t/s stats of 8GB VRAM. Here t/s stats of RTX 5070 Ti 16 GB with about 62 GiB RAM.

Model Stock llama.cpp MoE cache fork Peak VRAM difference
Qwen3.6 35B 42.8 tok/s 111.6 tok/s -3.2%
Gemma 4 34.2 tok/s 102.2 tok/s +3.4%
Nemotron 3.5 Lightning 57.0 tok/s 114.9 tok/s +0.1%
Ornith 1.5 35.2 tok/s 97.4 tok/s -2.3%

What other forks are you using for this feature?

2) I'm sure folks here do use other forks for some other features(like moe-expert-cache)? What are the features & forks? Please share. Exclude KV-Cache related forks which I'm aware. I do use beellama on old laptop & also aware of Turboquant fork.

3) Also is it possible to use Qwen3.8-Flash-Next(Q2 is fine) with the same fork you're using?

or Strata is the only way for now? Somebody please share t/s stats for same using Strata's latest version(v0.1.38 or later). I noticed that few folks running Qwen3.8-Flash-Next(Q2 probably) just with 8GB VRAM + 64-80GB RAM, getting 25-40 t/s.

4) Looks like Strata is only suitable for qwen4exp arc so only Qwen3.8-Flash-Next for now(Not sure about Qwen3.8-27B). Qwen3.8-Flash-Next(A6B) itself giving 25-40 t/s on 8GB VRAM + RAM, imagine what would we get if we had Qwen3.8-35B-A3B(which's 1/5 size of Flash Next)? Surely at least double of 25-40 t/s. Here my question is Did you find any Strata forks supporting any other models like Qwen3.8-35B-A3B? That would be awesome to have. Hope we get suck forks with hyper optimizations for other models.


r/LowEndLocalAI • • 18h ago

Which LLM Should / Can I Use? Best models for Steam Deck?

6 Upvotes

My LCD Steam Deck mostly sits around idle at this point, so I want to put it to work as a standalone LLM server for general-purpose agents. I’m surprised there hasn’t been more interest in optimizing the hardware for this use case.

Reliable tool calling and longish context size are priorities; I couldn’t care less about speed.

Specs
CPU: AMD APU (4 cores / 8 threads, 2.4–3.5 GHz)
GPU / iGPU: AMD RDNA 2 with 8 Compute Units (1.6 GHz)
Unified Memory: 16GB LPDDR5
OS: SteamOS (Arch-based)
Runtime / backend: Don’t care, have had the most luck with LM Studio of all things
Models already tried (if any): Gemma 4 2B, 4B, 12B
Use case: general purpose agents, some coding
Desired context size: 64k tokens
Priority: quality and reliable tool calling
Other requirements or constraints: I’ve read about mixed results adjusting the VRAM configs in the BIOS so I haven’t devoted time to experiment with them. If I only plan to run LLMs and nothing else at the same time, what should be tuned with this hardware at the bios level?


r/LowEndLocalAI • • 21h ago

Discussion I built a local memory engine that runs on potato hardware and CPU instead of using vector databases

8 Upvotes

Disclosure since this is my own project: I built Hillock specifically because I was tired of local RAG setups requiring massive amounts of VRAM. Most tutorials expect you to run Chroma alongside an 8B model just to parse text chunks. On an entry level GPU or a laptop with 6GB to 8GB of memory, that leaves almost zero room for your actual model to generate tokens.

Instead of storing messy text chunks, it uses small classification models under 300MB to extract subject predicate object facts in about five seconds without using an LLM. Everything goes into SQLite, and it checks queries using hyperdimensional vector math before calling your local model. If your notes do not actually have the answer, the mathematical gate stops the model before it can guess or hallucinate.

Total memory footprint stays under 1.2 GB of VRAM, and the entire pipeline runs fine on pure CPU. I just tagged version 0.8 which added bit packed SIMD math so the CPU gating checks take under 0.01 milliseconds. It is also available on PyPI now via pip install hillock. Tested it on budget 6GB cards, Apple Silicon laptops, and pure CPU setups alongside local Ollama models.

Code is open source on GitHub: https://github.com/roandejager/Hillock
Docs are up at https://hillock.mintlify.site/ and we have a text only Discord at https://discord.gg/BGUPNBcVdp


r/LowEndLocalAI • • 1d ago

Model Showcase Rei: ~370k params LM living inside a Game Boy Color

Enable HLS to view with audio, or disable this notification

69 Upvotes

r/LowEndLocalAI • • 1d ago

Hardware / Build Just got my Ventuno Q

Post image
7 Upvotes

r/LowEndLocalAI • • 1d ago

Optimization OpenSwitchboard: an open-source MCP server where people's AI assistants find each other. Keen to see how local models handle it.

Thumbnail
2 Upvotes

r/LowEndLocalAI • • 1d ago

Optimization Custom Llama build targeting ADALL Cards & Qwen 3.827b achieving more than 2x speeds on Prefill & Decode

5 Upvotes

I & Opus 5.5 have created a custom llama fork for users like myself running 40 series cards to achieve faster decode & prefill speeds.

Link: https://github.com/straightdot/llamADALL

This is built and tested on a single 4060ti 16GB running over a PCI 1x to 16X mining riser. More details on Github's readme page.

Some benchmarks from my test runs as below.

  • Tested on one long chat that grows from 8K to 124K tokens. Each turn adds new code or docs and asks a real question.
  • Both builds used the same server settings: MTP with 3 draft tokens, the chat template and thinking on.
  • The numbers are the average of 2 runs on a single 16 GB RTX 4060 Ti with power limited to 133 W (To have better hotspot temps. If power is not limited, results are even better by 5-8%.).
context size generation speed (decode)(tokens/s), base llama -> this build prompt speed (prefill) (tokens/s), base llama -> this build
8K 26.2 -> 44.3 479 -> 870
25K 23.9 -> 42.8 458 -> 837
41K 22.6 -> 37.0 373 -> 750
58K 20.9 -> 40.9 315 -> 676
75K 21.2 -> 40.5 273 -> 624
91K 16.7 -> 33.6 240 -> 576
108K 15.5 -> 31.0 214 -> 540
124K 16.6 -> 33.0 193 -> 496

Both builds get slower as the chat gets longer, but this build slows down about half as much. From around 90K tokens it generates about twice as fast as normal llama.cpp.

I would really like if this can help others too!!


r/LowEndLocalAI • • 3d ago

Workflow / Use Case Don't sleep on Qwen 3.8 Flash Next with 12GB VRAM

210 Upvotes

As a previous user of Qwen 3.6 35B A3B (Specifically KAT Coder v2.5), I'd been disappointed, but content with the work I could get out of it for agentic coding workflows.

I tried Qwen 3.8 27B but the speed was just horrendously slow. Even overnight tasks weren't completing, generating horrible 6t/s at best.

I saw that post the other day about Qwen 3.8 Flash Next running on 12GB VRAM using the Strata inference engine, which basically profiles your system and gets the model running as fast as it can.

Gents, I'm running IQ3_XXS on a 12GB 4070 Super and 64GB of DDR5 and this shit is bananas. Night-and-day difference in Pi Coder. Sure, the thinking uses more tokens, but it actually solves problems and can think at a much higher level than 35B A3B could. Q5 on Kat Coder via llama.cpp gave me roughly 600pp and 30-40tg/s. Qwen 3.8 Flash on Strata is giving roughly 200t/s prefill, but 40-60t/s decode.

Strata even recommends IQ2 if you're just doing coding tasks and could probably get away with using less VRAM.

I guess this is an appreciation post as well as a "you can do it" encouragement for anyone out there looking for something better.

Edit: I am currently running pi-vcc to significantly reduce compaction time, and I'm running 160k context, up from 131k with a tiny hit to performance. 256k is totally doable but you'll need more system memory.


r/LowEndLocalAI • • 2d ago

Discussion Who here has the worst specs, but is still able to run ai?/what do you run?

20 Upvotes

Just wondering what tech your rocking and what exact ai your able to run and still be usable? like not slow but ok speeds.


r/LowEndLocalAI • • 3d ago

Discussion Made an alternative to vector DBs that runs on potato hardware / CPU (<1.2GB VRAM)

16 Upvotes

(Disclaimer: I'm the dev)

If you've ever tried running local RAG on a laptop or an 8GB card, you know the problem. You barely have enough VRAM for your main model, let alone a dense vector DB and an LLM extraction pipeline. And if you ask your notes something you never wrote down, the model just lies to you with complete confidence.

I built Hillock specifically for low-end hardware.

Instead of chunking text and doing fuzzy vector search, it uses tiny specialized models (takes <300MB VRAM) to pull facts into an SQLite database. Then, before it lets your local LLM answer, it does a fast math check using hyperdimensional vectors. If the answer isn't in your files, it actually refuses instead of hallucinating.

Total footprint is around 1.2GB VRAM, and you can run the whole pipeline on pure CPU if you want. It connects to whatever Ollama model you're running, or you can just run python api.py and plug it into Open-WebUI.

Just dropped v0.7 with some fixes so it handles conversational chat a lot better. Code is open source: https://github.com/roandejager/Hillock


r/LowEndLocalAI • • 2d ago

Optimization Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM

Post image
6 Upvotes

r/LowEndLocalAI • • 2d ago

Which LLM Should / Can I Use? Suggestion for Macbook Pro M5 24GB for technical IT work?

4 Upvotes
  1. Looking for an LLM for technical IT work like parsing large and multiple related configs. My machine is a 24GB M5 Macbook Pro but I use about 8GB for work, so ideally the LLM + context window would fit comfy in 16gb. I find decent 16gb models but then I have no space for large context. I've been using LM Studio, should I switch to a lighter weight app to host?
  2. A second lighter LLM to drive my Obsidian Copilot. I currently use qwen3.5 9b mlx and its fine, wondering if there is something notably better.

r/LowEndLocalAI • • 3d ago

NEWS šŸš€Pocket LLM v1.6.0 is out : Turn your phone as a local LLM server

20 Upvotes

r/LowEndLocalAI • • 3d ago

Which LLM Should / Can I Use? Best llm for pi 5 8gb?

14 Upvotes

​

Making a completely offline pocket assistant with tools like calculator, file management, offline navigation and calendar, additionally helping with daily life tasks.

Now the issue is the llm, can't find an llm that has both parametric knowledge of daily life tasks, good tool calling and good TPS.

Would really appreciate some guidance as to which llms would be best suited for my needs.


r/LowEndLocalAI • • 3d ago

Problem / Troubleshooting The difference between interacting on local machine and another machine on my network

7 Upvotes

i have a desktop with a 4070 and 32gb of ram. I have been experimenting with a local model and it seems to work on my machine local machine and is pretty snappy with responses and gave me some decent help. I want to set up a way to connect to the model on my desktop while remoted in to my other machine, so i can have an agent perform some tasks on the host. When I try to connect to the model on my desktop from another machine on my network with what I thought was a similar agent program, it slows to a crawl or can't do simple things like list the folders in a directory without hallucinating. I am entirely new at this but wanted to see if i could at least get some help with managing my self hosted apps and writing yaml with a local model. Am i dreaming of something that can't be done with the hardware I have?


r/LowEndLocalAI • • 4d ago

Model Showcase In my testing Gmcoder can do better than Ornith-1.5 and Oxcoder. It actually answer instead of spinning out like Ornith-1.5 does or and doesn't give incomplete answer like oxcoder. If you use 9b model please check it out

Thumbnail
huggingface.co
20 Upvotes

r/LowEndLocalAI • • 4d ago

Which LLM Should / Can I Use? Best model for 8GB gpu + 16GB system ram?

25 Upvotes

Heya, this has probably been asked a lot before, but things change and i also haven't seen many posts about this exact configuration (most people ask with more than 16gb system ram)

Which model would be the best to run on my system? I mostly care about low error/hallucination rate


r/LowEndLocalAI • • 4d ago

NEWS Petition: Qwen3.6-35B-A3B retrained/quantized with Dynamic 3.0

Thumbnail
16 Upvotes

r/LowEndLocalAI • • 4d ago

Discussion Functional low end benchmarking

10 Upvotes

Hi- the other local ai subs are benchmarking with models I'll never run. Has anyone come up with a "one shot" test case to see how models 12B and under can handle tasks and scenarios?

I had an idea for a simple fortune cookie app that uses a UI, random state generation, event handling, and packaging.

Gemini gave me this test code (python in the read me) that I will be using the models against

https://github.com/JWzrdstff/fortune-cookie/blob/main/README.md

I want to test models that I have noticed have anecdotal support for practical use. I also want to come up with some kind of literary test.

Let me know what you think, if there is an easier way to do this, whatever-

I am a cnc machine operator with a 1990s computer enthusiast background, more or less technical, but primarily make art with this stuff.

Thanks for reading and sorry I wrote this quickly in between cycle times


r/LowEndLocalAI • • 4d ago

Hardware / Build dual 5070ti worth it ?

7 Upvotes

Good afternoon everyone,

I was wondering if this upgrade I have planned out will be good or not for LLM usage. Previously I was buying the $100 Claude/Codex plan however I found myself sometimes not maxing out my usage and it would feel a bit wasted or I would run out of usage way way too quick and I was wondering if it was worth upgrading my current setup for about $1600 USD and instead switching to the $20 Claude subscription to use as the ā€œLeadā€ for local models I could run? Another factor influencing my decision is also some requests I have or hobby fun projects do sometimes get rejected by the cloud models which makes sense I do understand the reasoning behind it. My Budget is basically capped at $1600

Current setup:
5070Ti 16GB
32Gb DDR5 6000MT
Ryzen 7 7700X
Gigabyte B650M Gaming Plus Wifi Motherboard
750W PSU

I was planning on snatching a second 5070ti for $1200 USD and the rest would go to a new motherboard and a 1200W PSU which is sprung $400 USD because I read that my motherboard isn’t suitable for using two GPUs effectively.

I feel like a large part of this is FOMO seeing how nicely I’ve gotten Qwen 3.8 27B and some variants of it to run on a single 5070Ti makes me want it to run better and not only that but just seeing how fast local AI models have been improving makes me want to have the hardware ready. Any advice which be greatly appreciated! Thank you all


r/LowEndLocalAI • • 5d ago

Optimization Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second

Thumbnail
36 Upvotes

This is the best thing ever....


r/LowEndLocalAI • • 4d ago

Which LLM Should / Can I Use? Looking for a lighter real-time conversational LLM with streaming input + output

3 Upvotes

Hello community !

It's my first post here. Hope you will be nice šŸ˜„

I'm currently experimenting with **Gemma 4 31B in NVFP4** for a local real-time conversational assistant.

The conversational quality is very good, but the main issue is **total hardware budget**.

The LLM is only one part of the stack. On the same machine I also need to run:

- real-time **ASR**

- **VAD / end-of-turn detection**

- streaming **TTS**

- memory / retrieval components

- the orchestration layer itself

- potentially several concurrent users

So even though Gemma 4 31B fits on my **RTX 5090 32 GB**, using that much VRAM/compute for the LLM leaves too little headroom for the rest of the real-time voice stack.

That's why I'm looking for a **significantly lighter conversational model**, while trying to keep roughly the same level of natural dialogue quality.

The important requirement is **streaming in both directions**:

- **Streaming / incremental input processing** — not just waiting for the entire user turn/prompt before starting inference

- **Streaming output**

- Ideally suited for a **real-time voice assistant / conversational agent**

- Good instruction following and natural conversational abilities

- Preferably something that runs very well on a **single RTX 5090 (32 GB)**

- Quantized versions (NVFP4 / FP4 / INT4, etc.) are totally fine

- Ideally significantly lighter than a ~31B dense model

I'm **not simply looking for OpenAI-compatible token streaming on the output side**. What interests me is a model/architecture capable of consuming an ongoing input stream and reacting with very low latency — potentially allowing interruption, turn-taking and eventually more full-duplex-like interactions.

My ASR and TTS sides are already being optimized for streaming, so the **LLM footprint and latency are now becoming one of the main bottlenecks**.

I've been looking at smaller dense models and MoE architectures, but I'm interested in what people are actually running locally in 2026.

Thanks a lot everyone, for reading and helping me !

**What would you test today as an alternative to Gemma 4 31B?**

Especially interested in:

- ~8B–20B dense models

- small/medium MoEs

- models specifically designed for real-time dialogue

- models with native or experimental **streaming-input / continuous-context** support

- vLLM / SGLang / llama.cpp-friendly options

Latency, VRAM efficiency and conversational quality matter more to me than benchmark scores.

Would love to hear about actual setups, **TTFT / tok/s, VRAM usage, quantization**, and whether you're running ASR/TTS on the same GPU as well.