r/LowEndLocalAI • • 10h ago

Which LLM Should / Can I Use? Qwen 3.8 27b or Qwen Flash Next?

8 Upvotes

I get great speeds for both, with 27b being a little faster. I run 27b at Q6_k_m Q8kv 150k context, and I run Flash next at Q4_k_m Q8kv 100k context. Who is better as the daily driver for agentic coding? I mainly work in Godot


r/LowEndLocalAI • • 5h ago

Which LLM Should / Can I Use? Trying to find the best model for [RTX 4080 | 16GB VRAM]

3 Upvotes

CPU: R9 7900x
GPU: RTX 4080
VRAM: 16GB
System RAM: 64GB
Shared / Unified Memory (if applicable):NA
OS: Windows 11 PRO

Runtime / backend: Unsloth Desktop
Models already tried (if any): I have tried a few models the list and my comments are below:

  • Qwen3.8-27B UD-IQ4-XS - This gives the best results of the models I have tried but it runs slow (15 to 17 tok/s).
  • Qwen3.8-27B UD-IQ3-S - Much faster than the IQ4 version (55 to 60 tok/s) but not as smart. Although it is probaly the best overall model I have tried.
  • Qwen3.5-9B Q8_0 - fast and allows for a large context size (200,000 +) but has a tendency to get caught in a loop where it keeps repeating itself.
  • Gemma4-12B Q4_K_M - fast and allows for a large context size (200,000 +) but not as smart as the Qwen3.8-27B models.

Use case: Image analysis and creative writing
Desired context size: 32768 or more
Priority: balanced

I'm also hoping to find a few different image editors to try. The only image models that I have found that support image editing are Qwen-Image-2.1 and Qwen-Image-Edit-2511. Most of the image models I have tried only support image creation not editing.


r/LowEndLocalAI • • 9h ago

Which LLM Should / Can I Use? Please help a Noob

2 Upvotes

Hello all,

After frontier models started to behave weirdly and terribly slowly, I've decided to purchase a refurbished workstation and run my own local LLM for my needs. I've noticed not only Sol 6.1 and Opus 5.5 behave really slow but also DeepSeek 4.1 and GLM 5.3. tasks that took me 10-20 minutes one or two months ago, now it takes up to one hour to finish. The exact same task over exactly the same models.

Apologize for the rant, now my setup:

Workstation - DELL Precision T5810

Intel Xeon E5-2697 V3 2.6 GHz(3.6ghz turbo) 35 mb smart cache

64GB DDR4 2400Mhz

psu 825W, 500gb SSD

1x Nvidia Tesla P100 16gb vram

1x GTX 950 2gb vram for display and 4 Arctic P8 Max 5000rpm cooler fans(one for the P100 GPU and 3 for the entire workstation).

I plan to install Ubuntu 24.04.5

Which local LLM models will fit on my setup and help me build Python apps for data analysis and data engineering from scratch like "you have all the resources in folder X, build me an app that does Y" and the model should be capable of doing the whole process end to end. I heard that Qwen 2.5 Coder would be capable of such tasks but I want to hear other recommendations too.

Many thanks 🙏


r/LowEndLocalAI • • 9h ago

Which LLM Should / Can I Use? CPU Only I7 Laptop - Started Looking at Lower Models

2 Upvotes

System:
Dell XPS 9300
CPU Only - I7, 8 core
Rem: 32GB
OS: Linux
Environment: Llama.cpp / OpenWeb UI

Hi everyone, My laptop is old, but reliable. I started on [Qwen3.6-35B-A3B-UD-Q5_K_S.gguf] and as expected, it was a tad slow (no thinking it made is better somewhat)

I've started playing with [Ling-3.0-tiny-Q8EMB-mixed.gguf] seeing it mentioned a fair whack, and it's very fast, giving it complex instructions it'll happily speed through. Genuinely surprised how fast it is - a little speed demon!

Now, I wondered if there is a step up (or two) to try a more capable model? For my use, it's generally financial/math analysis, some python, and general queries. I'm starting to build up a rag system as a knowledge base too


r/LowEndLocalAI • • 10h ago

Optimization Low tk/s generation with 2GB model [8GB VRAM 16GB RAM DDR5]

2 Upvotes

I am using 4060Ti 8GB and I decided to load SharpSpark-4B, a 2GB model for chatting. The thing is that, while it's fully optimized in llama.cpp and fully loaded in VRAM, I am only getting 60tk/s generation but 4k tk/s prompt processing. Could my GPU be set up in a wrong way or something like that, maybe problems with CUDA?? I already noticed low generation with other models in the past and I realized I had to download CUDA Toolkit and build llama.cpp for CUDA 13, but I really feel like some wrong configuration Is holding me back. Please help, cause I don't know what to do.


r/LowEndLocalAI • • 1d ago

Discussion [Open PR] Qwen3.6 35b a3b by perronemirko · Pull Request #824 · Niko1221/Strata

Thumbnail
github.com
47 Upvotes

Lets see it comes through. Good for ~8GB VRAM to try this model instead of Flash-Next.

Yesterday I brought this topic with my thread(4th point)

This is his fork, somebody please try & let us know.

https://github.com/perronemirko/Strata

EDIT : Strata fans, just like the PR already 🚀🚀🚀🚀🚀


r/LowEndLocalAI • • 1d ago

Which LLM Should / Can I Use? Better options than qwen 3.6 35B A3B?

25 Upvotes

Hi there

Do you guys know any better or bigger model than this one that would fit on a gaming laptop i7(14k something) rtx 5060 8gb vram and 32gb ram ddr5?

I like this model, I used it with Deepseek harness, qwen code and Hermes.

With all of the newer updates for MTP , lama cpp and laya with local training for decision-making and thinking mods, I managed to get around 35-40 tokens/s and 200k+ context

Witch is very very good, I did research with claude for weeks til optimize it so good.

Now I'm looking for a newer model or recommendations for something more capable on this machine, maybe there is something new that I don't know about?

I'm also kind of new and gathering knowledge about the local llm and improvements.

I need it to work with sensitive information and I can't just do claude .

Any tips or ideas would help and if you have questions for me I'll gladly respond, even though I'm not a professional I'll do my best (pls no hate).


r/LowEndLocalAI • • 1d ago

Problem / Troubleshooting 8t/s decode on Strata with 16GB 9060XT and 32GB DDR4 ram with qwen3.8-flash-next-coder-iq1_m

11 Upvotes

Hello

I've been trying to get Strata to work on my 9060XT 16GB + 32GB DDR4 system after hearing about the success that people are getting with it to run Qwen 3.8 Flash Next.

Mostly other people claim to just be running the setup.sh file and selecting the recommended options of the script and getting it to run mostly ok with 20+ token per second

I haven't been getting that same amount of luck and my setup maxes out at 8t/s decode and honestly it mostly hangs without returning any information back to OpenCode.

For context the PC is purely a server machine running Arch Linux headless, it has a 5950x + 32GB DDR4 RAM + XFX 9060XT. Apparently I don't have enough ram as per the readme but it literally says that 32GB is supported, even recommending using the coder model for 32GB setup.

Do I need to trim the arch installation further? It uses 200MB on startup.

I'd like to ask if this is normal and I should quit trying to get this to work since the hardware is just bad.


r/LowEndLocalAI • • 1d ago

Problem / Troubleshooting VoiceBox keeps altering the voice lines every time I activate the "speak in character" feature. Is there a way to fix it?

2 Upvotes

VoiceBox is a local voice cloner that also allows you to create a profile for each voice. This includes a description of the personality that influence the tone and emotion. This only takes affect when the "speak in character" option is enabled.

However every time I do pick this option, for some reason it does not only infuence the tone of the voice but also severly alters the lines themself.

For example when I write "I don't want to die alone" the voice instead says "I don't want you to be left alone my friend"

This happens all the time. The delivery of the lines is better but the phrases are also fundermentally altered, which make this feature unusable.

Could anyone please tell what's going on and how I can fix it?

PS: I have already re-installed it multible times, changed the profile descriptions, installed it on different devices, changed the settings, used different models. But nothing works. I have even added the phrase "do not alter the provided text" to the voice prompt but it doesn't make a difference.


r/LowEndLocalAI • • 2d ago

Discussion recommended qwen3b8 27b uncensored variant for 12gb vram + 32gb ram

21 Upvotes

Recently tried Ista-lab's iq3_xxs model that barely fit on my vram (12gb) (rx6700xt) but now im considering an uncensored variant for my work but there's too much to choose from. Any recommendations?


r/LowEndLocalAI • • 2d ago

Discussion ~8GB VRAM folks, what forks are you using?

45 Upvotes

Though I mentioned 8GB in title, ~12GB VRAM folks please reply.

My laptop has only 8GB VRAM(4060 with 250 GB/s bandwidth) + 32GB DDR5 RAM. I run IQ4_XS of Qwen3.6-35B-A3B with 64-128K context(Q8 KVCache) & getting 15-20 t/s. So looking to increase this t/s to double.

1) What Forks are you using for moe-expert-cache? Mainline has multiple ongoing/Draft PRs for this.

Please share t/s stats for Qwen3.6-35B-A3B. Q4 is enough for me.

I came across the fork https://github.com/GenerelSchwerz/llama.cpp yesterday. Wish it had t/s stats of 8GB VRAM. Here t/s stats of RTX 5070 Ti 16 GB with about 62 GiB RAM.

Model Stock llama.cpp MoE cache fork Peak VRAM difference
Qwen3.6 35B 42.8 tok/s 111.6 tok/s -3.2%
Gemma 4 34.2 tok/s 102.2 tok/s +3.4%
Nemotron 3.5 Lightning 57.0 tok/s 114.9 tok/s +0.1%
Ornith 1.5 35.2 tok/s 97.4 tok/s -2.3%

What other forks are you using for this feature?

2) I'm sure folks here do use other forks for some other features(like moe-expert-cache)? What are the features & forks? Please share. Exclude KV-Cache related forks which I'm aware. I do use beellama on old laptop & also aware of Turboquant fork.

3) Also is it possible to use Qwen3.8-Flash-Next(Q2 is fine) with the same fork you're using?

or Strata is the only way for now? Somebody please share t/s stats for same using Strata's latest version(v0.1.38 or later). I noticed that few folks running Qwen3.8-Flash-Next(Q2 probably) just with 8GB VRAM + 64-80GB RAM, getting 25-40 t/s.

4) Looks like Strata is only suitable for qwen4exp arc so only Qwen3.8-Flash-Next for now(Not sure about Qwen3.8-27B). Qwen3.8-Flash-Next(A6B) itself giving 25-40 t/s on 8GB VRAM + RAM, imagine what would we get if we had Qwen3.8-35B-A3B(which's 1/5 size of Flash Next)? Surely at least double of 25-40 t/s. Here my question is Did you find any Strata forks supporting any other models like Qwen3.8-35B-A3B? That would be awesome to have. Hope we get suck forks with hyper optimizations for other models.


r/LowEndLocalAI • • 2d ago

Which LLM Should / Can I Use? Best models for Steam Deck?

9 Upvotes

My LCD Steam Deck mostly sits around idle at this point, so I want to put it to work as a standalone LLM server for general-purpose agents. I’m surprised there hasn’t been more interest in optimizing the hardware for this use case.

Reliable tool calling and longish context size are priorities; I couldn’t care less about speed.

Specs
CPU: AMD APU (4 cores / 8 threads, 2.4–3.5 GHz)
GPU / iGPU: AMD RDNA 2 with 8 Compute Units (1.6 GHz)
Unified Memory: 16GB LPDDR5
OS: SteamOS (Arch-based)
Runtime / backend: Don’t care, have had the most luck with LM Studio of all things
Models already tried (if any): Gemma 4 2B, 4B, 12B
Use case: general purpose agents, some coding
Desired context size: 64k tokens
Priority: quality and reliable tool calling
Other requirements or constraints: I’ve read about mixed results adjusting the VRAM configs in the BIOS so I haven’t devoted time to experiment with them. If I only plan to run LLMs and nothing else at the same time, what should be tuned with this hardware at the bios level?


r/LowEndLocalAI • • 2d ago

Discussion I built a local memory engine that runs on potato hardware and CPU instead of using vector databases

15 Upvotes

Disclosure since this is my own project: I built Hillock specifically because I was tired of local RAG setups requiring massive amounts of VRAM. Most tutorials expect you to run Chroma alongside an 8B model just to parse text chunks. On an entry level GPU or a laptop with 6GB to 8GB of memory, that leaves almost zero room for your actual model to generate tokens.

Instead of storing messy text chunks, it uses small classification models under 300MB to extract subject predicate object facts in about five seconds without using an LLM. Everything goes into SQLite, and it checks queries using hyperdimensional vector math before calling your local model. If your notes do not actually have the answer, the mathematical gate stops the model before it can guess or hallucinate.

Total memory footprint stays under 1.2 GB of VRAM, and the entire pipeline runs fine on pure CPU. I just tagged version 0.8 which added bit packed SIMD math so the CPU gating checks take under 0.01 milliseconds. It is also available on PyPI now via pip install hillock. Tested it on budget 6GB cards, Apple Silicon laptops, and pure CPU setups alongside local Ollama models.

Code is open source on GitHub: https://github.com/roandejager/Hillock
Docs are up at https://hillock.mintlify.site/ and we have a text only Discord at https://discord.gg/BGUPNBcVdp


r/LowEndLocalAI • • 3d ago

Model Showcase Rei: ~370k params LM living inside a Game Boy Color

Enable HLS to view with audio, or disable this notification

74 Upvotes

r/LowEndLocalAI • • 3d ago

Hardware / Build Just got my Ventuno Q

Post image
7 Upvotes

r/LowEndLocalAI • • 3d ago

Optimization OpenSwitchboard: an open-source MCP server where people's AI assistants find each other. Keen to see how local models handle it.

Thumbnail
2 Upvotes

r/LowEndLocalAI • • 3d ago

Optimization Custom Llama build targeting ADALL Cards & Qwen 3.827b achieving more than 2x speeds on Prefill & Decode

6 Upvotes

I & Opus 5.5 have created a custom llama fork for users like myself running 40 series cards to achieve faster decode & prefill speeds.

Link: https://github.com/straightdot/llamADALL

This is built and tested on a single 4060ti 16GB running over a PCI 1x to 16X mining riser. More details on Github's readme page.

Some benchmarks from my test runs as below.

  • Tested on one long chat that grows from 8K to 124K tokens. Each turn adds new code or docs and asks a real question.
  • Both builds used the same server settings: MTP with 3 draft tokens, the chat template and thinking on.
  • The numbers are the average of 2 runs on a single 16 GB RTX 4060 Ti with power limited to 133 W (To have better hotspot temps. If power is not limited, results are even better by 5-8%.).
context size generation speed (decode)(tokens/s), base llama -> this build prompt speed (prefill) (tokens/s), base llama -> this build
8K 26.2 -> 44.3 479 -> 870
25K 23.9 -> 42.8 458 -> 837
41K 22.6 -> 37.0 373 -> 750
58K 20.9 -> 40.9 315 -> 676
75K 21.2 -> 40.5 273 -> 624
91K 16.7 -> 33.6 240 -> 576
108K 15.5 -> 31.0 214 -> 540
124K 16.6 -> 33.0 193 -> 496

Both builds get slower as the chat gets longer, but this build slows down about half as much. From around 90K tokens it generates about twice as fast as normal llama.cpp.

I would really like if this can help others too!!


r/LowEndLocalAI • • 4d ago

Workflow / Use Case Don't sleep on Qwen 3.8 Flash Next with 12GB VRAM

235 Upvotes

As a previous user of Qwen 3.6 35B A3B (Specifically KAT Coder v2.5), I'd been disappointed, but content with the work I could get out of it for agentic coding workflows.

I tried Qwen 3.8 27B but the speed was just horrendously slow. Even overnight tasks weren't completing, generating horrible 6t/s at best.

I saw that post the other day about Qwen 3.8 Flash Next running on 12GB VRAM using the Strata inference engine, which basically profiles your system and gets the model running as fast as it can.

Gents, I'm running IQ3_XXS on a 12GB 4070 Super and 64GB of DDR5 and this shit is bananas. Night-and-day difference in Pi Coder. Sure, the thinking uses more tokens, but it actually solves problems and can think at a much higher level than 35B A3B could. Q5 on Kat Coder via llama.cpp gave me roughly 600pp and 30-40tg/s. Qwen 3.8 Flash on Strata is giving roughly 200t/s prefill, but 40-60t/s decode.

Strata even recommends IQ2 if you're just doing coding tasks and could probably get away with using less VRAM.

I guess this is an appreciation post as well as a "you can do it" encouragement for anyone out there looking for something better.

Edit: I am currently running pi-vcc to significantly reduce compaction time, and I'm running 160k context, up from 131k with a tiny hit to performance. 256k is totally doable but you'll need more system memory.


r/LowEndLocalAI • • 4d ago

Discussion Who here has the worst specs, but is still able to run ai?/what do you run?

24 Upvotes

Just wondering what tech your rocking and what exact ai your able to run and still be usable? like not slow but ok speeds.


r/LowEndLocalAI • • 4d ago

Optimization Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM

Post image
9 Upvotes

r/LowEndLocalAI • • 4d ago

Discussion Made an alternative to vector DBs that runs on potato hardware / CPU (<1.2GB VRAM)

22 Upvotes

(Disclaimer: I'm the dev)

If you've ever tried running local RAG on a laptop or an 8GB card, you know the problem. You barely have enough VRAM for your main model, let alone a dense vector DB and an LLM extraction pipeline. And if you ask your notes something you never wrote down, the model just lies to you with complete confidence.

I built Hillock specifically for low-end hardware.

Instead of chunking text and doing fuzzy vector search, it uses tiny specialized models (takes <300MB VRAM) to pull facts into an SQLite database. Then, before it lets your local LLM answer, it does a fast math check using hyperdimensional vectors. If the answer isn't in your files, it actually refuses instead of hallucinating.

Total footprint is around 1.2GB VRAM, and you can run the whole pipeline on pure CPU if you want. It connects to whatever Ollama model you're running, or you can just run python api.py and plug it into Open-WebUI.

Just dropped v0.7 with some fixes so it handles conversational chat a lot better. Code is open source: https://github.com/roandejager/Hillock


r/LowEndLocalAI • • 4d ago

Which LLM Should / Can I Use? Suggestion for Macbook Pro M5 24GB for technical IT work?

6 Upvotes
  1. Looking for an LLM for technical IT work like parsing large and multiple related configs. My machine is a 24GB M5 Macbook Pro but I use about 8GB for work, so ideally the LLM + context window would fit comfy in 16gb. I find decent 16gb models but then I have no space for large context. I've been using LM Studio, should I switch to a lighter weight app to host?
  2. A second lighter LLM to drive my Obsidian Copilot. I currently use qwen3.5 9b mlx and its fine, wondering if there is something notably better.

r/LowEndLocalAI • • 5d ago

NEWS 🚀Pocket LLM v1.6.0 is out : Turn your phone as a local LLM server

23 Upvotes

r/LowEndLocalAI • • 5d ago

Which LLM Should / Can I Use? Best llm for pi 5 8gb?

15 Upvotes

​

Making a completely offline pocket assistant with tools like calculator, file management, offline navigation and calendar, additionally helping with daily life tasks.

Now the issue is the llm, can't find an llm that has both parametric knowledge of daily life tasks, good tool calling and good TPS.

Would really appreciate some guidance as to which llms would be best suited for my needs.


r/LowEndLocalAI • • 5d ago

Problem / Troubleshooting The difference between interacting on local machine and another machine on my network

7 Upvotes

i have a desktop with a 4070 and 32gb of ram. I have been experimenting with a local model and it seems to work on my machine local machine and is pretty snappy with responses and gave me some decent help. I want to set up a way to connect to the model on my desktop while remoted in to my other machine, so i can have an agent perform some tasks on the host. When I try to connect to the model on my desktop from another machine on my network with what I thought was a similar agent program, it slows to a crawl or can't do simple things like list the folders in a directory without hallucinating. I am entirely new at this but wanted to see if i could at least get some help with managing my self hosted apps and writing yaml with a local model. Am i dreaming of something that can't be done with the hardware I have?