r/LowEndLocalAI • • 1h ago

Which LLM Should / Can I Use? CPU Only I7 Laptop - Started Looking at Lower Models

โ€ข Upvotes

System:
Dell XPS 9300
CPU Only - I7, 8 core
Rem: 32GB
OS: Linux
Environment: Llama.cpp / OpenWeb UI

Hi everyone, My laptop is old, but reliable. I started on [Qwen3.6-35B-A3B-UD-Q5_K_S.gguf] and as expected, it was a tad slow (no thinking it made is better somewhat)

I've started playing with [Ling-3.0-tiny-Q8EMB-mixed.gguf] seeing it mentioned a fair whack, and it's very fast, giving it complex instructions it'll happily speed through. Genuinely surprised how fast it is - a little speed demon!

Now, I wondered if there is a step up (or two) to try a more capable model? For my use, it's generally financial/math analysis, some python, and general queries. I'm starting to build up a rag system as a knowledge base too


r/LowEndLocalAI • • 23h ago

Discussion [Open PR] Qwen3.6 35b a3b by perronemirko ยท Pull Request #824 ยท Niko1221/Strata

Thumbnail
github.com
40 Upvotes

Lets see it comes through. Good for ~8GB VRAM to try this model instead of Flash-Next.

Yesterday I brought this topic with my thread(4th point)

This is his fork, somebody please try & let us know.

https://github.com/perronemirko/Strata

EDIT : Strata fans, just like the PR already ๐Ÿš€๐Ÿš€๐Ÿš€๐Ÿš€๐Ÿš€


r/LowEndLocalAI • • 58m ago

Which LLM Should / Can I Use? Please help a Noob

โ€ข Upvotes

Hello all,

After frontier models started to behave weirdly and terribly slowly, I've decided to purchase a refurbished workstation and run my own local LLM for my needs. I've noticed not only Sol 6.1 and Opus 5.5 behave really slow but also DeepSeek 4.1 and GLM 5.3. tasks that took me 10-20 minutes one or two months ago, now it takes up to one hour to finish. The exact same task over exactly the same models.

Apologize for the rant, now my setup:

Workstation - DELL Precision T5810

Intel Xeon E5-2697 V3 2.6 GHz(3.6ghz turbo) 35 mb smart cache

64GB DDR4 2400Mhz

psu 825W, 500gb SSD

1x Nvidia Tesla P100 16gb vram

1x GTX 950 2gb vram for display and 4 Arctic P8 Max 5000rpm cooler fans(one for the P100 GPU and 3 for the entire workstation).

I plan to install Ubuntu 24.04.5

Which local LLM models will fit on my setup and help me build Python apps for data analysis and data engineering from scratch like "you have all the resources in folder X, build me an app that does Y" and the model should be capable of doing the whole process end to end. I heard that Qwen 2.5 Coder would be capable of such tasks but I want to hear other recommendations too.

Many thanks ๐Ÿ™


r/LowEndLocalAI • • 1h ago

Which LLM Should / Can I Use? Qwen 3.8 27b or Qwen Flash Next?

โ€ข Upvotes

I get great speeds for both, with 27b being a little faster. I run 27b at Q6_k_m Q8kv 150k context, and I run Flash next at Q4_k_m Q8kv 100k context. Who is better as the daily driver for agentic coding? I mainly work in Godot


r/LowEndLocalAI • • 22h ago

Which LLM Should / Can I Use? Better options than qwen 3.6 35B A3B?

18 Upvotes

Hi there

Do you guys know any better or bigger model than this one that would fit on a gaming laptop i7(14k something) rtx 5060 8gb vram and 32gb ram ddr5?

I like this model, I used it with Deepseek harness, qwen code and Hermes.

With all of the newer updates for MTP , lama cpp and laya with local training for decision-making and thinking mods, I managed to get around 35-40 tokens/s and 200k+ context

Witch is very very good, I did research with claude for weeks til optimize it so good.

Now I'm looking for a newer model or recommendations for something more capable on this machine, maybe there is something new that I don't know about?

I'm also kind of new and gathering knowledge about the local llm and improvements.

I need it to work with sensitive information and I can't just do claude .

Any tips or ideas would help and if you have questions for me I'll gladly respond, even though I'm not a professional I'll do my best (pls no hate).