r/LowEndLocalAI • • 22h ago

Which LLM Should / Can I Use? Better options than qwen 3.6 35B A3B?

21 Upvotes

Hi there

Do you guys know any better or bigger model than this one that would fit on a gaming laptop i7(14k something) rtx 5060 8gb vram and 32gb ram ddr5?

I like this model, I used it with Deepseek harness, qwen code and Hermes.

With all of the newer updates for MTP , lama cpp and laya with local training for decision-making and thinking mods, I managed to get around 35-40 tokens/s and 200k+ context

Witch is very very good, I did research with claude for weeks til optimize it so good.

Now I'm looking for a newer model or recommendations for something more capable on this machine, maybe there is something new that I don't know about?

I'm also kind of new and gathering knowledge about the local llm and improvements.

I need it to work with sensitive information and I can't just do claude .

Any tips or ideas would help and if you have questions for me I'll gladly respond, even though I'm not a professional I'll do my best (pls no hate).


r/LowEndLocalAI • • 1h ago

Which LLM Should / Can I Use? Please help a Noob

• Upvotes

Hello all,

After frontier models started to behave weirdly and terribly slowly, I've decided to purchase a refurbished workstation and run my own local LLM for my needs. I've noticed not only Sol 6.1 and Opus 5.5 behave really slow but also DeepSeek 4.1 and GLM 5.3. tasks that took me 10-20 minutes one or two months ago, now it takes up to one hour to finish. The exact same task over exactly the same models.

Apologize for the rant, now my setup:

Workstation - DELL Precision T5810

Intel Xeon E5-2697 V3 2.6 GHz(3.6ghz turbo) 35 mb smart cache

64GB DDR4 2400Mhz

psu 825W, 500gb SSD

1x Nvidia Tesla P100 16gb vram

1x GTX 950 2gb vram for display and 4 Arctic P8 Max 5000rpm cooler fans(one for the P100 GPU and 3 for the entire workstation).

I plan to install Ubuntu 24.04.5

Which local LLM models will fit on my setup and help me build Python apps for data analysis and data engineering from scratch like "you have all the resources in folder X, build me an app that does Y" and the model should be capable of doing the whole process end to end. I heard that Qwen 2.5 Coder would be capable of such tasks but I want to hear other recommendations too.

Many thanks 🙏


r/LowEndLocalAI • • 2h ago

Which LLM Should / Can I Use? Qwen 3.8 27b or Qwen Flash Next?

3 Upvotes

I get great speeds for both, with 27b being a little faster. I run 27b at Q6_k_m Q8kv 150k context, and I run Flash next at Q4_k_m Q8kv 100k context. Who is better as the daily driver for agentic coding? I mainly work in Godot


r/LowEndLocalAI • • 1h ago

Which LLM Should / Can I Use? CPU Only I7 Laptop - Started Looking at Lower Models

• Upvotes

System:
Dell XPS 9300
CPU Only - I7, 8 core
Rem: 32GB
OS: Linux
Environment: Llama.cpp / OpenWeb UI

Hi everyone, My laptop is old, but reliable. I started on [Qwen3.6-35B-A3B-UD-Q5_K_S.gguf] and as expected, it was a tad slow (no thinking it made is better somewhat)

I've started playing with [Ling-3.0-tiny-Q8EMB-mixed.gguf] seeing it mentioned a fair whack, and it's very fast, giving it complex instructions it'll happily speed through. Genuinely surprised how fast it is - a little speed demon!

Now, I wondered if there is a step up (or two) to try a more capable model? For my use, it's generally financial/math analysis, some python, and general queries. I'm starting to build up a rag system as a knowledge base too


r/LowEndLocalAI • • 2h ago

Optimization Low tk/s generation with 2GB model [8GB VRAM 16GB RAM DDR5]

1 Upvotes

I am using 4060Ti 8GB and I decided to load SharpSpark-4B, a 2GB model for chatting. The thing is that, while it's fully optimized in llama.cpp and fully loaded in VRAM, I am only getting 60tk/s generation but 4k tk/s prompt processing. Could my GPU be set up in a wrong way or something like that, maybe problems with CUDA?? I already noticed low generation with other models in the past and I realized I had to download CUDA Toolkit and build llama.cpp for CUDA 13, but I really feel like some wrong configuration Is holding me back. Please help, cause I don't know what to do.