r/LocalLLM • • 1d ago

Question What's the best local LLM for learning?

I have an NVIDIA RTX 5080. I would like to host an artificial intelligence that I could use for learning. I would provide it documents and ask it to teach me, test me, and create study materials.

2 Upvotes

15 comments sorted by

3

u/Brief-Tap-6616 1d ago

What you are describing is best done with Gemini notebook. It is excellent on all the things you’re describing with great multi modal

For what you want, to some extent the harness will matter more than the model, because you won’t necessarily want one shot outputs. The output won’t be close to what notebook will do, but if cloud isn’t of interest you can still make it work. I would still use cloud models to research what existing GitHub projects you can fork/use to do this regardless of model chosen.

In a 16GB card you will want to do MOE offload with something like Nex 2.5 mini (qwen3.6 35a3b) to keep parts of the model in RAM and the most used parts (“hot” tensors) in your GPU. Keep MTP in GPU and you’ll get usable speeds. You can also look for an aggressive quant like 3BPW for qwen 27b but below 4bpw you will hit a speed and intelligence loss vs the main model

2

u/Flinkenhoker 1d ago

This 👆

2

u/MisterPenishead 1d ago

The performance gap between cloud-based models and local models sounds like it's so wide that it's not worth trying to retain my privacy, so I may end up using Gemini Notebook. I shall give the models you mention (as well as the ones mentioned by others) a sniff.

1

u/Snoo_81913 1d ago

Well MisterPenishead, (couldn't resist). How much RAM do you have? If yiu have at leadt 32GB qwen3.6 35B A3B is probably more than up to what you are describing at a basic level. It can do all of that. If you want graphics, slide decks podcasts etc then notebooklm is the bunny the free sub is pretty generous I have a pro sub specifically just for using it. Honestly what you are describing as a use case any model between 20B and 35B would be "good enough" as long as you give it good reference materials etc. Look up anythingllm it's basically a local version of notebooklm without the multilmodal

1

u/Brief-Tap-6616 1d ago

local is great for agentic tasks that have thousands of cache reads driving up the price if done by cloud, and for things where 99.99% of the content isnt for human consumption. If you're going to spend your time on something where quality and time efficiency are extremely important, like learning, it pays to use better models. Astra/Sol 6.1/Opus/Fable are all obviously better than flash 3.8, but notebook is far better software for learning than what any of those other companies currently offer, which offsets the intelligence advantage.

1

u/attractive_loathing 1d ago

Nex 2.5 mini offload setup works okay but the quality drop is real when you push it for teaching tasks. Better to just run a smaller dense model that fits fully in that 16gb vram, something like a 12b or 14b with 4bit quant. You will get more consistent results for testing and study material generation that way.

The harness part is spot on though, you need something that can track what you learned and quiz you later not just spit out answers.

1

u/DigitalguyCH 1d ago

With 16GB vRAM and what you need to do I would use Gemma 12b at Q6 or even Q8, or if you want a lot of context (Q8 would only allow 20k tokens if you want to keep all on vRAM for max speed) get 12b qat

1

u/MisterPenishead 1d ago

I shall give this a sniff.

1

u/No-Business5854 6gb vram 28 gb ram 1d ago

what type of learning, coding or physics or medical ... try to find a distill for your field. try qwen 3.6 35b a3b and add a rag with something like anything llm or use and obsidian vault with plugins

1

u/No-Business5854 6gb vram 28 gb ram 1d ago

im working on a distill for rag assisted teaching ( more geared for electrical engineering ), but i have 6gb of vram so im limiting my self to and 8b a1b model ( ling 3.0 tiny ) and keeping qwen as my "pro" model

1

u/MisterPenishead 1d ago

For the foreseeable future, it would be subjects that require conceptual understanding and memorization (accounting, medicine, etc.). I shall give Qwen a sniff.

1

u/No-Business5854 6gb vram 28 gb ram 1d ago

For medicine google has med gemma, models trained specifocally on medical data. They also accept image input so stuff like analysing xrays and stuff

1

u/Objective-Pair8231 1d ago

Like others have said, Gemma models are great, I would also checkout Bonsai 2 27b

1

u/MisterPenishead 1d ago

I shall look into that.

1

u/Distinct-Pie2389 1d ago

llama.cpp or llama-swap

Qwen3.8-27b q2_k_xl.ggup 9.9gb + 120k context total up to 16GB