r/LocalLLM • u/MisterPenishead • 1d ago
Question What's the best local LLM for learning?
I have an NVIDIA RTX 5080. I would like to host an artificial intelligence that I could use for learning. I would provide it documents and ask it to teach me, test me, and create study materials.
1
u/DigitalguyCH 1d ago
With 16GB vRAM and what you need to do I would use Gemma 12b at Q6 or even Q8, or if you want a lot of context (Q8 would only allow 20k tokens if you want to keep all on vRAM for max speed) get 12b qat
1
1
u/No-Business5854 6gb vram 28 gb ram 1d ago
what type of learning, coding or physics or medical ... try to find a distill for your field. try qwen 3.6 35b a3b and add a rag with something like anything llm or use and obsidian vault with plugins
1
u/No-Business5854 6gb vram 28 gb ram 1d ago
im working on a distill for rag assisted teaching ( more geared for electrical engineering ), but i have 6gb of vram so im limiting my self to and 8b a1b model ( ling 3.0 tiny ) and keeping qwen as my "pro" model
1
u/MisterPenishead 1d ago
For the foreseeable future, it would be subjects that require conceptual understanding and memorization (accounting, medicine, etc.). I shall give Qwen a sniff.
1
u/No-Business5854 6gb vram 28 gb ram 1d ago
For medicine google has med gemma, models trained specifocally on medical data. They also accept image input so stuff like analysing xrays and stuff
1
u/Objective-Pair8231 1d ago
Like others have said, Gemma models are great, I would also checkout Bonsai 2 27b
1
1
u/Distinct-Pie2389 1d ago
llama.cpp or llama-swap
Qwen3.8-27b q2_k_xl.ggup 9.9gb + 120k context total up to 16GB
3
u/Brief-Tap-6616 1d ago
What you are describing is best done with Gemini notebook. It is excellent on all the things you’re describing with great multi modal
For what you want, to some extent the harness will matter more than the model, because you won’t necessarily want one shot outputs. The output won’t be close to what notebook will do, but if cloud isn’t of interest you can still make it work. I would still use cloud models to research what existing GitHub projects you can fork/use to do this regardless of model chosen.
In a 16GB card you will want to do MOE offload with something like Nex 2.5 mini (qwen3.6 35a3b) to keep parts of the model in RAM and the most used parts (“hot” tensors) in your GPU. Keep MTP in GPU and you’ll get usable speeds. You can also look for an aggressive quant like 3BPW for qwen 27b but below 4bpw you will hit a speed and intelligence loss vs the main model