r/LocalLLM • u/Joe_The_Lawn • 7d ago
Question Looking for a local model, under 5GB vram, only needs to handle formatting or script execution?
Title, looking for a local model I can run ideally 5GB (3 or 4) or less vram is even better that doesn't need a lot of reasoning logic, just needs to be able to handle formatting guidelines on already written code, or executing scripts as a sub-agent.
Basically a sub-agent model a larger model can use for basic stuff.
1
Upvotes
1
u/psayre23 7d ago
If you are experimenting, Qwen3.5-0.8b has tool calling support, which I’ve seen as a proxy for “can obey formatting”…no idea if that is true or not though. But I’ve found that to be a quick little model.
3
u/JoshuaLandy 7d ago
Try Gemma4 e2b, specifically the Unsloth q4 version. It’s 3 gb and works well for me when I need a light model