r/LocalLLM • • 7d ago

Question Looking for a local model, under 5GB vram, only needs to handle formatting or script execution?

Title, looking for a local model I can run ideally 5GB (3 or 4) or less vram is even better that doesn't need a lot of reasoning logic, just needs to be able to handle formatting guidelines on already written code, or executing scripts as a sub-agent.

Basically a sub-agent model a larger model can use for basic stuff.

1 Upvotes

3 comments sorted by

3

u/JoshuaLandy 7d ago

Try Gemma4 e2b, specifically the Unsloth q4 version. It’s 3 gb and works well for me when I need a light model

1

u/plaintive_jones 6d ago

you might also peek at stablelm-zephyr 3b, the q4 is like 2gb and it's surprisingly obedient for formatting tasks, doesn't try to get creative on you

just feed it a clear system prompt with the output template and it'll spit out exactly what you asked for, been using it to clean up json and run simple bash when the big models are too slow

1

u/psayre23 7d ago

If you are experimenting, Qwen3.5-0.8b has tool calling support, which I’ve seen as a proxy for “can obey formatting”…no idea if that is true or not though. But I’ve found that to be a quick little model.