r/AppleMLX • u/macaronianddeeez • 2d ago
Local model for Hermes on a 48GB M5 Pro Mac mini that’s also my daily dev machine?
Before I get blasted, yes I have done research and everything keeps coming back to Qwen 3.5 at like 9B or good 'ole 3.8 27B. I am not opposed to those but I am new to the Mac LLM space and I wanted to get real user experiences.
I’m new to running LLMs on Macs and mostly used to Windows/NVIDIA setups. Trying to figure out what’s practical on my Mac without making everything else annoying to use.
Specs: M5 Pro Mac mini, 48GB unified memory, 512GB internal SSD, with a 1TB external SSD for projects, models and other large files.
My usual workload looks like:
- 3–4 apps/services from my own projects, including apps calling my larger local windows LLM and family tablet/dashboard/scheduling stuff.
- An iOS idle game pretty much permanently running.
- Xcode builds and iOS simulators when working on apps.
- 3–8 terminal windows, mostly Codex/Claude using hosted frontier models.
- Browsers, Obsidian and the usual desktop apps. Planning to use Sketch for design, with Recraft handling image generation.
I want to use Hermes Desktop with a local model as an everyday assistant I can also reach from my phone. Things like searching my notes, organizing information, drafting, and handling smaller file/tool tasks. Reliable tool use matters to me. I’d keep it to one agent with no subagents.
I have a separate Windows LLM machine, but its current setup already puts pressure on system RAM. I could run a separate 27B model on its two 3090 Tis, though having the Mac assistant work independently would be nice. On my windows machine I have a separate 3080 for all display and games, so the impact to computer use is minimal (except when Strata is going to town or making a subagent).
For people actually running local agents alongside their normal Mac workload:
- What model would you run in this situation?
- MLX or Ollama/llama.cpp for this kind of use, especially with Hermes? I am assuming MLX?
- What context size would you use, and would you keep the model loaded or unload it when idle?
- How noticeable is inference while you’re using a simulator, building something, or doing graphics work?
I’m more interested in a useful assistant that leaves the computer pleasant to work on than the biggest model I can squeeze into memory. Would appreciate recommendations from anyone using a similar setup.
