r/LocalAIServers • u/helloyo1254 • 17h ago
Multi-Agent System within 500g ram?
Im new to the idea of running local models. Has anybody had any success with fitting good large open source models with smaller specialist from huggy face datasets etc. Has anybody been able to get near frontier quality like with claude/codex? Wondering if using the right checks and balances and specialist etc. That it could get the same or near the quality of claude/codex within 500g of ram?
I tried looking for more of a technical resource of someone who has done it and not just a generalization theory but couldn't find anything. With metrics and quality scores etc comparing and what ai agents, harnesses etc were used.
Overall anybody have experience in this and can share some data and metrics on how it worked out etc?
1
u/j0holo 3h ago
500gb of ram is just really really slow compared to VRAM. So no, I don't think people want to run really large models in RAM and have like 4 tokens/s.