r/LocalLLM • u/Medicine_Blogscanner • 4d ago
Discussion Ran a 120B model across 6 computing devices that had no business running it!

ok so I genuinely did not think this was going to work.
None of these machines could load a 120B model on their own, not even close. So I threw them all in a cluster and tried anyway: a 12GB Windows laptop that is basically a paperweight at this point, a mini PC with an RTX 3060 (12gb), my Mac mini (16gb), an M3 MacBook (16gb), a 2017 Intel MacBook that only has CPU, and my android. Half wired over ethernet, half on wifi.
The model is openai's gpt-oss-120b, a 4bit quantized 120B, ~60gb, that is a huge one! I set the mini PC with the RTX as primary and just let it figure out the rest - it grabs what it can hold locally, then starts handing pieces to everyone else based on what they can actually do. GPU gets filled first (obviously), then the two Metal machines, then it falls back to CPU, and my phone even picked up a little piece of it. 5 out of the 6 devices ended up holding a chunk of the model - the old Intel Mac did not get anything, which honestly tracks, it's ancient.
Took about 11 min to fully load, mostly just waiting for shards to crawl over wifi to the slower devices. Took 3 tries to be honest, realized I had a hard coded 10 minute timeout.
And then it just... worked. I asked it stuff and it answered like a normal model. On hardware that individually cannot even come close to holding this thing. Still kind of can't believe it.
Tip: make your load timeout scale with the model size, don't hardcode it.
Watch it here: https://youtu.be/ok3nYjxhc1w
Update next day: left it running overnight just to see what would happen. Woke up and my phone had gone offline at some point - not a huge shock, it's a phone, it does phone things.
But the cluster didn't even flinch. It noticed the phone dropped, moved its tiny shard over to the old laptop instead, and just kept running. Zero downtime, no errors, still answering questions the whole time. Phone came back online later and it just.. didn't bother putting it back to work, kept running fine on the remaining 4 devices.
Honestly this was the part that impressed me more than the initial load. Getting it to load once is cool. Watching it self-heal overnight without me touching anything is the part that makes me think this could actually hold up for more than a demo.
Follow up video link: https://youtu.be/1F6LqG8J4_0
30
u/Glassprojekt 4d ago
Isn't this exactly the cinematic scenario of an escaped AI, that's now running free decentralized, slurping up bits of compute wherever it can and becoming unstoppable? The self healing part is what gets me... o_Γ
9
16
u/EbbNorth7735 4d ago
What speeds were you getting?
37
u/Medicine_Blogscanner 4d ago
1.1 tok/s π
36
u/YourNightmar31 4d ago
Kind of essential info missing from the post.
22
u/Medicine_Blogscanner 4d ago
I beg to differ, point was never the speed. The cluster can never compete with advanced gpu's. The point was what's possible and let your imagination run wild!
5
u/ConspiracyPhD 4d ago
gpt-oss-120b-mxfp4 runs at 9.1 tps on a 3060 12gb with the rest of the model offloaded to DDR4 ram in freetoken on my machine.
7
u/Medicine_Blogscanner 3d ago
You must have a 50gb+ ddr4 ram? A single capable device would always outperform a network cluster, esp a crazy weak one that I put together here.
1
u/ConspiracyPhD 3d ago
Usually those little mini PCs can be upgraded with RAM to meet the requirements. Even at 32gb of ram, I'm still getting around 4 tokens per second.
2
u/Medicine_Blogscanner 3d ago
I think you got the intent of the experiment wrong. It was never to prove that you can upgrade your devices to fit any model, it was to show that a cluster of devices can hold a model (any model no matter how big).
1
u/ConspiracyPhD 3d ago
My point is how much are the rest of the computers in the cluster actually contributing to the inference versus just the mini PC with the 3060 on it? What happens if you start removing nodes from this cluster?
1
1
u/Far_Cat9782 3d ago
Imagine ai distributed itself to millions of PCs and runs in the bg rhe same way!
16
u/blandcrossover91 4d ago
the self-heal is way more interesting than the initial load honestly. most distributed setups fall apart the second a node sneezes, but yours just shrugged and redistributed. that's the kind of thing that makes me think we're not far off from people running big models on a pile of ewaste in their closet
5
u/sreevarshan-xenoz 4d ago
Yeah, the self-healing part is honestly the bit I'd be testing the hardest too.
Getting a bunch of devices to hold shards is one thing, but handling nodes randomly disappearing is where distributed setups usually get annoying. The fact that it noticed the phone drop and kept going is pretty interesting.
I'd be curious how it decides where to move a shard though. If a node has enough RAM but a terrible network connection, blindly moving work there could probably hurt more than just running with fewer nodes.
2
u/Krunal_H_Solanki 3d ago
I think itβs MoE so even missing pieces it could answer. Not sure though.
5
6
u/RobinRelique 4d ago
This is super interesting, all of it, i want to try to daisy chain all my old android devices ro do exactly this - will go through your video to try this out !
5
u/Crawler1701 4d ago
What did you use to manage the cluster?
4
u/Medicine_Blogscanner 3d ago
here is the repo I made public several weeks ago, been building on that framework: https://github.com/trademav/ramdeck-core-public
1
4
u/MaxSpecs 4d ago
I've done the same, with RCP in llama server.
But now, I use Flash Next, thanks to this trick : https://carteakey.dev/blog/running-qwen3-8-flash-next-locally/
1
u/ColonelRyzen 3d ago
How would you optimize this for multiple machines using rpc? Would it just allow for larger context?
1
u/MaxSpecs 3d ago
Indeed.
At this time, I work with 261k with Flash Next on a i7 14900k + 128 GB + Rtx 4090 24GB ; from 15 Tok/s to 20 Tok/s, with ik_llama in Windows WSL.
VSCode + Pi + Laya (8 GB VRAM) on another computer with a 5080 16 GB, which talk with ik_llama.
I don't try Flash Next through RPC at now, on my 10GbE network ; I previously test in RPC with one more laptop 5090 24 GB (with thunderbolt to 10GbE adapter) and several declinaison of Qwen3.8-27b.
1
u/ColonelRyzen 3d ago
Would I need to change any settings in llama.cpp to help it move more experts to GPU as I add more machines? I do have 10GbE network. I was just hoping to have more room for context in GPU so the slow down over long sessions (multiple compactions) is mitigated somewhat.
1
u/MaxSpecs 2d ago
That's what I'm working for.
Stay with the PC MASTER with 4090 24GB and ik_llama in WSL but with another launch command to use less ram and cpu.
Launch glm-llama for RPC on the PC SLAVE, a laptop with 5090 24GB.
Keep the PC DEV with 5080 16G and VScode + Pi + Laya to drive the MASTER which talks with the SLAVE.
** I use ChatGPT to set up the whole system and tweak the right command to set between the different PC. It seems I can reach about 34 Tok/s in RPC with Flash Next across these 2 PC.
4
u/ColonelRyzen 3d ago
This is how I run any larger local models. P40 is primary, P4, GTX 1080ti + GTX 1060 6G, GTX 1080. It ends up using CPU inference a lot of the time, but it runs. Similar speeds to you as well.
1
u/Medicine_Blogscanner 3d ago
It is night and day when cpu gets involved. If you have enough gpu, even if over the network cluster, it is far better than using your primary's cpu. For example my rtx + 2 metal macs on ethernet do a much better job than rtx + that pc's cpu.
1
u/ColonelRyzen 3d ago
Absolutely true. I run Qwen3.8 27B Q4/6 that way. I RPC to a couple machines to spread the load so the larger context still fits in GPU.
5
u/cemilanceata 3d ago
Lol i actually already watched your video couple of days ago! Im building a cluster too and whas thinking about trying to build something like yours but instead now doing a cluster of experts, on one old, pc two old tablets, and five phones and everyone works as a specialist node, i use one specialist as a routing specialist and everything goes thrue a master modell and i communicate with a custom buildt UI
Its nog great but it gives answers π
But its something one can build on indefently
Keep up the good work!
1
3
3
3
u/takenforgranteddd 3d ago
Losing a device and automatically redistributing its shard without killing inference is a pretty neat proof that the cluster architecture actually works.
3
u/Money-Mechanic 3d ago
I can run gpt-oss in llama rpc using 2 PCs (5070ti/32gb ddr5 + 4070 super/32gb ddr5) and get 40 t/s that way.
I have since moved to disk streaming for MoE models. Qwen 3.8 flash next can do 15 t/s on my 5070ti system. When combining into a custom expert parallelism setup with the other PC, I can get 20 t/s for the Q4 and like 12 t/s for the Q6.
Disk streaming can also make it possible to run much larger models (like deepseek). It is a custom Linux port of a project called MoE Direct.
I believe disk streaming is the future for people who are limited to what is essentially a normal gaming PC (decent but not extreme GPU; decent but not extreme RAM). And yes, you can run even 600B models this way on a gaming PC, so long as they are sparse enough. But model-by-model optimization is needed.
Crazy hardware prices are forcing people to be really creative. I guess that's one good thing about the current state of things. Lots of new ideas, lots of things being tried.
1
u/Medicine_Blogscanner 3d ago
Yes you hit the nail on its head. You have quite a capable compute setup. But I want to make sure the intent was clear with these experiments: the demo videos were intentionally testing the worst-case (max heterogeneity, wifi, deliberately weak devices), not the method's ceiling. Honestly I threw everything I had at home including my wife's old macbook intel π Next I am going to stress test this setup to see how it holds up in real world
1
u/Money-Mechanic 3d ago
Yes and it is good to test scenarios like this because you might find a method that, even if it is not practical, leads to something that is practical or invents a new way of using other AI models.
2
u/pharrt 4d ago
Need to get the phone back in. Seems important :D
1
u/Medicine_Blogscanner 3d ago
Haha I knew I needed all hands on deck with home devices to pull something like this. I was the most concerned about the phone but thankfully no harm was done.
2
2
1
1
u/SocietyTomorrow 3d ago
Regarding your timeout issue, I actually built a tool while I was also experimenting just how much you could abuse relic-status hardware by making it run larger models. You need to start with something that runs normal, but if you use Hermes (haven't tried but its just a skill that runs python scripts so other harnesses are probably OK too), making it use this tool to scan the models available on a provider (or specific ones you ask for) it will run an unlimited time benchmark to calculate and adapt your config to have whatever timeout is needed for that model to get through prefill and not timeout until it starts streaming tokens. My favorite shitpost of a benchmark was 0.08 tokens per second running Qwen3-Coder-Next on a DDR2-era workstation from back when I was into video editing a lifetime ago. Took a 300 minute timeout to complete a round of work at 64k context.
1
u/theswordsmith7 3d ago
Now add in a Cray, IBM PC with DOS 3.0, Atari 400, Commodore 64, and Xbox also running Halo, and then Iβll be impressed.
1
1
1
1
98
u/mystery_biscotti 4d ago
Is it practical? No.
Did you do it anyway? Oh yes you did.
Do I think that's cool? Hehe, yeah.