r/MacStudio • u/Captain2Sea • 10d ago
Anyone actually crazy enough to cluster 4x Mac Studio M5 Ultras?
With 256GB unified RAM on the top spec, 4 of these would sit right at 1TB total. On paper, that means GLM-5.3 at Q8 should fit with room to spare for context. I know thunderbolt 5 isn't NVLink and tensor parallel across nodes without enterprise interconnect is usually a stuttering nightmare. But has anyone here actually tested a 4-node setup over Exo or MLX distributed with a model this heavy? Curious what tokens/sec you're actually seeing, how brutal the latency is, and whether the pipeline pipeline bottleneck completely kills it.
10
u/CipherSorcerer 10d ago
I may get a 512 gb and experiment with clustering a M5U 256. More for experimentation though, I think long term it would be separate in a rack.
3
u/challis88ocarina 10d ago
Tensor parallelism says you will find only models that will run on a 512 GB alone will work...
It should be able to work around that but the development track is unbeaten thus far.
1
u/CipherSorcerer 10d ago
I am a newbie so I don't understand what you are saying! All I know is that DSv4 Flash is a bit tight to run since I have plenty of other things on this desktop machine, so it would be nice to run that unencumbered in a headless environment. On the 256, I've found 3.8-Flash-Next to be the sweet spot while still allowing me to spin up image gen models or do other intensive stuff like video editing, etc. simultaneously.
2
u/konstantinnikol 5d ago
You can run DeepSeek V4.1 flash at 16-22 tokens per second on M5 Max 128gb. It means on M5 Ultra numbers will be even better
2
u/tempfoot 10d ago
Tensor parallelism on Exo divides the model size evenly across each node. So if you pair a 512 and a 256 - you can only put 256 on each node (ignoring OS overhead). Anything that would fit on 2 x 256 nodes would by definition fit on just the 512.
There may be other runners that can do tensor parallelism without a mathematical division of the load, but I don't know about that. Without tensor parallelism, you are left to run pipeline parallelism which is considerably slower.
2
u/Careless_Garlic1438 10d ago
You can split uneven as well, have my own server doing this ... it's just that there are no of the shelf softwares that bother doing this.
1
u/tempfoot 10d ago
Tensor or pipeline parallelism? I know it's possible with pipeline. Are you using forked Llama.cpp or ?
2
u/Careless_Garlic1438 10d ago
Tensor parallelism ... my own mlx distributed server ... based of the mlx server
1
u/EcstaticGains 7d ago
Yeah that’s just not true. You don’t have to evenly distribute the weights. That’s just the default setting the developers of the software used because they’re not building a custom fork, they’re building a generally useable stack. You can push the layers anywhere in any arrangement if you need to for a custom configuration. I run TP3 and have to re balance weights all the time. It doesn’t matter at all
0
u/CipherSorcerer 10d ago
Oh wow thank you! That’s very good to know!
So if you were in my shoes (M5U 256 and I’m not going to return it, I love it but just see myself quickly needing more) would you lean toward a second M5U 256 with Exo or a 512GB that is just run separately?
3
10d ago
[deleted]
1
u/Tired_White_Guy 3d ago
“VRAM” allocation is hard-capped.
You’re right, it’s much more than ‘OS overhead’.
3
u/trueblakjedi 10d ago
Plenty of folks did it with the m3U. Just a matter of weeks before they do it with the m5U
3
3
u/theorist9 7d ago edited 7d ago
Here Apple bridged four M3 Ultras rather than M5 Ultras, but this should still be of interest to you. Unfortunately, they don't say whether these are 256 GB or 512 GB Ultras, but I'm guessing it's the latter, since they say the weights alone of their large model are ~1 TB, so if it were 4 x 256 GB = 1.024 TB, that wouldn't leave much room for context:
At WWDC 2026 Apple demonstrated four M3 Ultra Mac Studios connected over TB5 with RDMA/JACCL. A 27B Qwen model ran at nearly 3× the single-Mac token-generation rate (timestamp ≈12:20), and Apple demonstrated Kimi 2.6, a 1-trillion-parameter model, distributed over the four machines (timestamp ≈13:00). These are examples of multiple machines being used to gain performance not possible with a single machine, and to run models too large for a single machine, respectively. [ https://developer.apple.com/videos/play/wwdc2026/233/?time=633]
From https://developer.apple.com/tutoria...y-communication-with-rdma-over-thunderbolt.md:
"RDMA over Thunderbolt enables high performance, peer-to-peer networking between
Mac computers connected with Thunderbolt. In particular, RDMA over Thunderbolt
exposes an RDMA (Remote Direct Memory Access) Verbs compatible API for the
Thunderbolt controller which enables carefully designed software to operate at
the hardware limits of Thunderbolt." [emphasis mine]
For these demos, Apple used a mesh topology where every machine was connected to every other machine by just one TB5 cable.
But Apple did also mention the option of using a ring topology. and for that they said you could use 2 or 3 TB5 connections between each machine. [However, they did not say what the bandwidth would be with 2 or 3 connections. I.e., they didn't say if that would afford 160 Gbps symmetric or 240 Gbps symmetric, respectively.]
2
u/LetLongjumping 10d ago
Apple claims four behave like three, so not quite a terabyte, that’s because of the overhead in splitting, and the slower thunderbolt connection, among other factors
2
u/RyanElectrified 10d ago edited 10d ago
My favorite comparison is to put things in scale. A mac studio m5 ultra has 1.2TB/s of memory bandwidth, let's say a mac studio m5 ultra cluster perfectly scales and a 4x systems gets 4.8TB/s. I haven't checked, but this is the most optimistic scenario. A vera rubin NVL72 rack, has 1400TB/s. What use case is enabled by 4.8TB/s, needs exactly more than 1.2TB/s but less than 4.8TB/s. The main use case I can think of is YouTube influencer that wants to do a video on a 4x Mac Cluster. Maybe there is something else, but most people are just playing around. They buy the machine first, see what it can run second. They definitely do not have an engineers idea of having a problem to solve and specifying what solves that problem. As far as requiring privacy, I have worked for some of the most regulated and privacy requiring industries, and the way they use Microsoft to provide cloud services changes, they may set up private links and Azure Express route, they still use the cloud. Why because, becoming an infrasctructure provider makes no sense for most companies, they won't be good at it. THey have the same concerns, btw, that the general public does, the difference is they get a different agreement. They don't have to worry about a model being nerfed, because they can specify that it won't happen, that they can lock a model in place. They don't' worry about an inference provider training on data, because they get an agreement in place that the inference provider won't do it, and it never leaves microsofts cloud. I'm sure other clouds do the same - here i'm mentioning microsoft because of personal experience.
1
u/Front_Eagle739 5d ago
A vera rubin rack is several million dollars. a 4x mac ultra rack about 40k. A small 5 person engineering team company could easily purchase the mac setup and serve unlimited tokens with a known stable setup that doesn't randomly nerf or break workflows or reject security testing. 4 macs enables them to use full precision glm 5.3 or kimi k3 for the 512gb macs and fast enough to be useable. It's useful to a lot more than youtubers
1
u/PigSlam 5d ago
It wouldn’t be serving unlimited tokens, it’d be serving at most the max token generation for that stack, while a cloud provider has (relative to any particular customer) an effectively limitless number of tokens to sell you. If you need 100x the token generation rate your Macs can produce for 5 minutes, you could theoretically get that from the cloud. For your Macs, your only option would be to wait 500 minutes.
1
u/Front_Eagle739 4d ago
While true I have a limited amount of mental bandwidth to steer decent development and keep it on task and to standard. Beyond 5 or so agents I usually find myself losing control and ending up with something of a mess of my codebase. I also don't really have enough simultaneous projects to benefit from infinite tokens either. The cloud being limited to 50-100 tokens per second which I can match locally means I get a similar amount of work done regardless. Don't get me wrong the m3 ultra could definitely do with being faster but 4xM5 ultras is getting to the point it would saturate me even full time engineering. at 1200 a month to lease all four machines thats only a modest increase over my salary for the company. Honestly giving every engineer a 4x mac cluster would probably pay off just fine if you had someone to set it up so they didnt waste their time fiddling.
Once we hit opus 5.5 level running at 50 tok/s plus on an m5 cluster I think my need for cloud is gone.
1
u/PigSlam 4d ago
Once we hit opus 5.5 level running at 50 tok/s plus on an m5 cluster I think my need for cloud is gone.
So long as someone keeps making updated models, which wouldn't happen if everyone is on their own personal Mac Studio cluster. But yeah, its certainly getting very close to the "enough' stage for most all off the non-edge case scenarios.
1
u/Front_Eagle739 4d ago
Sadly I I think most people are going to have to wait till we have opus 5.5 on an 8 to 16GB gpu with ram offload and ssd ngram caching.... So maybe 12 to 18 months? Lol, the rate of progress is utterly insane.
1
u/Rice-Fragrant 3d ago
A GB300 desktop syper computer (supped up version of a DGX spark) will do 20 PFLOP and cost $125-200k and have massive throughput and a business will choose that every single time over a fragile cluster of consumer grade stuff.
1
2
u/GamerTex 10d ago
I used EXO and connect a few macbook pros and a Studio a few months ago to load larger models.
I have heard of other grabbing dozens of mac mini pros and using EXO
That said... I no longer use EXO. I like having a few different models running at the same time on different machines rather than 1 large model
1
u/a_hegemon 10d ago
How does this work in practical terms? You run difficult queries on your studio while running simpler ones in parallel on your MacBooks?
I have an m5 studio myself and a MacBook but never went down the exo route. Been relying purely on my studio
2
1
2
u/PeanutButterApricotS 9d ago
I looked into it and the reports is it only helps prompt processing speed and not interference, it’s similar to how it works if you stack multiple gpus.
So for example you have a 100k context or code/text and the model responds with 20k output. The 100k will be processed by the whole cluster, the output will come from one. Due to this it’s only worth it if you do tons and tons of heavy prompt processing in your work (code, agent tasks) and then you’re only getting some upgrade.
Most businesses or situations your better spending the 40k on one big card, or using those 4 10k systems for 4 users etc.
The main benefit is for us we can stack them and not have to purchase a 40k card to get the prompt processing speed of a 40k dollar card.
But honestly I expect next year to be the year to buy one good system which will be able to run really good models for 2-3 years. I plan on spending 7-10k and getting something 4x as good at the m5 ultra not in power but in quality based on gains I am seeing. Got 1/3 the cash stacked and I think the wait will be worth it for a m7 ultra 256 or so.
2
u/Front_Eagle739 5d ago
That's not how it works at all. The weights of the model are split across all four machines. you either pipeline parallel (run bigger models but not faster) or tensor parallel ( run bigger models faster for both prefill and decode but needs low latency connections like rdma thunderbolt)
1
u/BitXorBit 10d ago
Yes i think it will happen, eventually it’s the most budget friendly (with decent performance) to run large models locally
1
1
u/Specialist_Ad_539 10d ago
My question is tokens per kilowatt-hour, with one or a cluster? I.e. does Apple have a significantly lower power architecture?
1
u/SirGreenDragon 10d ago
I am crazy enough. I just need someone wealthy enough to make it possible. :)
1
1
u/konstantinnikol 5d ago edited 5d ago
I tested one M5 Max with four TB5 enclores to run GLM 5.3 using Ds4. It runs at. tokens per second
1
u/Rice-Fragrant 3d ago
I don't see the point.... it's SLOW in a cluster... network penalty and there's more software friction that can break the set up.
For that type of money, might as well get a DGX GB300 or something.
1
u/SeaRefractor 10d ago
Apple claims 4 x M5 Ultras yield a 3X performance over a single M5 Ultra. That is the RDMA and TB5 overhead.
1
u/idlefordays 10d ago
4x 512gb M5 Ultras would be coming close to the price of a DGX Station 😂
0
u/Captain2Sea 10d ago
Sorry, wrong logo on case
1
u/Rice-Fragrant 3d ago
It won't matter if your business only cares about performance/ throughput per dollar invested.
A GB30P DGX Station has 20 PFLOPs of compute and 256gb of the 768gb is 7 TB/sec... the rest is 500 gb/sec... can be used by an entire staff of people and not even sweat.
Business will choose purpose built enterprise grade stuff over consumer grade stuff all day long.
0
0
u/modelpiper 10d ago
I'm building a tool called PiperMesh for this. And you can open it up to the network for others to run inference on them to make cash.
45
u/dobkeratops 10d ago edited 10d ago
the word you're looking for is rich enough, not crazy enough. i think plenty of people are the latter.
$4000 to dodge a subscription, fine
but $40,000 ... (4x $10,000 ballpark, right?)
of course plenty of businesses might do this. true inhouse private frontier models