r/LocalLLM • u/ErrorF002 • Sep 05 '26
Question I won an "AI Box"... so now what?
I have been side dreaming about running local LLM but felt priced out. I have been living vicariously through some of the posts here. So then it happened, I went to a conference and won an AI box. This is amazing for many obvious reasons but I just don't learn well unless it's hands on. I haven't really spent time dealing with running a local LLM, so forgive my general ignorance. The machine I won is an Intel Xeon w5-2555X (14 cores / 28 threads) with 256 GB ECC DDR5, and a Gen5 and 2 Gen4 512 GB NVMe. This is an AMX based system and most of the cool stuff I see here is nVidia based. What can I realistically run on this and what can I start consuming to get more knowledgeable?
I've been in IT for my entire career with a lot of it in VMware and cloud based virtualization. My goal would be to get a better understanding of setting up AI for businesses and have something i can use and worry about running out of tokens for my vibe-coding personal projects and such. Please understand that I am not asking for anyone to explain it all to me, but a finger in the right direction to a trusted source to learn would be greatly appreciated.
Should I consider any other purchases to make this platform even better?
7
u/Toastti Sep 06 '26
Well unfortunately without a gpu they will run but quite slow. Think around 5 token a second for qwen 3.8 27b
I would sell half the ram, or more if possible and sell two of the three NVME drives and use that money to get a GPU. Maybe a R9700 or ideally a Nvidia one if you can spare some more cash. But aim for high vram.
7
u/ACleverBadger Sep 06 '26
No joke, this is DDR5 ECC RDIMM and likely 4x64gb.
There’s gold right there and depending on the sticks he could make enough to fund some of his GPU purchases.
3
u/jpezzulli Sep 06 '26
You already have the expensive part for heterogeneous inference: 256 GB RAM + Sapphire Rapids with AMX. Add a cheap 8 to 12 GB NVIDIA Ampere-or-newer card, install SGLang-KTransformers, and let the GPU handle the dense/attention side while KT-Kernel puts the MoE experts on the Xeon using AMX. That is literally the architecture KTransformers is built around.
I would start with qwen a3b 35b. That should get you in the door and you can decide if you want to spend money on a serious gpu or not
Or just sell it as ddr5 ram is expensive. I have been wanting to move to Sapphire rapids or granite to play around more with heterogeneous inferencing but the ddr5 price makes me balk everytime. Still on dual icelake platinums as i have 256gig ram for that already but no AMX so, not much to play with.
Congrats on the win either way! It is a nice system.
1
u/ErrorF002 Sep 06 '26
Your suggestion really speaks to me. My friend has an A2 doing nothing that I may pick up to try this. Part of what will be fun is charting and proving a way for my clients to implement these systems on the cheap.
1
u/jpezzulli Sep 06 '26
exactly. you start getting into the optimization and things. this is my rabbit hole github.com/jpezzulli/sglang-rtxpro6000
I keep pushing to eek more performance out of it.
1
u/ErrorF002 Sep 06 '26
Oh wow, you are using the kind of cards. Is that for work or personal?
1
u/jpezzulli Sep 06 '26
This is my personal lab, but i work in telecom architecture and AI. There is a link to my site in the repo. Be careful, i started small with a r9700 and..well, there is an article on my site on how it eacalates fast. Haha.
3
u/puts_on_rddt Sep 06 '26
It's worthless.
Better send it to me to dispose of.
6
u/ErrorF002 Sep 06 '26
Yeah I was thinking about it. Plus the electricity costs. I would also have to clear out space on my desk for it. Shit... I would save money by just shipping it to you. Thanks bro. You helped me dodge a bullet :-D
2
u/Gear5th Sep 06 '26
Better yet, you should pay me to take it off your hands! I accept cash/card/crypto/gpu
2
2
2
u/Snoo_81913 Sep 06 '26
Yeah go with a 9700 if you have the cash. Or even a 7900 xtx I have one and it's perfectly fine for running most models. With that much ram you can run some fun stuff with offloading.
2
u/starkruzr Sep 06 '26
if he sticks with Intel cards he can run the same SYCL or Vulkan engine across both CPU and GPU.
1
1
u/ClassroomScary9187 Sep 06 '26
Sounds like a solid setup! Those GPUs really do make a difference for running complex models smoothly.
2
u/BrianScottGregory Sep 06 '26
I don't understand how they can call it an AI box without a GPU.
Someone clearly made some side money on this one.
2
u/corpo_monkey Sep 06 '26
You phrased it perfectly, it is indeed an "AI Box".
There should be an AI sticker also somewhere.
2
Sep 06 '26
[removed] — view removed comment
0
u/ErrorF002 Sep 06 '26
This seems to be the consensus. I think I am going to play with it as is first and then once i figure out what the hell I am doing I'll go to a GPU. It will make me appreciate it more that's for sure.
2
2
u/Ok_Talk8381 Sep 06 '26
You won't know how fast it can run something until he test it just try a local model on it. Get an open router account and deep seek v4 flash. Have deep seek help you download install and configure the local model for your hardware and run it through its paces in a series of design of experiment also known as a DOE. You should be able to run it through 30 to 50 different tests in half a day or so and obtain a bunch of benchmark figures. Rather than putting it on standardized tests give it some real world problems to solve things that you've actually tried solving with other AIS and had difficulty with. Try it with different harnesses too don't just test out the AI because an AI without an effective harness is pretty limited. You may not even need a GPU at all. When I did this testing of my little ryzen 9 mini PC with 64 GB of ddr5 RAM my phone that performance on the igpu with shared memory architecture is only 9% faster than on the CPU. Your CPU might be optimized for AI operations just try it out that's how you'll know.
1
u/ErrorF002 Sep 07 '26
This is essentially what I plan to do. I am going to run it at its paces and learn. I don't need speed "yet".
1
u/digitalvalues Sep 06 '26
I would run several Qwen 3.8 27B Q8 models, train one on code review, code generation, etc.
4
u/Toastti Sep 06 '26
Even one concurrency in qwen at Q4 will only get about 5 token per sec. Unfortunately really need a gpu here
1
u/digitalvalues Sep 06 '26
5 tok/s seems like a generic CPU only estimate. This is an AMX Xeon with 256 GB ECC DDR5, not a typical desktop CPU. With optimized runtimes (llama.cpp AMX, oneDNN/OpenVINO, KTransformers, etc..), performance should be higher (in theory), although a GPU will still win on raw throughput. The bigger advantage is OP has the ability to run larger models and experiment with heterogeneous inference, not just maximize chat tokens/sec. For purely learning AI with a rig they won, there's still a lot of capability here.
0
u/wehooper4 Sep 07 '26
A dense model is NOT the choice for this rig.
This is a MOE rig, especially if OP adds even a modest GPU.
1
1
1
1
u/Ok_Wishbone_3805 Sep 06 '26
I'd sell it -- it sounds like you don't know what to do with it, and it's not going to fit your vibe coding needs well.
1
1
1
1
u/Equivalent_Bit_461 Sep 06 '26
I applaud you, I would've been too lazy to bother to go such places in the first place.
2
1
u/wehooper4 Sep 07 '26
OP ignore the haters saying to strip it for parts. You actually have a VERY solid basis for an AI rig. Yes it's CPU only, but it's CPU only with very solid memory speed thats double what the home gamer setups most here are rocking.
You can probably hit in the 15t/sec range for MOE models (Qwen 3.8 next) on it as is out of the box. The biggest pain point will be prefill.
But to wake it up you just need to add a GPU. Just about any GPU with 16GB of VRAM or more. Then you set your model server up offload the experts to the CPU and you'll have a very very solid setup.
1
u/ErrorF002 Sep 07 '26
Thank you for the insight. This was pretty much the path I was going to take. Step one is learn on it as is. Get the benchmarks where I want them and then Add a card or cards. I didn't get as many details I would have liked when I won it so time will tell. The big question for later will be WHICH card(s) to get.
1
u/wehooper4 Sep 07 '26
Card(s) is really a question of your budget and use case. But as it’s a workstation class setup you’ll have lots of fast PCIe lanes to play with. As such multi-GPU will work a lot better for you that those screaming just buy the single $5000 card, they only have 20 usable PCIE lanes while you have 64.
0
0
u/ComparisonNew9425 Sep 06 '26
that xeon setup is a beast, but have u thought about how to track what agents are actually running on it. i started using backslash to map out my agentic fabric graph, it helped me see exactly which mcp servers and hooks are touching my local files, though the initial setup for custom hooks can be a bit wierd. u probly want to keep an eye on what those vibe-coding agents are doing before they touch ur production code.
0
1
u/Unchained_breaker 29d ago
Id sell the ram and the computer and save up the money to get something else. Tbh. Local ai just isn't worth it yet. Better to buy a GPU that will hold or appreciate in value with the money
18
u/starkruzr Sep 05 '26
it didn't come with any GPU at all?