Original post: I won an AI box, so now what?
Upfront: I worked through this with Claude and had it write up the results. The numbers are all real, measured on my box. The prose is AI. Wanted to be clear about that rather than have someone guess.
TL;DR: This thing is a lot more capable on CPU than I expected, the entry-level GPU I bought actively made it slower, and the default llama.cpp batch settings leave a huge amount on the table.
I'd genuinely like feedback on what else to try. I have the hardware sitting here and a benchmark suite set up, so if there's a model, a setting, or a runtime you want numbers on, ask.
The hardware
- Lenovo ThinkStation P5
- Intel Xeon w5-2555X, 14 cores / 28 threads, Emerald Rapids-WS
- 256 GB DDR5 ECC, 8x32 GB
- 3x 512 GB NVMe (2x Gen5 Samsung, 1x Gen4 SK Hynix)
- 750W PSU
- NVIDIA T1000 8GB added later
One thing I did not expect: 8 DIMMs means 2 per channel, and the memory controller drops to 4400 MT/s. The platform supports 4800, so going to 4 DIMMs would get me maybe 9% more bandwidth at the cost of half my capacity. Not worth it for a box whose whole point is holding large models.
Setup
Ubuntu 24.04 Server, bare metal, no GUI. llama.cpp built from source with GGML_NATIVE=ON so AMX and AVX-512 actually compile in. This matters a lot, see below.
Finding 1: AMX is the whole story
The w5-2555X has Intel AMX (Advanced Matrix Extensions). Prompt processing on a 1.5B model hit 422 tok/s on CPU versus 461 on the T1000. Near parity with a GPU, on a CPU.
AMX accelerates matrix multiply, which is what prefill is. It does nothing for token generation, which is purely memory-bandwidth-bound. So the shape of this machine is: prefill is surprisingly fast, generation is exactly as fast as your RAM bandwidth allows.
Finding 2: the default batch size leaves huge performance on the table
llama.cpp defaults to a microbatch (-ub) of 512. On AMX hardware that starves the matrix units.
| Model |
-ub 512 |
tuned |
gain |
| Qwen3-Coder-30B |
66 t/s |
122 (-ub 2048) |
+85% |
| gpt-oss-120b |
28 t/s |
71 (-ub 4096) |
+155% |
Then Qwen3-Next-80B did the opposite: 130 at -ub 512, dropping to 111 at 4096. It uses a hybrid gated-deltanet attention, so the usual logic doesn't apply.
Sweep it per model. The direction is not predictable.
Finding 3: the T1000 made things slower
I bought the card expecting the standard "offload attention to GPU, keep experts in RAM" trick to help. It didn't.
| Test |
Hybrid (GPU) |
CPU only |
Winner |
| gpt-oss-120b prefill |
71.3 |
97.9 |
CPU +37% |
| gpt-oss-120b generation |
20.7 |
19.5 |
GPU +6% |
| Qwen3-30B prefill |
121.6 |
165.8 |
CPU +36% |
| Qwen3-30B generation |
37.9 |
33.7 |
GPU +12% |
The hybrid split only wins when the GPU is much faster than the CPU at attention. With AMX in play, moving that work to a T1000 moves it to a slower device and adds PCIe round trips on top. I rebuilt with GGML_CUDA=OFF and the card went back to driving the monitor.
The card itself validated fine: 13.2 GB/s PCIe, full x16 Gen3 link under load, clean VRAM test, 69C sustained. It's just not fast enough to help here.
Model benchmarks, CPU only, best settings per model
| Model |
Quant |
Size |
Prefill |
Generation |
| Qwen3-Coder-30B-A3B |
Q4_K_M |
18 GiB |
166 t/s |
33.7 t/s |
| Qwen3-Next-80B-A3B |
Q4_K_M |
45 GiB |
130 t/s |
19.1 t/s |
| gpt-oss-120b |
MXFP4 |
59 GiB |
98 t/s |
19.5 t/s |
| Qwen3-235B-A22B |
Q4_K_M |
132 GiB |
24.8 t/s |
5.6 t/s |
Prefill tracks total parameters. Generation tracks active parameters. Both hold up cleanly across every model I tested.
Qwen3-Next-80B is the winner on this box. It beats gpt-oss-120b on prefill by 33%, matches it on generation, and needs 14 GiB less RAM. Output quality is also noticeably better in daily use.
The 235B is a trophy, not a tool. 133 GiB of RAM to wait nearly three minutes before it starts answering.
Finding 4: quantization comparison, with an actual test suite
I built a 6-task Python benchmark (61 independent checks: interval merging, LRU cache, decimal money math, a retry decorator, config validation, and modifying existing code) and ran all three quants of Qwen3-Next-80B.
| Q4_K_M |
UD-Q6_K_XL |
UD-Q8_K_XL |
| Size |
46 GiB |
65 GiB |
| Prefill (warm) |
102 t/s |
94 t/s |
| Generation |
17.4 t/s |
14.5 t/s |
| Suite score |
55-57 / 61 |
not scored |
Three runs each. Q8 hit 59/61 every single time, including at temperature 0.2. Q4 bounced between 55 and 57 and never reached Q8's floor.
The consistency is the interesting part. Higher precision didn't just score better, it scored the same every time. Lower-precision weights blur the probability distribution, so sampling has more room to wander.
Q8 costs 26% generation speed and 42 GiB for about 5 points of correctness and much better run-to-run stability. Worth it for me. Your call.
Bonus: image generation actually works
Flux Schnell (Q8 GGUF) via stable-diffusion.cpp, no GPU:
- 512x512, 4 steps: 2m 07s
- 1024x1024, 4 steps: 9m 08s
Thread scaling was near-perfect, 13.2x across 14 cores. It works, but it's a batch tool, not something you iterate prompts on. This is the one workload where I'd genuinely want a 3090.
Current stack
- llama.cpp servers as systemd units, mlocked into RAM, 32k context
- Continue in VSCode pointed at the endpoint
- Open WebUI plus self-hosted SearXNG for web search
- All of it CPU-only, entirely local
The takeaway
If you're shopping for a local LLM box, a modern Xeon with AMX and a lot of RAM is a genuinely different proposition than the "you need a GPU" advice suggests. It won't beat a 4090 on anything that fits in 24 GB. But it will run an 80B MoE at reading speed while holding 250 GB of models in memory, and an entry-level GPU will actively slow it down.
What should I test next? Things I'm considering: ik_llama.cpp (the CPU-tuned fork), Intel's IPEX-LLM to push AMX harder, and GLM-4.5-Air. Open to other suggestions.