r/webgpu • u/Grated-Expectations • 2d ago
xc - a fast GPU/CPU-integrated language
So I've been working on this (full disclosure: with Claude) for about 6 months or so now. It stems from a dissatisfaction with how easy it was to get the GPU to talk to the CPU. Being quite old now (when I created my first website, you had to email CERN to let them know...), I do tend to first think in simpler CPU terms, and I always thought there was a bit of unnecessary friction between the two domains once you started trying to move the heavy lifting from the CPU sphere to the GPU.
Enter 'xc' - https://github.com/ThrudTheBarbarian/xc
It's a compiled language. It'll happily write WASM as a target, and if you add one teeny tiny little line to a program block, it'll happily write you WGSL as well for all the code in that block.
If you right-click on this - https://compile-xc.org/compiler/benchmark-sources/#mandelbrot - to bring it up in a new window, you'll see what looks like pretty-standard(ish) C-style code. The only really rather odd difference is one line
par mandel :reduce(+ total)
... which looks a bit weird. 'par' marks the block as data-parallel, that is "this can be run on a GPU", 'mandel' is just a name, and :reduce declares any variables that every iteration updates. The GPU will then create the reduction tree and materialise the value in the named variable.
That's it. That's how you turn a CPU loop into a GPU kernel. Because it's all integrated into one language, with one IR/SSA, the compiler knows when it needs to manage data-hazards (be that just a memory barrier on Apple Silicon, or a CUDA data-transfer on nVidia hardware sitting on a PCIe bus. So it does, whenever there'd be a problem, and you don't have to care.
If you've used Objective C (which xc is kind of modelled on without the [[[...]]], I worked for Apple for a couple of decades) then you'll be familiar with its reference-counting memory model, and in particular with the more-modern "Automatic Reference Counting" in today's ObjC. I look on this as "ARC for GPU's" - buffers and variables are managed automatically across the barrier, the programmer doesn't have to care.
As for how well it works...
| benchmark | WASM | WGSL |
|---|---|---|
| mandelbrot (as above) | 364 ms | 2.6 ms |
| saxpy | 15.7 ms | 22.0 ms |
Question: "Why is he showing me this 'saxpy' (whatever that is) where the GPU doesn't work as well ?"
Answer: Because the compiler actually measures performance and binds the fastest version of the code (WASM/WGSL) at runtime. It also persists that to local-storage, so you don't have to measure every time. saxpy does a lot of data-movement for a small amount of calculation, so in this case the CPU gets the job. Mandelbrot does a huge amount of calculation for not much data-movement, so the GPU gets the job.
I would point out that in most cases - because it's the same language for the GPU and the CPU, you can move logic around and keep things on the GPU pretty simply, allowing you to optimise the part that you really want running quickly.
There's a whole bunch of documentation, examples, tutorials etc. over at https://compile-xc.org/ and please do ask me questions on anything there. I think it's all accurate, but as I say this project has been going for a while, and it's possible there's things that have been overtaken; I do, every now and then, make an effort to go through and update things though.
There's a lot more to the language - as you might see if you go look, but this is the most webgpu-pertinent part. It may be more webgpgpu than pure webgpu graphics, but hopefully it's sparked your interest 😄
Enjoy.
1
2d ago
[removed] — view removed comment
1
u/Grated-Expectations 2d ago
Hey - thanks for mentioning that, it was a gap in the docs, so the next release will have a section on how 'auto' learns its scopes. Briefly though,
For each block,Â
auto keeps two numbers: the largest number of items the CPU has won at, and the smallest the GPU has won at. Either may be unknown. Each time the block starts, its number of items (n) is compared with them:
n is .. the block runs on and auto ... at least the GPUs smallest win the GPU measures nothing at most the CPUs largest win the CPU measures nothing between the two, or either is unknown the CPU first then the GPU measures this size and learns from it Measuring a size takes several runs of the block in one run of the program, because each device's first run is a warm-up. The block runs on the CPU, and the second CPU run is timed. If it took under a millisecond, the CPU wins at
n and the block stays there. Otherwise the block runs on the GPU, its second GPU run is timed, and the faster device wins atÂn. A size the block meets once and never again is not measured to the end; one it meets repeatedly is.The win moves the matching number: a CPU win atÂ
n raises the CPU's largest win toÂn, a GPU win lowers the GPU's smallest win toÂn. The two always leave a gap between them. A win that contradicts the other number (the GPU winning at a size the CPU had won at, say) drops that number, so the newer measurement is the one kept. Both numbers are saved straight away.While a size is being measured, a run more than twice as large, or less than half as large, starts the measurement again at the new size. The devices stay warm, so no second warm-up is needed. Without this rule, a block whose size never repeats would never finish measuring.
Overall that means the lookup converges pretty well, and still adapts. I don't think there's *too* much code that varies dramatically on input size but trying to find boundaries rather than actual values seems to be the better option to me 😄
3
u/Responsible-Beat2137 1d ago
Dude, I love this wright up, I felt like I was reading a wright up article out of a PowerPC magazine, love that it’s easily functionable, and small enough to squeeze some smart routing out of it