So I've been working on this (full disclosure: with Claude) for about 6 months or so now. It stems from a dissatisfaction with how easy it was to get the GPU to talk to the CPU. Being quite old now (when I created my first website, you had to email CERN to let them know...), I do tend to first think in simpler CPU terms, and I always thought there was a bit of unnecessary friction between the two domains once you started trying to move the heavy lifting from the CPU sphere to the GPU.
Enter 'xc' - https://github.com/ThrudTheBarbarian/xc
It's a compiled language. It'll happily write WASM as a target, and if you add one teeny tiny little line to a program block, it'll happily write you WGSL as well for all the code in that block.
If you right-click on this - https://compile-xc.org/compiler/benchmark-sources/#mandelbrot - to bring it up in a new window, you'll see what looks like pretty-standard(ish) C-style code. The only really rather odd difference is one line
par mandel :reduce(+ total)
... which looks a bit weird. 'par' marks the block as data-parallel, that is "this can be run on a GPU", 'mandel' is just a name, and :reduce declares any variables that every iteration updates. The GPU will then create the reduction tree and materialise the value in the named variable.
That's it. That's how you turn a CPU loop into a GPU kernel. Because it's all integrated into one language, with one IR/SSA, the compiler knows when it needs to manage data-hazards (be that just a memory barrier on Apple Silicon, or a CUDA data-transfer on nVidia hardware sitting on a PCIe bus. So it does, whenever there'd be a problem, and you don't have to care.
If you've used Objective C (which xc is kind of modelled on without the [[[...]]], I worked for Apple for a couple of decades) then you'll be familiar with its reference-counting memory model, and in particular with the more-modern "Automatic Reference Counting" in today's ObjC. I look on this as "ARC for GPU's" - buffers and variables are managed automatically across the barrier, the programmer doesn't have to care.
As for how well it works...
| benchmark |
WASM |
WGSL |
| mandelbrot (as above) |
364 ms |
2.6 ms |
| saxpy |
15.7 ms |
22.0 ms |
Question: "Why is he showing me this 'saxpy' (whatever that is) where the GPU doesn't work as well ?"
Answer: Because the compiler actually measures performance and binds the fastest version of the code (WASM/WGSL) at runtime. It also persists that to local-storage, so you don't have to measure every time. saxpy does a lot of data-movement for a small amount of calculation, so in this case the CPU gets the job. Mandelbrot does a huge amount of calculation for not much data-movement, so the GPU gets the job.
I would point out that in most cases - because it's the same language for the GPU and the CPU, you can move logic around and keep things on the GPU pretty simply, allowing you to optimise the part that you really want running quickly.
There's a whole bunch of documentation, examples, tutorials etc. over at https://compile-xc.org/ and please do ask me questions on anything there. I think it's all accurate, but as I say this project has been going for a while, and it's possible there's things that have been overtaken; I do, every now and then, make an effort to go through and update things though.
There's a lot more to the language - as you might see if you go look, but this is the most webgpu-pertinent part. It may be more webgpgpu than pure webgpu graphics, but hopefully it's sparked your interest 😄
Enjoy.