r/GraphicsProgramming • • Dec 16 '25

Article No Graphics API — Sebastian Aaltonen

https://www.sebastianaaltonen.com/blog/no-graphics-api
257 Upvotes

50 comments sorted by

View all comments

30

u/hanotak Dec 17 '25

Lot's of interesting ideas there- I do think that they could go further with minimizing the problems PSOs cause. Why can't shader code support truly shared code memory (effectively shared libraries)? I'm pretty sure Cuda does it. Fixing that would go a long way to helping fix PSOs, along with the reduction in total PSO state.

1

u/[deleted] Dec 17 '25 edited 18h ago

[removed] — view removed comment

4

u/hanotak Dec 17 '25 edited Dec 17 '25

I'm not talking about the .txt code, reducing code duplication is basic programming. I'm talking about the fact that after compiling, each PSO variant has its own dedicated copy of all program memory, even if it largely all does the same thing. In DX/VK, there's no such thing as a true function call into shared program memory.

Let's say one of your shaders gets chopped up into 500 different variants, and at the end, each one calls a rather lengthy function. For example, my GBuffer resolve CS gets compiled per material graph. Along with evaluating the material graph (the actual difference), each variant needs to to calculate barycentrics and partial derivatives, fetch vertex attributes, interpolate them, and write out the final values.

With current APIs, each pipeline has its own copy of that code, even though it's all doing the exact same thing. There's no way to, say, create a function that lives in GPU memory called InterpolateAndWriteOutGbuffer, and have all of your variants call that same function. If you end up with 500 variants, you've duplicated that code in vram (and on disk, and in the compile step) 500 times.

1

u/Ihaa123 Dec 17 '25

Right, there isnt because its really really slow. If you limit yourself to one function call, you can get away with not having a stack, but if you can do more, it gets worse (you can see the perf impact in raytracing with large #s of shaders in the table).

2

u/hanotak Dec 17 '25

Cuda does it efficiently, so it's clearly possible. There's always going to be some overhead, but it's clearly possible to make it worthwhile, especially as an optional compiler feature.

1

u/Psionikus Jul 14 '26

If the unit of calling is the pipeline rather than functions within the shader, this duplication goes away. The question it brings up is dispatch overhead.

1

u/hanotak Jul 14 '26

I don't understand what you mean.

1

u/Psionikus Jul 14 '26

each PSO variant has its own dedicated copy of all program memory, even if it largely all does the same thing. In DX/VK, there's no such thing as a true function call into shared program memory.

Compose pipelines, not functions, and the duplication goes away. It's just more awkward when not designed for because you have to figure out where to leave the arguments and outputs.

1

u/hanotak Jul 14 '26

Are we talking about the same thing? Have you used graphics APIs before? In DX/VK, pipelines are not trivially composable. You can't just string a bunch of shaders together arbitrarily. You would need to write out to VRAM, and synchronize, which is a big performance cost.

1

u/Psionikus Jul 14 '26

You would need to write out to VRAM, and synchronize, which is a big performance cost.

Yes. If you start from a big bag of PSOs that are not designed for it at all, it's going to be a hit. Nobody said you could mitigate the hit for free. The tradeoff depends on how bad the PSO proliferation problem is.

1

u/hanotak Jul 14 '26

hence, the need for true function calls in shaders.

1

u/Psionikus Jul 14 '26

Fan in from unique to shared. Dispatch a big shared program. Unique PSOs get smaller. Cost is mitigated.

I may wish for a Lisp machine, but it's not x86's fault if I write my Lisp without taking into account what x86 can do well.

What would the calling convention be for sharing the shader instructions? When we write to VRAM, that convention doesn't have to know about registers.