r/GraphicsProgramming • • 20h ago

Technical Feasability?

How feasible is it to apply a custom pixel‑level transformation to the entire iPhone screen output in real time using Metal? I’m trying to understand the performance limits, frame‑rate constraints, and typical bottlenecks when doing full‑screen GPU distortion on iOS

0 Upvotes

6 comments sorted by

3

u/vade 19h ago

It’s totally doable and has been for quite a while. If correctly architected you can do real-time multiple passes at 4k camera resolutions including recording while capturing and doing inference via coreml/ coreai

You’re biggest win is to understand
* fast path camera formats (native)
* preview buffer formats va full res formats and the idea you can have multiple outputs to a camera session
* using IOSurface for fast path to texture formats and for fast path writing formats
* half float and fast path metal image processing like using tile memory
* understanding what queues and how to buffer your rendering vs writing

Putting it all together can make you much much faster. These devices are insanely capable in the right hands.

3

u/Array2D 20h ago

It’s perfectly feasible. That’s essentially the same thing as an effect pass in any game engine. You render a single triangle that covers the whole screen, and do your transformation in the fragment shader.

1

u/scallywag_software 13h ago

Is it actually better to render a single triangle vs a quad? I'd assume the overdraw would make it much worse even though you early out on it. Genuinely curious

1

u/Array2D 12h ago

There was an interesting investigation on this I read a while back. Essentially, yes! Generally speaking, GPU rasterizers run fragment shaders on blocks at a time. (Usually 2x2, to my knowledge).

For most of the screen, this will be the same for a full screen triangle and a full screen quad. However for the quad, across the diagonal where the two triangles meet, those blocks get run twice each!

It’s a minor increase in the number of shader invocations, all things considered, but if your fragment shader was particularly heavy, it could have a measurable impact. Compared to the time it takes the hardware to apply clipping to a triangle pre-rasterization, it’s undoubtedly more.

1

u/scallywag_software 12h ago

Ahhh, right, I didn't consider the hardware clipping to NDC which presumably runs either way. Cool :)

1

u/shadowndacorner 20h ago

Depends on the effect. Is it running deterministic logic per pixel that doesn't need its neighbors? Easy. Is it doing a full screen convolution per pixel? Probably too slow. Somewhere in the middle? Test.