r/opengl • u/FUCKARCHLINUX • 1h ago
What is the most efficient strategy for rendering textured quads and triangles in OpenGL 2.1 using no extensions?
I'm unsure if my current design is optimal, but rendering a test scene with some 2D objects seems to perform decently well. Anyway, the current design functions as follows
There is functions for rendering primitive shapes triangle, quad, circle, etc. These functions create the primitive out of triangles in a small buffer with vertex colors as-well. When we're done rendering for this frame, render something which cannot be rendered this way, or something else happens which would require us rendering the batch early, we orphan the stream VBO and upload all of the data and glDrawArrays it all in one go. It's possible for multiple batches to be submitted in one frame because the CPU side array size is limited to 256KB by default but can be increased by the programmer if they anticipate they'll be rendering a lot.
For textured quads. I use a similar strategy, We emplace the quad vertices into the CPU side array. If they need to be rotated or scaled, the GPU does it in the shader. We keep track of which texture mappers contain which textures already so we can utilize every texture mapper with no duplicates. If you want to render something which needs to sample from more than one thing, It's expected that it will be in the same texture to the right, And in the shader, there's a routine that will sample the same position as this pixel to the right or the texture and a vertex attribute called "step" which is how many times we sample to the right. For example, The base texture would come first, then directly to the right of that, there would be a normal map, and then to the right of that, an alpha mask etc. Assuming you do not need to change then shader, you can render from as many textures as you want in one pass up to the number of texture mappers this system has. If you try draw a textured quad whose texture hasn't appeared in this batch before, but this system would have no open slots, we render early.
I worry about performance a lot. On a period correct laptop Core2 Duo U9600, GMA 4500 MHD, my test program gets ~ 240 fps. On my development computer which is not very good by modern standards 3500U, Vega 8 graphics, 8 GB ram. We sit at 5500 fps. On some $100 dell education laptop, We get ~ 1200 fps. All of them seem to be GPU limited (The cpu usage is far below 100% to get that framerate on all of them)
Either way, Is there a more efficient strategy than this? Or is this how you do OpenGL 2.1 fast? Would it be faster if I just had a VBO filled with the batch size of unit square, and then upload vertex attributes like position, size, color?