Rendered at 16:25:42 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Pannoniae 4 hours ago [-]
Nice article :) Yeah this is basically a tradeoff between CPU and GPU power. There are different types of culling. There's the basic stuff like backface culling (supported in hardware, don't render triangles facing away from you) and frustum culling (don't render objects which your camera doesn't see). These are used in just about every game.
For occlusion culling it's a bit more tricky because you can either do it on the CPU in broadly two ways, either do low-res raycasting / software rendering like in the article on the CPU and cull based on that. This is an adaptive workload, you can give it more threads or CPU power and it scales for better culling which results in less pixels rendered on the GPU.
You can also use GPU culling but that's more complicated to do and that uses the GPU which creates a catch-22 - you want to use GPU culling to reduce GPU load but integrated GPUs don't cope well with compute shaders and memory bandwidth in general, so doing a culling pass might wipe out any culling gains you might have.
And dedicated GPUs have the raw power and memory bandwidth to just submit everything in your frustum and get most of it depth-rejected.
I wonder if it would be even faster to create a connectivity graph on the CPU, like each chunk knows whether a neighbour is visible and vice versa. On rendering the chunk graph is walked and the visible chunks are submitted, kind of like a primitive garbage collector to determine liveness. The culling would be worse but I presume traversing a fairly small (few thousand elements) list is quite a bit cheaper than rendering the "mipped" occlusion boxes, but do let me know if this is wrong.
avaer 26 minutes ago [-]
If you can turn the problem into a small kernel operating on a heap of data (or a hierarchy), the GPU almost always wins for culling, especially if you pipe the cull into the draw with GPU-driven rendering.
If the author used GPU culling it would likely be faster on modern hardware, they just said they can't because of platform restrictions. But that's what AAA games do.
The modern non-nanite techniques here are basically to regularize to grids, cull the grids on the GPU, and then cull more with a HZB. That's your "mipped occlusion boxes", except it's actually very cheap to do this because you're reusing depth you already had from previous frames, the test for each object just a few texture samples, and it all stays on the GPU. As a bonus, you can do your LODs on the GPU at the same time, saving even more CPU work and memory bandwidth.
Also, depth rejection is not going to help with the problems that culling solves; draw, vertex processing, raster, and then pixel tests is much more expensive than a cull test before doing any of this. And if you're forward rendering w/expensive fragment shader, the overdraw of relying on the Z buffer to do your "culling" can kill you.
Pannoniae 18 minutes ago [-]
>that's what AAA games do
yeah where the pixel work and the vert count is magnitudes higher. I forgot to mention in my comment that I was talking about low-poly / pixel stuff like this, with a very simple PS and being pretty much API bound or memory-bound in perf
tranceylc 17 minutes ago [-]
I’ve been having a lot of fun implementing most of these on the GPU, and being okay with saying “sorry” to integrated graphics.
The pure bandwidth on a GPU is insane. 1000fps with infinite render distance.
corbinvachal 1 hours ago [-]
[dead]
keyle 6 hours ago [-]
Wow a tech article on HN. What is happening. Are we back in 2022? /s
fullstackwife 6 hours ago [-]
January 2026, thats light years in AI psychosis time units, thats before harness, and auto hill climbing era
For occlusion culling it's a bit more tricky because you can either do it on the CPU in broadly two ways, either do low-res raycasting / software rendering like in the article on the CPU and cull based on that. This is an adaptive workload, you can give it more threads or CPU power and it scales for better culling which results in less pixels rendered on the GPU.
You can also use GPU culling but that's more complicated to do and that uses the GPU which creates a catch-22 - you want to use GPU culling to reduce GPU load but integrated GPUs don't cope well with compute shaders and memory bandwidth in general, so doing a culling pass might wipe out any culling gains you might have.
And dedicated GPUs have the raw power and memory bandwidth to just submit everything in your frustum and get most of it depth-rejected.
I wonder if it would be even faster to create a connectivity graph on the CPU, like each chunk knows whether a neighbour is visible and vice versa. On rendering the chunk graph is walked and the visible chunks are submitted, kind of like a primitive garbage collector to determine liveness. The culling would be worse but I presume traversing a fairly small (few thousand elements) list is quite a bit cheaper than rendering the "mipped" occlusion boxes, but do let me know if this is wrong.
If the author used GPU culling it would likely be faster on modern hardware, they just said they can't because of platform restrictions. But that's what AAA games do.
The modern non-nanite techniques here are basically to regularize to grids, cull the grids on the GPU, and then cull more with a HZB. That's your "mipped occlusion boxes", except it's actually very cheap to do this because you're reusing depth you already had from previous frames, the test for each object just a few texture samples, and it all stays on the GPU. As a bonus, you can do your LODs on the GPU at the same time, saving even more CPU work and memory bandwidth.
Also, depth rejection is not going to help with the problems that culling solves; draw, vertex processing, raster, and then pixel tests is much more expensive than a cull test before doing any of this. And if you're forward rendering w/expensive fragment shader, the overdraw of relying on the Z buffer to do your "culling" can kill you.
yeah where the pixel work and the vert count is magnitudes higher. I forgot to mention in my comment that I was talking about low-poly / pixel stuff like this, with a very simple PS and being pretty much API bound or memory-bound in perf
The pure bandwidth on a GPU is insane. 1000fps with infinite render distance.