HN.zip

Software rendering in 500 lines of bare C++

191 points by mpweiher - 32 comments
articulatepang [3 hidden]5 mins ago
I went through this a few months ago in Rust. I wrote all the code by hand, no LLMs. Then I went ahead and added a small "game" on top, plus some special effects like pixelization shaders and chromatic aberration at the edge of a flashlight.

https://github.com/kshitijl/tinyrenderer-rs

if anyone is interested! The repo has lots and lots of in-progress screenshots so you can see the renderer come to life, plus all the hilarious visual bugs along the way.

I learned a lot! My biggest lesson, other than the specifics of how rendering works, was that modern CPUs are really fast: a single-threaded CPU renderer can definitely run an interactive 3D game with some fancy special effects.

0x1ceb00da [3 hidden]5 mins ago
Why does it pull in wgpu if it's a software renderer?
grovesNL [3 hidden]5 mins ago
In this case wgpu is just providing the surface texture for the window that the software rendered pixels are drawn into.
articulatepang [3 hidden]5 mins ago
Yes, exactly this -- it's a fast and convenient way for my code to just write the pixels into a spot in RAM (the CPU's RAM, not VRAM) and have those pixels end up on screen. On modern architectures this nearly always goes through a GPU, so even if you're not using the GPU to accelerate your 3D rendering/math, you gotta deal with it just to put pixels on a screen at the end of the day. So that's what wgpu does for me.
john01dav [3 hidden]5 mins ago
Softbuffer?

(I originally wrote this a few years ago but other people maintain it now)

LoganDark [3 hidden]5 mins ago
I like softbuffer, but it owns its window and doesn't support as many events / integrations as winit.
LoganDark [3 hidden]5 mins ago
You do often need a graphics context to create the surface that displays the framebuffer.
nkanaev [3 hidden]5 mins ago
This resource, along with Mathematics for Computer Graphics by John Vince [1], was truly indispensable when I wrote my own software renderer [2]. This was long before LLMs, so the whole process took me at least a couple months - most of it trying to wrap my head around math behind computer graphics and tracking down C segmentation faults. Fun times.

[1]: https://www.amazon.co.uk/Mathematics-Computer-Graphics-John-...

[2]: https://github.com/nkanaev/tipsy

bob1029 [3 hidden]5 mins ago
I wish we could have just one of these tutorials properly cover the concern of triangle clipping. This is the part that I struggle with the most in a software renderer. If you are going to be building a practical one, this is something you will eventually have to deal with, even for super basic scenes. Any time geometry intersects the view frustum you need to clip those triangles.
ggambetta [3 hidden]5 mins ago
aninteger [3 hidden]5 mins ago
Sorta unrelated but I really enjoy (I don't want to use the past tense as I often refer back to it) your article on Emulator Backed Remakes (https://gabrielgambetta.com/remakes.html). Thanks!
ggambetta [3 hidden]5 mins ago
Thanks for saying that :) Have you built something based on that idea? I still haven't, but I'm thinking on throwing some Claude at it for the most mechanical parts.
thechao [3 hidden]5 mins ago
You only need to clip triangles is you're worried about attribute interpolation for very large triangles. There's two ways to handle this: (1) discard (fast but not a great user experience); or, (2) primitive synthesis. Just frustum clipping is enabled by point picking in the local tile. Primitive synthesis requires some FP kung fu; but, is easiest done in barycentric space against a reverse transformed clipping rectangle. This lets you carefully control clip rounding error using either doubles or (better) fixed point. Abrash likes to use integer fixed point, but that is historical — modern fixed point can be handled with careful control of the fp unit in the mantissa. The major issue is regenerating the Z and the 1/Z values for the new vertices of the synthesized primitives. Everything else should flow down the pipe naturally, assuming a deferred attribute synthesis rasterizer.

There are examples in the open source version of my rasterizer: OpenSWR.org.

delta_p_delta_x [3 hidden]5 mins ago
I'm clearly super old-school when it comes to rasterisation, and a lot of this has flown above my head. I have a ton of questions; I hope you can answer them.

> (1) discard (fast but not a great user experience)

What are we discarding here, and why is it fast but not a great user experience?

> (2) primitive synthesis

I assume this is retriangulating clipped triangles that are now no longer triangles?

> reverse transformed clipping rectangle

Which spaces does this reverse transformation map from and to? I assume the clipping rectangle here is the triangle's AABB in raster space (or as you say, barycentric space).

> integer fixed point, but that is historical — modern fixed point can be handled with careful control of the fp unit in the mantissa

So we are no longer doing 16.16 fixed point, but tweaking the FP representation itself?

> The major issue is regenerating the Z and the 1/Z values for the new vertices

Why is this a major issue?

> deferred attribute synthesis rasterizer

I assume this means attributes are perspective-correct interpolated in raster space.

thechao [3 hidden]5 mins ago
Sorry I was so terse! Let's ignore guard band, for now. Our rasterization surface is a rectangle. We don't need to "geometrically clip" a triangle to the rectangle's surface. Instead, we just walk the surface of the rectangle and ask two questions: (1) does the triangle capture this (sub)sample; and, (2) what's the interpolated value of the attributes at this (sub)sample. In practice, for software rasterizers, we're working on tiny subrectangles (the tile), e.g., 4x4 or 2x8, whatever. So, we can be a little "inefficient" with our walking with respect to the edge testing for interiority. (This is roughly how HW works at the 2x2 level, as well.) Because the (sub)sample is "pulling" the attribute, we don't need to geometrically "clip" the triangle to the tile: the tile's (x,y) subsamples do that "for free".

On the flip side, if a triangle is 'really big' we need a guard band to either reject or subdivide triangles. The first one is fairly cheap -- we're throwing away the triangle! -- the second one is a better user experience, but requires synthesizing primitives. (The worst case is that a single triangle becomes 5 triangles, I think.) Each of those triangles needs its Z and 1/Z calculated in fixed precision. The precision of that fixed precision (though) can be clamped to the local tile; so, even though the global precision might need to be 25.25 (or whatever), the tile-local precision is only 4.9 (or whatever), with an intermediate 24.24 that can be handled with a float-float patch-up. The computation should all occur in the triangle's barycentric space: that means you need the inverted barycentric mapping to invert the guard band into the triangle's barycentric space. You do that because it lets you control the fixed point calculations better. (You can leave off the inverted determinant multiplication until the last moment.)

When I say "deferred attribute synthesis" I mean that we don't calculate attributes in the vertex shader. Instead, we calculate the barycentric, Z, and 1/Z values and pass those along. When we fire up the tile walker for the triangle (in general we only need 1-3 tiles), we calculate the attributes "on the fly, as they're used" and then let the compiler do CSE to fold down the replicated constructions.

ack_complete [3 hidden]5 mins ago
My own knowledge of GPU rasterization may be dated, but IIRC GPUs tend to rely on guard band clipping up to a guard band threshold before using geometric clipping. The guard band clipping involves rejecting 2D coarse rasterization blocks that are fully outside of the scissor rect. This is just a quick rectangle check, but the tradeoff is that a larger guard band means more GPU time lost in over-rasterization and potential higher precision requirements for rasterization values (which could be fixed point). Beyond the guard band, the triangles are clipped in floating point against the frustum clip planes.
Sharlin [3 hidden]5 mins ago
My renderer attempts always got stuck on the "should implement clipping" phase too, until I finally bit the bullet and managed to write a working one without much effort, independently "rediscovering" the Sutherland–Hodgman algorithm [1] as I found out later (googling it beforehand would've been cheating, of course).

The algorithm itself is fairly straightforward and intuitive, I think the biggest mental block is the weirdness of the projective space and working with homogeneous coordinates (actually the only frustum plane that you have to clip against in P₃(ℝ) is the front plane, the rest could be clipped after the perspective division, but no reason not to do it all at the same time while you're at it). The plane equations in the clip space are super simple, basically the six equations of the form ax + by + cz = w simplify to

  x = ±w 
  y = ±w 
  z = ±w.
Meaning, for example, that if the x coordinate of your vertex is greater than the w coordinate, that vertex is outside the right clipping plane. The Sutherland–Hodgman itself goes something like this:

  # Returns true if point is inside the half-space defined by plane
  def point_inside_plane(point, plane) -> bool:
    # single dot product, can be further simplified

  # Returns t such that the edge (p1, p2) intersects plane at lerp(t, p1, p2)
  def edge_intersect_plane(edge: (Point, Point), plane) -> float:
    # single dot product, can be further simplified

  # Given the vertices of a simple polygon and a plane,
  # returns the part of the polygon fully inside the plane 
  def clip_against_plane(poly: [Vertex], plane):
    let result: [Vertex] = []
    let [(v_1, v_2), (v_2, v_3), ..., (v_n, v_1)] = poly.edges()
    for each (v_i, v_j) of the edges:
      let i_inside = point_inside_plane(v_i, plane)
      let j_inside = point_inside_plane(n_j, plane)

      if i_inside and j_inside:
        # v_j will be pushed on the next iteration!
        result.push(v_i) 
      else if not i_inside and not j_inside:
        pass # Nothing to do!
      else:
        # One is inside, the other is not, we have to clip
        let t = edge_intersect_plane((v_i.pos, v_j.pos), plane)
        # Synthetize a new vertex straddling the plane
        let v_new = Vertex(
          pos = lerp(t, v_i.pos, v_j.pos),
          # For each vertex attribute
          attrib = lerp(t, v_i.attrib, v_j.attrib)
        )
        if i_inside:
          result.push(v_i); result.push(v_new) # discard v_j
        else:
          result.push(v_new); result.push(v_j) # discard v_i
      
   return result

Then you just call this for all the planes so that the output of one call becomes the input for the next call! The end result of this process is a convex polygon (of at most nine vertices for a triangle against six planes), which can be trivially triangulated. You can make the whole process faster by precomputing so-called outcodes which allow you to avoid clipping triangles known to be entirely outside at last one plane, or entirely inside every plane.

[1]: I. Sutherland and G. Hodgman. 1974. "Reentrant polygon clipping." Communications of the ACM, Volume 17, Issue. Available: https://dl.acm.org/doi/10.1145/360767.360802

blt [3 hidden]5 mins ago
This is a great comment. As someone who is familiar with R^3 geometry but not projection or clipping, I feel like I could almost implement it from your comment alone.

Can you explain more about clipping before/after the divide by w operation? Is this about avoiding coordinates with w = 0? Otherwise all the geometric ideas seem to make sense in both the world/view-space frustum and the ±1 box.

espetro [3 hidden]5 mins ago
Finally an engineering feat that's not built in Rust
0x1ceb00da [3 hidden]5 mins ago
Ok I'm following along but how do I view the tga files that this program produces? Windows can't open them.
mhoela [3 hidden]5 mins ago
Irfanview?
t1234s [3 hidden]5 mins ago
I jumped when the first image loaded on the page.. didn't expect that.
hnlmorg [3 hidden]5 mins ago
What image was that? The first image I get is a headshot.

Personally I’d have gone with a teapot as my demo but I don’t see anything shocking about the model the author chose.

mito88 [3 hidden]5 mins ago
> I jumped when the first image loaded on the page.. didn't expect that.

mbappe'?

anthk [3 hidden]5 mins ago
There was a software rendered game in vein of Tomb Raider 1 graphics for both DOS and Unix, but I can't remember its name. It was an exploration game, modern, a bit cyberpunkish, in 3D.
brcmthrowaway [3 hidden]5 mins ago
How does this differ from POVRay?
42days [3 hidden]5 mins ago
POVray is a declarative language used to render raytraced scenes. This is a tutorial that demonstrates how the rasterization technique works from a programmatic perspective.
fithisux [3 hidden]5 mins ago
This is a very good article.
uncivilized [3 hidden]5 mins ago
I thought this would be something new but it’s just ssloy’s tinyrenderer. Article should have a date since it’s old as dirt
mootdentures [3 hidden]5 mins ago
Just because you've seen something before doesn't mean it isn't new for some readers! This is my first time seeing it
chrisjj [3 hidden]5 mins ago
"Comments

An error occurred: API rate limit already exceeded "

Sharlin [3 hidden]5 mins ago
I've been wanting to write a series of articles exactly like this for a long time. Good that someone else has done it now =D