When I first started working on games, I used high level frameworks that would give you a Camera that just kind of worked. But as I learned more and wanted t...
39 comments
What this lets you do is composing multiple transforms into one matrix multiplication instead of a sequence of multiplications and additions; that's what dramatically increases performance, on modern computing devices but most especially on ancient ones, where we were fighting for every MUL.
More details: https://gabrielgambetta.com/computer-graphics-from-scratch/1...
But you do need homogenous coordinates if you want to include translations (fixed distance shifts, such as moving the camera) in the set of operations you can represent as linear transformations and thereby gain all the benefits of linear algebra, including the benefit of being able to include them in a series of operations that you can represent with a single transformation matrix.
Consider a transformation f where we wish to move x-y coordinates s units to the right. In 2d, we could express it as:
f(x, y) = (x + s, y)
But that transformation is affine not linear. There is no way to generate the value s as a linear combination of the inputs x and y. So, our workaround is to embed the x-y plane into 3d space at z=1. Then we can move (x,y,1) points in that plane s units to the right using this transformation:
f(x, y, z) = (x + s*z, y, z)
This new transformation is linear: it maps (0,0,0) to itself. But it maps our embedded 2d plane's origin (0,0,1) to (s,0,1), shifting it right by s units, as we want.
The matrix form of that transformation is:
[[1 0 s]
[0 1 0]
[0 0 1]]
The same scheme would work if we had embedded the plane at any fixed z=r for nonzero r. We would only have to rescale the s in the matrix to s/r. Again, however, if r=0, this scheme will not work, as 1/r has gone to infinity.[0]: Here’s a visual: https://gunn-gatm.github.io/textbook/gatm.pdf#page=28
so as the other 2 helpful commenters also just said: 3d shears using linear algebra degenerate to 2d affine transformations when z=1 (or w in 4d)
The killer feature is that you can put the two together for a rotation axis (~ vector) and angle (~ scalar).
With just three gimbals (rotating circles), if gimbals A and B are aligned, you only have two degrees of freedom (rotating A is the same as rotating B). Because of this, interpolating angles is unwieldy in vector space. 'Gimbal lock' confounds animation (in hilarious but unrealistic ways) but also aerospace (four hours before 'one small step for man', just after landing, Collins joked he would like a fourth gimbal for Christmas).
I've used that when teaching short "Graphics 101" (not in the first session though) and the math comes out more intuitive than the usual "here's how to calculate a perspective matrix, don't ask where these numbers come from" version.
Students can easily imagine the pyramid from an eye to the window frame, and that it keeps going.
Then I point to an object outside the window and say imagine strings going from the corners of the object to your eye. To do this they would have to go through the window, where do they do that. 3d graphics is finding out where on the window the strings go through so you can stick a picture of thing outside onto the window and it looks exactly the same.
A lot of 3D graphics can be derived pretty easily just from knowing a few basics like divide by depth. I think knowing how to construct a transformation matrix from a coordinate system basis is another one -- that would remove the need to look up how to construct a perspective matrix, for example.
A few things like that will get you pretty far and you can kind of take the same journey of discovery as early 3D pioneers. That's one of the best ways to learn because you're much more likely to remember something you figured out compared to something you just read about.
It gets tricky with perspective-correct textures, but running into issues like that on your own (even if not solved on your own) is part of the fun of learning, I think.
I still remember the rush that I got when I got the rotating cube in my QBasic IDE in my 14inch monochrome monitor. The intuition I used was derived from how light casted shadows of 3d objects in a 2d wall..From there, this x' = x/z y' = y/z, follows naturally..
Maybe consider clipping the ball properly on the edges of the sides of the view frustum ?
Similarly the near and far could also be clipped; i know this is not true of the math necessarily but is the expected result in 3d graphics applications.
Now-a-days maybe you can ask an LLM for something exactly at your level and anything you don’t get you can ask it that too. That wasn’t the case until recently and many still don’t use them
I suck at math myself so can relate. But actually the math you need to learn to be effective at graphics is mostly about understanding _a very small subset_ of a bunch of concepts (such as how 2x2, 3x3 and 4x4 matrices work) and one is set for life,
so I think it’s good to give the formal mathemathical explanation and then Gilbert Strang what it all means.
But camera transform is so unintuitive it really helps if the student tries to work it out for themselves why it works.
For me it’s really weird anyone would write code and _not_ start from reading all of introductory computer graphics literature because it’s so awesome but I realize I’m the outlier here.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 1 day ago
- Hacker News · 2 points · 9 days ago
- Hacker News · 1 points · 3 days ago
- Hacker News · 1 points · 2 days ago
- DEV Community · 17 points · 7 days ago
- The Verge · 0 points · 6 days ago