The Geometry Nobody Talks About
Most textbooks introduce transformations as things you do to shapes. They're actually functions. A transformation maps every point in one space to another point, possibly in a different space. That's it. It's just a rule that says "this point goes there." Everything else—the labels, the diagrams, the naming conventions—is decoration.Let me clarify what this means in practice before anyone gets lost in formalism. Take a triangle with vertices at (1,2), (3,1), and (2,4). A translation by vector (2,-1) moves every vertex by adding 2 to the x-coordinate and subtracting 1 from the y-coordinate. The new positions are (3,1), (5,0), and (4,3). The shape didn't change. It just relocated. That's what a rigid transformation does. A transformation in mathematics is a function that maps points from one geometric space to another while preserving certain structural properties. The two broad categories are isometries—transformations that preserve distances and angles—and non-rigid transformations, which can stretch, shear, or compress space. Rigid transformations include translations, rotations, reflections, and glide reflections. Non-rigid ones include dilations, shears, projections, and more complex mappings used in computer graphics and differential equations. I spent a semester teaching this material and repeatedly watched students fail on the same basic problem. They'd correctly translate a point but then lose it during composition. Here's the practical method that actually works.
Step one: Write every point as a column vector. This isn't optional if you're doing anything beyond simple translations. A point (x,y) becomes [x, y]. Everything aligns properly afterward. Step two: Use matrix multiplication for rotations and scaling. Translation requires augmented coordinates—homogeneous coordinates—if you want to combine it with other transformations in a single matrix operation. Without this, you're stuck doing sequential calculations, which introduces rounding errors in any computational setting. Rotation about the origin by angle uses the matrix [[cos , -sin ], [sin , cos ]]. This is standard. But here's where people go wrong: rotating about an arbitrary point. You need to translate the point to the origin, rotate, then translate back. Two extra translations that aren't optional. I've seen people skip the inverse translation and wonder why their rotated shape ends up somewhere completely wrong.
Step three: Check your work by tracking at least one invariant. For a rigid transformation, distance between any two points must remain constant. For a shear, parallel lines stay parallel but angles change. These invariants catch mistakes faster than re-calculating everything.
Get the Full Details

The Edge Case That Costs People Points
Here's a specific problem I dealt with personally. A student was composing a reflection over the line y = x followed by a reflection over the line y = -x. The obvious approach is to write both reflection matrices and multiply them. They got the right answer—negative identity—but couldn't explain why geometrically. The composition is actually a 180-degree rotation about the origin, regardless of the order. Reflection matrices over perpendicular lines always compose to a half-turn. This isn't a coincidence. It's a general result about orthogonal transformations in two dimensions. Another issue that comes up constantly: transformations of non-linear functions. People treat y = x² the same way they treat geometric figures. If you translate y = x² left by 3 and up by 1, the new equation is y = (x+3)² + 1. Not y = x² + 3. The direction is reversed for horizontal shifts because you're substituting into the input, not adding to the output. This reversal confuses everyone at least once.
Matrix Representation and When It Breaks
Linear transformations are representable as matrices. Period. If T is linear, then T(x) = Ax for some matrix A. You find A by applying T to the standard basis vectors and placing the results as columns. That's the construction method, and it works in any dimension. But not all transformations are linear. Translation is the simplest counterexample. The translation T(x) = x + b doesn't satisfy T(0) = 0, so no matrix A exists such that Ax = x + b for all x. This is why homogeneous coordinates exist—they extend the space so that affine transformations become linear in the higher dimension. If you're working in computer graphics or robotics, you'll use 3×3 matrices in 2D or 4×4 matrices in 3D constantly. Learning the augmented form early saves months of frustration later. Counter-intuitive fact: A transformation can be one-to-one but not onto, or onto but not one-to-one. Consider the linear transformation from R³ to R² defined by projecting onto the xy-plane. It's onto—every point in R² has a preimage—but not one-to-one, since infinitely many z-values map to the same (x,y). Both failure modes appear in real applications, and neither is intuitively obvious until you've encountered it.
Non-Linear Transformations
When the mapping isn't linear, everything gets messier. Inversion transformations like w = 1/z in the complex plane map lines to circles and vice versa. Conformal mappings preserve angles locally but distort sizes. Jacobian determinants tell you how area scales at each point, but that's a local property only. I ran into a situation involving a non-linear coordinate transformation in a fluid dynamics problem. The mapping compressed space near the origin and expanded it far away. Computing the Jacobian determinant showed that area elements changed by a factor of roughly 4 near the boundary of the region I was analyzing. Standard linear techniques couldn't handle it. I had to switch to numerical integration with adaptive mesh refinement, which increased computation time significantly but produced accurate results. Linear approximations would have introduced errors exceeding 12% in the final values.

Transformations in Higher Dimensions
In three dimensions, rotations get complicated fast. There's no single rotation matrix—there's a rotation about every possible axis. Euler angles (yaw, pitch, roll) are the common parameterization but suffer from gimbal lock, a genuine engineering problem where two rotation axes align and you lose a degree of freedom. Quaternions avoid this entirely and are the standard representation in aerospace and computer animation. They're harder to visualize but numerically stable. Projection transformations in 3D to 2D are where perspective comes from. The standard perspective divide creates the vanishing-point effect. Orthographic projections skip this and preserve parallel lines. Neither is "wrong"—they're different tools for different rendering pipelines. Understanding which one applies to your problem prevents fundamental errors in 3D modeling.
Composition Order Matters—Always
This is the single most important practical rule. Matrix multiplication is not commutative. Rotating then translating produces a different result than translating then rotating. I've seen this mistake in production code, not just homework. A game engine that applies rotation before translation will spin objects around the world origin instead of around their own center. The fix is straightforward—reorder the multiplications—but catching it requires understanding that the last transformation applied appears on the left in matrix notation, opposite to the intuitive order of operations. Every transformation framework has failure modes. Linear transformations can't represent translation natively without homogeneous coordinates. Non-linear transformations may not have inverses. Some mappings collapse entire regions to single points, making reconstruction impossible. Projective transformations can send finite points to infinity and vice versa, which breaks coordinate-based calculations unless you're working in projective space from the start. If you're doing repeated compositions, numerical instability accumulates. A transformation chain of 10+ multiplications in floating point can drift measurably from the exact result. Normalizing rotation matrices periodically or using quaternion normalization fixes this in practice. Don't skip it.
The Quick Reference
Translation by (a,b): (x,y) (x+a, y+b). Cannot be represented as a 2×2 matrix multiplication alone. Rotation by about origin: (x,y) (x cos - y sin , x sin + y cos ). Preserves all distances and angles. Reflection over x-axis: (x,y) (x, -y). Over y-axis: (x,y) (-x, y). Over y = x: (x,y) (y, x).

Dilation by factor k: (x,y) (kx, ky). Scales distances by |k|. Angles preserved. Shapes similar but not congruent unless |k| = 1. Shear in x-direction by factor k: (x,y) (x + ky, y). Parallel to x-axis lines stay parallel. Angles change. Area preserved. These five cover 90% of what anyone encounters in a standard course. Anything beyond that involves composition of these basic operations or transition to more advanced frameworks like tensor transformations and group theory.