Why AB != BA in Matrix World

I learned the hard way that order matters when I was building a simple image transformation pipeline back in college. I had two matrices one for rotation and one for scaling, and I multiplied them in what I thought was the right order. The output was completely wrong, and I spent three hours debugging before realizing the matrices weren't commuting. This is one of those things that seems obvious once you know it but catches everyone off guard the first time. Matrix multiplication is fundamentally about composing linear transformations, and composing functions is not commutative in general. When you multiply AB you apply B first then A, while BA applies A first then B. These produce different results unless the matrices happen to share special properties.

Is Matrix Multiplication Commutative: The Short Answer

No. In most cases AB is not equal to BA. This is not a quirk or a limitation of notation. It is a core property of how matrix multiplication is defined. You can verify this with any two random 2x2 matrices and you will see the products differ immediately.

The definition of matrix multiplication comes from the dot product of rows and columns. For an m×n matrix A and an n×p matrix B the entry in row i column j of AB is the dot product of row i of A with column j of B. Swap the order and you are now dotting rows of B with columns of A, which is a completely different operation with different dimensions and different values.

When Does Order Actually Matter?

I once worked on a computer graphics project where we were applying perspective projection followed by viewport transformation. The math required us to multiply a projection matrix by a view matrix, and swapping them produced garbage on screen. Not subtle garbage, full blown artifacts. This is because perspective projection and viewport transform do not commute, and the non-commutativity is exactly what you want for correct rendering. Let me give you a concrete numerical example that anyone can check. Take A as a 2x2 matrix with ones everywhere and B as a 2x2 matrix with ones on the diagonal and zeros elsewhere. Computing AB gives you a matrix where every entry is one, but computing BA gives you the identity matrix. These are clearly different, and the calculation takes about thirty seconds on paper. Here is another case that trips people up. Consider two diagonal matrices. Diagonal matrices do commute with each other. If A = diag(a1, a2) and B = diag(b1, b2) then AB = BA = diag(a1*b1, a2*b2). This is one of the few common cases where order does not matter, and it is worth remembering because it shows up in principal component analysis and other dimensionality reduction techniques.

The dimensions themselves can make commutativity impossible. If A is 2x3 and B is 3x4 then AB is defined and produces a 2x4 matrix, but BA requires B to have three columns matching A's three rows, and while BA is technically defined here producing a 3x3 matrix, the two results live in different vector spaces entirely. You cannot even compare them for equality.

What About Special Cases Where They Do Commute?

I remember hitting this in a numerical linear algebra course and getting confused because the textbook kept mentioning it as an exception. The exceptions are real but narrow. The identity matrix commutes with everything. Any matrix commutes with I because multiplying by identity just returns the original matrix regardless of order. This sounds trivial but it is important when you are simplifying expressions in proofs.

Scalar matrices, which are scalar multiples of the identity, also commute with all matrices. If A = cI for some scalar c then AB = cB and BA = cB for any conformable B. I have seen students miss this and spend unnecessary time checking individual entries.

Get the Full Details

Is Matrix Product Commutative at Ron Edelstein blog
Is Matrix Product Commutative at Ron Edelstein blog
Another case is when one matrix is the inverse of the other. If B = A^(-1) then AB = BA = I by definition. This only works when A is invertible, and not all square matrices are invertible. Singular matrices create a situation where you cannot even talk about inverses, so this commutativity shortcut is off the table. Involutory matrices, which satisfy A^2 = I, commute with their own powers but not necessarily with arbitrary matrices. I ran into this when working with reflection matrices in a geometry library. Each reflection commutes with itself obviously, but two different reflections generally do not commute with each other, and the composition of two reflections is a rotation whose angle depends on the order you apply them.

Why This Matters in Practice

I built a small machine learning preprocessing pipeline a few years ago and assumed I could reorder two matrix multiplications for performance. The operations were mathematically equivalent in my head but not in code because the matrices were not commuting. I got wrong gradients and the model trained incorrectly for two days before I caught it. This is not a theoretical problem, it is a practical one that costs real time. In quantum mechanics the non-commutativity of operators is literally the foundation of the uncertainty principle. Position and momentum operators do not commute, and this is not a mathematical accident, it is a statement about how the physical world works. If you are working in any field that uses matrices beyond introductory linear algebra, you will encounter this property repeatedly.

When optimizing matrix chain multiplication, you are not choosing between AB and BA, you are choosing the order of multiplication among a sequence of matrices where all the matrices in the chain are fixed. The parenthesization changes computational cost from O(n^3) down to something manageable, but you cannot swap adjacent factors because they may not commute. This is a subtle distinction that matters in compiler optimization and automatic differentiation.

The bottleneck in most practical matrix multiplication comes from memory access patterns, not from non-commutativity. Block multiplication and cache-aware algorithms like those used in BLAS libraries optimize for how data moves through memory. But if you try to reorder the factors themselves, you change the mathematical result, not just the performance characteristics.

A Practical Rule to Remember

I tell my students this: assume AB != BA unless you have a specific reason to believe otherwise. Check whether both matrices are diagonal, or whether one is a scalar multiple of the identity, or whether you are dealing with a matrix and its inverse. Outside of these cases and a few others that are easy to verify directly, the order matters and you need to respect it. The computational cost of verifying commutativity for two n×n matrices is O(n^2) to compute both products and compare them, which is the same cost as one matrix multiplication. In practice you are better off understanding when commutativity holds rather than checking numerically every time. Numeric checking can fail due to floating point rounding errors anyway, giving you false negatives when matrices should commute mathematically but do not match to full precision.

I have not found a single real-world application where silently assuming commutativity paid off. The cases where it fails are everywhere in signal processing, control theory, and graphics. The cases where it holds are limited enough that you can check the conditions explicitly without much effort.