The lecture notes you keep coming back to even after you think you know the material
I first encountered Numerical Linear Algebra By Trefethen And Bau when I was roughly two years into a graduate program that assumed everyone already knew how to make a matrix behave. The book is forty-eight lectures. It covers most of what you need and deliberately skips most of what you don't. That design choice has caused more arguments in my office than any single algorithm in the text. Lecture 4 is where the book earns its reputation. Condition numbers, sensitivity, the whole idea that a well-posed problem can still be numerically unkind. The derivation of the perturbation bound for linear systems is tight enough to be useful and short enough to actually read. I've assigned those pages to students who later came back and said they finally understood why their code produced garbage even though the math was correct.
Why Numerical Linear Algebra By Trefethen And Bau still matters in practice
The notes are not a reference manual. You won't find exhaustive tables of LAPACK call signatures or detailed discussions of block algorithms for modern GPUs. What you get is a clean account of the fundamental methods: Gaussian elimination with partial pivoting, LU factorization, orthogonal transformations, QR decomposition, the SVD, least squares, iterative methods for symmetric systems, and eigenvalue problems via the QR algorithm. Each topic gets roughly equal attention relative to its conceptual importance, not relative to its industrial usage volume. The pedagogical strength is the consistency. Trefethen writes every lecture in the same voice. When you move from Lecture 9 on QR factorization to Lecture 19 on the singular value decomposition, the definitions and notation do not shift. That matters more than it sounds. I have seen students bounce between five different textbooks before realizing they were reading three different conventions for the same operation.
How to actually use these notes instead of just reading them
Read actively. The exercises are not filler. The first thirty problems in each lecture directly test whether you can reproduce the main theorem without looking at the proof. I spend about forty-five minutes on Lecture 6 because the Gram-Schmidt orthogonalization exercises expose the difference between the classical and modified forms. If you skip those, you will later wonder why your projection code is slowly drifting toward linear dependence. Run the accompanying MATLAB examples. The book ships with a small set of code snippets and a companion web page with more. There is a reason people cite it alongside computation. The notes explain why partial pivoting works almost always, and then Lecture 7 shows you where it fails, with a concrete counterexample that is only about a thirty-by-thirty matrix but takes an hour to write down on paper. Do not treat the iterative methods section as exhaustive. Lectures 35 through 39 cover CG, GMRES, and basic preconditioning ideas. They are correct and well-motivated. They are also a starting point. If you are solving large sparse systems in production, you will need a library like PETSc or Trilinos afterward, but understanding what the notes describe first will save you months of debugging someone else's wrapper code.
Get the Full Details

The part nobody warns you about until it costs you a weekend
Here is a specific failure mode I ran into last fall. I was debugging a solver for a discretized boundary value problem where the coefficient matrix was symmetric positive definite but had condition number around ten to the eight. I used a straightforward CG implementation following the algorithm in Lecture 36. The residual dropped for roughly twelve iterations and then started climbing. The theoretical bound said convergence should occur in about eighty steps, but the arithmetic was clearly destabilizing. The issue was not the theory. The issue was that I was forming the matrix explicitly from a finite-difference stencil and then passing it to CG. Round-off in the assembly step introduced a tiny symmetric component that was not quite positive definite, something on the order of machine epsilon times the norm. CG is sensitive to that. The workaround was to switch to a matrix-free evaluation of the action A*x and use a shift-and-invert preprocessing step with a tolerance tight enough that the near-singularity did not affect the early iterations. The fix took about twenty lines of code and saved me from rewriting the entire solver. I mention this because the book assumes you are working in exact arithmetic long enough to absorb the theory, and that is appropriate for a course text. It does not teach you how to handle the edge cases that appear when floating-point arithmetic meets a near-singular problem.
Common pitfalls that beginners keep repeating
The first one is confusing backward stability with forward accuracy. Partial pivoting is backward stable for general nonsingular matrices. That does not mean the computed solution is close to the exact solution of the original system. It means the computed solution is the exact solution of a nearby system. The distance to that nearby system depends on the condition number, which brings us to the second pitfall. The third pitfall is treating QR as a universal replacement for Gaussian elimination. It is more stable, yes. It is also roughly twice as expensive in floating-point operations for a dense square system. For least squares problems, the QR approach through the normal equations is a bad idea unless you have a strong reason. Lecture 12 covers this directly. People ignore it anyway.
Where the notes are thin and what to read instead
The eigenvalue algorithms section stops at the QR iteration and does not deeply cover shifts, deflation strategies, or the implicitly shifted QR algorithm that every production code uses. If you need implementation detail there, Saad's Iterative Methods for Sparse Linear Systems or the LAPACK Working Notes are more useful, though neither is as readable. The numerical ODE side is absent. The notes are linear algebra. That is a feature, not a flaw, but it means you should pair this text with something that covers discretization error if your work spans both areas. The treatment of fast transforms is also light. Fast Fourier and Walsh transforms appear only in passing. If your interest is computational efficiency for structured matrices, you will need additional material, such as the work by Gustavson and others on banded and cyclic reduction methods.

A practical study sequence that works
Work through Lectures 1 through 14 first. This gives you factorizations, stability, and least squares. Then move to 15 through 24 for the SVD and its applications. Lectures 25 through 34 cover eigenvalue problems and orthogonal reduction to Hessenberg or tridiagonal form. The final fourteen lectures address iterative methods and some advanced topics like Krylov subspace techniques. Spend about six to eight hours per lecture if you are doing the exercises. The total commitment is roughly three hundred hours for a serious pass. That is reasonable for a text that condenses a full semester of graduate content.
The notes are free and easy to access legally
The official source is Lloyd Trefethen's personal web page at Oxford. The PDF is available at his university-affiliated site and has been archived by the Mathematical Sciences Research Institute. I recommend downloading the version linked from the Oxford department because the typesetting is cleaner than the mirrored copies that circulate on random document-hosting sites. There is also a companion video lecture series posted online that matches the notes almost lecture-by-lecture. Watching the videos while reading is useful the first time. The second time, skip the videos and just do the problems.
What this book will not give you and why that is fine
It will not teach you how to parallelize BLAS routines. It will not discuss GPU implementations. It does not cover randomised linear algebra, tensor methods, or the newer preconditioners that have appeared in the last decade. That is expected. The book targets foundational understanding, and it delivers that consistently. If you are looking for a coding manual, buy a different book. If you want to understand why the algorithms you use behave the way they do, these notes remain one of the clearest accounts available. I return to them every few years when I need to remember the precise statement of a theorem or reconstruct an argument from first principles. That repetition is the real measure of a textbook.