Why this textbook keeps showing up on every stats grad school syllabus

Statistical Inference by Casella and Berger is the standard reference for mathematical statistics at the graduate level. It covers measure-theoretic probability, point estimation, hypothesis testing, and confidence intervals with a level of rigor that most introductory texts skip entirely. If you are preparing for qualifying exams or need a serious treatment of sufficiency and completeness, this is the book. The official publisher is Chapman and Hall/CRC. The second edition came out in 2002 and remains in print. Most university libraries carry the physical copy. Online, you will find digitized versions circulating on academic file-sharing platforms, repository sites, and PDF hosting services. The book is approximately 669 pages in the second edition. It is dense. Do not expect to skim it. The first chapter covers sigma-algebras and probability spaces. This is not optional background material. The entire rest of the book assumes you are comfortable with measure-theoretic definitions. If you come from an applied statistics background where expectation is just an integral sign you move past quickly, chapter one will slow you down considerably. I spent roughly three weeks working through the first four sections before moving on because I kept hitting gaps in my understanding of convergence modes.

Chapter six on sufficient statistics and the Rao-Blackwell theorem is where the book earns its reputation. The Lehmann-Scheffe theorem follows naturally and gives you a clean path to finding uniformly minimum variance unbiased estimators. The proofs are self-contained but demanding. You cannot gloss over them and still expect to use the results correctly in practice. One specific problem I ran into involved applying the Neyman-Pearson lemma to a composite hypothesis where the parameter space had a boundary issue. The textbook example assumes an open parameter space and regularity conditions that do not hold in my case. I worked around it by constructing a reduced parameter space using a monotone likelihood ratio argument and then verifying the boundary condition separately. This added about two hours to a problem set that would normally take thirty minutes, but it was the only way to make the conclusion valid.

Common mistakes people make when working through this book

The most frequent error is treating the exercises as optional. They are not. The exercises extend the theory in ways the main text does not always cover explicitly. Skipping them leaves real gaps in understanding, particularly around exponential families and completeness arguments. Another mistake is reading the book linearly from start to finish without returning to earlier chapters. Chapter seven on estimation builds directly on chapter five's treatment of risk functions and admissibility. If you do not retain the decision-theoretic framework from chapter five, chapter seven becomes nearly impenetrable on a first pass. There is also a tendency to over-rely on the solutions manual without working through proofs independently first. The manual exists and is widely available, but using it prematurely turns active learning into passive verification. You will recognize the steps but not know when to apply them in a new context.

Get the Full Details

Casella Berger Statistical Inference | PDF
Casella Berger Statistical Inference | PDF

When this book is the wrong choice

If you need practical data analysis guidance, this is not the right resource. Casella and Berger do not cover computational methods, bootstrap procedures in depth, or modern machine learning connections. For applied work, pairing this text with something like Gelman's Bayesian Data Analysis or Hastie's Elements of Statistical Learning fills the gap. The book also assumes mathematical maturity that some graduate students do not yet possess. If real analysis or advanced calculus is unfamiliar, the notation and proof style will feel impenetrable regardless of how hard you work. In those cases, starting with Hogg and McKean's Introduction to Statistical Thought as a bridge before tackling Casella and Berger reduces the friction significantly.

How I use this book in practice

I keep it on my desk as a reference rather than reading it cover to cover. When I encounter a situation involving unbiased estimation under constraints or need to verify whether a statistic is complete, I look up the relevant theorem and re-derive the proof. This takes longer than consulting a shorter reference but reinforces the conditions under which each result holds. That matters when the problem you are solving deviates even slightly from the textbook assumptions. For exam preparation, I work through selected exercises from chapters two, four, six, and eight. These chapters contain the material most likely to appear on comprehensive exams. The problems on UMVUE and likelihood ratio tests repeat in variations across institutions. The second edition contains errata that are documented online. Checking the publisher's errata page before starting a difficult chapter saves time. A few incorrect equations in the first printing can derail an entire proof attempt if you follow them verbatim.