A Book That Actually Teaches You to Think About Data
Most introductory statistics textbooks build the house from the roof down. They pile on probability theory, conditional distributions, and limit theorems until you can derive the central limit theorem in your sleep but still don't know what to do when a dataset has a long right tail and seventeen missing values. John Rice Mathematical Statistics And Data Analysis flips that entire scaffold. It starts with real data, shows you what the data is doing, and then introduces the mathematical machinery only when you actually need it to explain what you are seeing. I first worked through this book around 2018 while designing an undergraduate lab sequence for a junior-level stats course at a mid-tier research university. My problem was that students would breeze through two semesters of probability theory and then freeze when handed a raw CSV file from a real experiment. They could compute a variance by hand but could not decide whether a transformation was necessary or how to justify it in writing. Rice's text forced the sequence backward. We started each module with a dataset, built intuition visually, and only then opened the math chapters.
The Data-First Approach in Practice
The third edition, published by Cengage and now in its later printings, runs roughly 650 to 700 pages depending on the format. The book is divided into four parts. Part I covers data description and exploratory analysis using scatterplots, stem-and-leaf displays, and basic summaries. Part II moves into probability fundamentals, but the probability material is anchored to examples rather than abstract measure theory. Part III handles estimation and hypothesis testing with a stronger emphasis on bootstrap methods and permutation tests than most competing texts allow at this level. Part IV covers regression, analysis of variance, and nonparametric methods. What makes the structure functional rather than merely different is how Rice treats probability. Instead of defining random variables, deriving moment-generating functions, and only later connecting them to anything empirical, he introduces distributions through shape, scale, and real-world behavior. The exponential distribution arrives after discussing waiting times and decay processes, not before. The chi-squared distribution appears naturally when analyzing categorical data and goodness-of-fit, which is how it actually shows up in practice. This ordering matters because students stop asking why they are learning these things and start seeing them as tools for specific questions.
Specific Features That Matter More Than You Would Expect
Rice includes substantial material on computational methods early, which is unusual for a book at this level. The bootstrap section in Part III does not just give you the algorithm and a couple of textbook examples. It walks through why bootstrapping works, when it breaks down, and how to interpret the output correctly. I had a student who used the percentile bootstrap on a dataset with heavy skew and got confidence intervals that extended below zero for a variable representing time-to-failure. The book does not explicitly flag this edge case, so we spent an extra lab session discussing BCa corrections and when the basic bootstrap is unreliable. That gap in the text turned out to be a useful teaching moment. The ANOVA chapters are also stronger than typical for this tier. Rice derives the F-test from first principles using likelihood arguments rather than just stating the test statistic and moving on. He connects the linear model framework to both regression and experimental design, which closes a conceptual loop that many students find mysterious until they take a separate course on design of experiments. The treatment of blocked designs and randomized complete blocks is thorough enough that a graduate student could use it as a reference.
Get the Full Details

A Concrete Problem I Ran Into and How I Worked Around It
One of my recurring issues with this book involves the treatment of multiple comparisons in the regression and ANOVA sections. Rice introduces Bonferroni and Tukey adjustments, which is correct but sometimes insufficient for real datasets. In a project involving a dataset with roughly forty predictors and correlated covariates, I had students apply stepwise selection procedures that the book mentions only briefly. The results were unstable and the p-values were meaningless after selection. The textbook does not cover regularization methods like lasso or elastic net, which would have been the appropriate tool for that kind of data. I supplemented the course with an outside module on penalized regression, drawing from Hastie, Tibshirani, and Friedman. Students who only had Rice in that situation would have been underprepared for any modern applied work involving high-dimensional data. Another friction point is the book's coverage of Bayesian methods. It exists, but it is limited to a small section near the end. If you are teaching a course that needs even a surface-level Bayesian module, you will need to add external materials. The frequentist treatment is solid, but the asymmetry between the two paradigms is noticeable.
What the Book Gets Wrong or Underweights
The book assumes a calculus background and a comfortable relationship with sigma notation. Students who have not taken multivariable calculus or who struggle with summation manipulation will find the derivations dense. This is not a flaw in the writing but a structural limitation. The book is positioned somewhere between a heavily theoretical text like Casella and Berger and a more applied data-science book like James et al., and it does not fully satisfy either camp. The computational exercises are generally done by hand or with basic software like R or Minitab. Modern courses that expect students to work in Python, use pandas, or touch scikit-learn will need to adapt the assignments. The data files are available online, but the book itself does not include code snippets. That means an instructor has to write or source the code, which is time-consuming if you are building a course from scratch. John Rice Mathematical Statistics And Data Analysis remains one of the better intermediate texts available, but it is not a complete package for a contemporary curriculum. It excels at building mathematical maturity and statistical reasoning. It is weaker on modern computational workflows, high-dimensional methods, and Bayesian content. If you pair it with a short module on regularization and a Python-based lab component, it works well for a one-semester upper-division course. If you expect it to cover the full range of applied statistics, you will be disappointed.
How to Use It Effectively
Start with Chapters 1 and 2 to establish the language of data. Do not skip the exercises on constructing and interpreting plots. Students tend to underestimate how much they learn from reading graphs carefully before touching a formula. Move into probability through Chapters 3 and 4, but spend extra time on Chapter 4 if your students have weak calculus backgrounds. The integration techniques used to derive expectations and variances are straightforward, but the mental shift from algebra to continuous mathematics is real and some learners need more practice here. The estimation and testing chapters, roughly Chapters 7 through 10, are the core. Work through them slowly. The bootstrap material deserves two or three lab sessions rather than a single lecture. The hypothesis testing chapters contain the most common pitfalls in student reasoning, particularly around p-value interpretation and the difference between statistical and practical significance. I recommend pairing each major concept with a dataset where the difference is dramatic, such as a large sample with a tiny effect size that is statistically significant but substantively irrelevant. For regression, use Chapters 11 and 12 as the foundation, then supplement with external material on diagnostics, multicollinearity, and model selection criteria like AIC and BIC if your course timeline allows. The book covers AIC briefly but does not integrate model selection into the main narrative, which is a missed opportunity for a text at this level.

Where It Falls Short and What to Use Instead
If your goal is rigorous mathematical statistics with full measure-theoretic foundations, Casella and Berger is the standard alternative, though it is significantly more demanding. If your goal is applied data analysis with modern computational tools, Introduction to Statistical Learning by James, Witten, Hastie, and Tibshirani is a better fit, especially for students who want to move into machine learning territory. Rice occupies a middle ground that is valuable but not decisive. It is best used as a backbone for a course that wants to emphasize understanding over procedure, with supplemental materials filling in the gaps for computation and contemporary methods. The third edition is available through Cengage and various academic booksellers. The solution manual exists for instructors. The data and code can be downloaded from the publisher's companion site. There is no official open-access version, so budget constraints may affect adoption in programs that cannot justify the textbook cost. In those cases, libraries and course reserves usually carry copies, and students can share them if the syllabus is shared across multiple sections.