Working Through Stock and Watson Empirical Exercises
The empirical exercises in Introduction to Econometrics by Stock and Watson are where most students hit a wall. The textbook walks you through theory cleanly, then drops you into messy datasets with instructions that assume you already know which regression commands to run first. I spent years helping grad students untangle these problems, and the pattern is always the same: they copy answers without understanding why the output looks wrong, then panic when their professor asks follow-up questions. Here is what actually happens when you sit down with these exercises. You load the dataset into R or Stata, run the regression your professor assigned, and get coefficients that don't match the expected values. Eight times out of ten, it is a variable transformation issue. Stock and Watson frequently ask you to work with logged variables, differenced data, or interaction terms, but the datasets come raw. The FRED data for instance includes nominal values that need inflation adjustment before any meaningful regression. I remember one student who spent six hours trying to match Chapter 4 exercise results. She was getting strange heteroscedasticity diagnostics. Turned out she had not dropped the missing observations before running the regression. The dataset had a few NA values scattered through the time series, and different software handles those differently. R includes them by default in some functions while dropping them silently in others. Once she ran na.omit() explicitly, her numbers matched perfectly.
The biggest mistake people make is treating these solutions as something to copy rather than something to dissect. Each empirical exercise in this textbook builds on the previous ones. Chapter 3 introduces simple regression. Chapter 4 moves to multiple regression with controls. Chapter 5 brings in binary regressors. If you skip ahead and try to solve Chapter 7 problems without understanding how the OLS assumptions shift between chapters, you will not recognize when your standard errors are wrong. Another thing nobody tells you about these exercises: the datasets have evolved across editions. The third edition uses different data files than the second. The macrodata.csv file from Chapter 4 is not identical between versions. If you are using older solution sets with newer textbook problems, your intercepts will be slightly off. Always check the date stamp on your dataset against the exercise requirements. When you are working through these on your own, start by running the basic descriptive statistics the exercise mentions before touching the regression. The textbook often hints at outlier issues in the data description. One exercise in Chapter 8 involving the relationship between test scores and expenditure ratios had a single district with unusually high spending that was driving your R-squared upward. Running summary() in R or summarize in Stata before regression saves you from chasing phantom relationships.
The solutions you find online tend to show the final output without the intermediate steps. That is the least helpful format. What matters is documenting your data cleaning process, your variable construction, and your diagnostic checks. Your professor can tell whether you actually worked through the problem or just pasted numbers. They also tend to ask about robust standard errors in follow-up questions, which means you should learn to compute those even when the exercise does not explicitly ask. Common pitfalls to watch for include omitted variable bias when you drop a control variable to simplify your model, and specification error when you use levels instead of growth rates for non-stationary time series data. The textbook mentions these issues but does not always flag them prominently within each exercise. If your t-statistics look impossibly large for macroeconomic data, check whether you have difference-transformed your variables appropriately. For those struggling with specific chapters, the clearest approach is to work through each exercise sequentially and keep a running log of your commands and outputs. The process usually takes two to three hours per chapter on the first pass. After you understand the patterns, it drops to about forty-five minutes. There is no shortcut around understanding the mechanics, but there is a shortcut around wasting time debugging syntax errors.
Get the Full Details
