A Practical Guide to Shadish Cook Campbell Experimental Design
The book everyone in program evaluation and causal inference keeps referencing but few people actually read cover to cover is Shadish Cook Campbell. Full title: Experimental and Quasi-Experimental Designs for Generalized Causal Inference. It was originally published in 1979 and revised in 2002. The authors are William Shadish, Thomas Cook, and Donald Campbell. This is not a quick read. It is the single most comprehensive treatment of causal inference design that exists, and it shows up constantly in methodology courses, grant review panels, and evidence synthesis work. The book catalogs every standard experimental and quasi-experimental design you will encounter in social science research. Internal validity threats, external validity threats, statistical conclusion validity, construct validity. The four validity framework is where most people get stuck. Campbell himself developed the original version, and Shadish and Cook expanded it significantly in the 2002 revision. The threat taxonomy runs through selection bias, maturation, history effects, testing effects, instrumentation drift, regression to the mean, attrition, diffusion of treatment, and a dozen others. Each one gets a full explanation of when it bites you and what design feature blocks it. The real value is not the list. It is the design matrix. The book organizes designs by how many groups they use, whether assignment is random or nonrandom, and how many pretest and posttest measurements exist. You get from two-group posttest-only designs all the way through time series with control groups and Solomon four-group designs. The comparison tables at the end of each chapter tell you which validity threats each design handles and which ones leave you exposed.
How I Actually Use This Book
I keep a printed copy on my desk. I do not read it sequentially. When a new evaluation lands on my desk, I go straight to the design catalog and match my situation to one of the templates. The 2002 edition has better coverage of quasi-experimental methods than the original, especially around difference-in-differences, regression discontinuity, and propensity score matching as design-enhancing strategies rather than standalone fixes. Here is a specific problem I ran into last year that the book helped me untangle. We were evaluating a workforce training program where participants self-selected into treatment and a comparison group came from a separate administrative database. The obvious move was a standard difference-in-differences approach. But when I mapped the timeline, I realized the comparison group had been exposed to a parallel policy intervention in the same month the treatment group started. That is a history threat, and it violates the parallel trends assumption that DiD requires. Shadish Cook Campbell does not give you a shortcut out of this. It makes you name the threat explicitly and then either find a different comparison source, add a second pretest wave to test parallel trends more rigorously, or accept that your internal validity is compromised and frame the results accordingly. I added a second pretest period and ran a placebo test on the pre-period coefficients. The placebo failed. We dropped the comparison group and switched to a regression discontinuity design based on an enrollment cutoff score in the program application. The redesign took about three weeks of additional work and cut our validity argument from "somewhat credible" to "conditionally credible with stated limitations."
Common Mistakes People Make With This Framework
The biggest mistake I see is treating internal validity as the only thing that matters. It is not. A design can eliminate every internal validity threat and still produce results you cannot generalize anywhere. External validity is where most evaluation reports quietly fail. The book spends considerable time on population validity, ecological validity, treatment manipulation validity, and setting validity. In practice, this means if you test an intervention in a controlled lab setting with motivated volunteers, your causal claim is tight but your ability to predict what happens when you roll it out in the wild is essentially zero. Another mistake is assuming random assignment solves everything. It does not. Random assignment handles selection bias. It does not handle attrition, it does not handle treatment contamination, and it does not handle measurement error. I have seen three separate programs where randomization was achieved but the treatment group migrated into the comparison condition within six weeks, completely contaminating the estimate. The design was randomized. The implementation was not. A third mistake is reading the threat list as a checklist to defeat rather than a framework for thinking. You cannot defeat all threats simultaneously. Every design choice that strengthens one validity dimension weakens another. Adding a second control group improves internal validity but complicates external validity because you now have two different comparison contexts. Adding pretests improves detection of selection bias but introduces testing effects that contaminate your outcome measure.
Get the Full Details
Quasi-Experimental Designs Explained Practically
The 2002 revision expanded the quasi-experimental section considerably. Non-equivalent control group designs remain the most common approach in field evaluation, and the book gives you the full taxonomy: nonequivalent groups posttest-only, pretest-posttest, multiple time series, and combinations thereof. The key insight that beginners miss is that quasi-experiments do not fail because they lack randomization. They fail when the researcher pretends the design provides the same causal leverage as a randomized experiment. The book is clear about this throughout. A non-equivalent control group design with a single pretest and posttest has a very narrow validity window. You need multiple pretest observations, a plausible mechanism, and a strong argument that the groups would have followed parallel trajectories absent the treatment. The regression discontinuity section is worth reading carefully. Many people treat RD as a standalone method. Shadish Cook Campbell frames it correctly as a design that borrows strength from both experimental and quasi-experimental logic. The causal identification comes from the treatment assignment rule, not from balancing covariates. The bandwidth choice, the polynomial order, and the manipulation test are where the design lives or dies. Most published evaluations get these wrong.
Where the Book Falls Short
The 2002 edition does not cover modern causal inference tools like principal stratification, instrumental variable designs with multiple instruments, or the synthetic control method. If your work involves panel data with unit-level fixed effects and heterogeneous treatment effects, you will need to supplement this with Angrist and Pischke or Imbens and Rubin. The book also predates the replication crisis literature, so it does not address p-hacking, publication bias, or the file drawer problem. Those issues are real and they change how you should interpret any single study's validity claims, regardless of design quality. Another gap is that the book treats each design in isolation. In practice, evaluations often combine features from multiple designs. A stepped wedge rollout, a clustered randomized trial with a non-equivalent comparison site, a difference-in-differences analysis layered on top of a regression discontinuity—these hybrids are common and the book does not give you a template for evaluating them as integrated units. You have to apply the validity framework piece by piece.
Where to Find the Book
The 2002 edition is published by SAGE Publications. You can purchase it through standard academic retailers. Libraries frequently carry it. Some universities provide institutional access through their library systems. There is no legitimate free PDF of this book available, and you should be cautious of any site offering it without a publisher link. The original 1979 edition is in the public domain in some jurisdictions and may be available through archive.org, but the 2002 revision contains substantial additions that are not in the first edition. If you are doing serious evaluation work, the 2002 edition is the one you need. First, a chapter on transportability and generalizability frameworks. The validity discussion is necessary but not sufficient for deciding whether your causal estimate applies to a different population or setting. Pearl's transportability theory and the work by Bareinboim and Pearson provide formal tools that complement the Shadish Cook Campbell framework without replacing it. Second, coverage of sensitivity analysis for unmeasured confounding. Rosenbaum's work on hidden bias is relevant here, and the book mentions it only briefly. In observational studies, which is what most real-world evaluation is, sensitivity analysis is not optional.

Third, a discussion of pre-registration and design transparency. The 2002 edition assumes a research culture where design decisions are made and reported after the fact. That assumption no longer holds. Pre-registering your design, your validity threats, and your planned analyses changes how reviewers should interpret your work, and the book has nothing on this. The core framework remains sound. The validity taxonomy is still the standard reference point in methodology programs across the social sciences. The design catalog is still the best single-source comparison of experimental and quasi-experimental options. It is dense, it is dry, and it is exactly what you need when a funder asks whether your evaluation design can support a causal claim.