Working with WWC Evidence Standards Is a Different Beast Than You Think
I spent three years helping research teams prepare datasets for WWC review before I started doing it myself. The math behind their evaluation framework isn't complicated in isolation, but it's incredibly unforgiving when your study design has even a small flaw. Most people approach this backwards. They try to make their data fit the WWC standards instead of understanding what the standards actually require before collecting a single data point. The core of What Works Clearinghouse Math centers on how effect sizes are calculated, how violations are weighted, and how studies get sorted into one of four evidence tier ratings. That's the easy summary. The hard part is everything that happens between the raw data and the final rating.
Understanding What Works Clearinghouse Math
WWC uses Cohen's d as their primary effect size metric, standardized across studies so comparisons are possible. They calculate it by taking the difference in adjusted posttest means between treatment and control groups, then dividing by the pooled standard deviation. Sounds straightforward until you deal with multiple outcome measures, attrition adjustments, or clustered designs. Here's where most people stumble. When a study has more than one outcome measure for the same construct, WWC doesn't average them. They pick a single primary measure based on predetermined rules, and getting that wrong can change your entire rating. I had a team once who submitted a reading intervention study with twelve different literacy outcomes. We spent two weeks arguing over which one qualified as primary under WWC guidelines. The study ended up getting a "Meets WWC Standards with Reservations" rating partly because we'd flagged the wrong measure early on, which cascaded into incorrect violation counts.
The Three Numbers That Actually Matter
Effect size. Statistical significance. And the balance check statistic. Those are the pillars everything else builds on. Effect size tells you the magnitude of the impact. WWC generally looks for effects of 0.05 or greater to consider a program potentially meaningful, though they don't set a hard cutoff. Studies with zero or near-zero effects simply won't move forward regardless of statistical significance. Statistical significance under WWC typically means p less than 0.05, and they apply Bonferroni corrections when multiple outcomes are tested. This is where things get tricky. A study might show a genuinely large effect on its primary outcome, but if you measured ten secondary outcomes and only one hits significance after correction, the overall narrative suffers. The math penalizes fishing expeditions aggressively.
Get the Full Details

The balance check is the one nobody talks about enough. It measures whether treatment and control groups were equivalent at baseline before any intervention occurred. WWC calculates this using standardized mean differences on key covariates. If baseline imbalances exceed certain thresholds, the study gets marked as having a major violation, and that can tank your rating even if the intervention effect looks strong.
Practical Walkthrough: From Dataset to Rating
I'll walk through what actually happens when a study enters WWC review. Your dataset comes in with pretest and posttest scores for treatment and control groups. First step is checking eligibility criteria: randomized controlled trial or quasi-experimental design with comparison group, sufficient sample size, appropriate outcome measures. Next comes the attrition analysis. WWC has strict thresholds for acceptable participant loss. If more than five percent of participants drop out and that dropout rate differs between groups, you're looking at a potential major violation. I've seen perfectly good studies get downgraded to "No Evidence" ratings purely because attrition wasn't properly documented or was unevenly distributed. Then the effect size calculation. For continuous outcomes, you use the pooled standard deviation from the posttest. For binary outcomes, they convert to odds ratios and then to Cohen's d using a standard logistic-to-normal approximation. This approximation introduces a small amount of error, which matters when you're working with very small effects near the relevance threshold.
After effect sizes are computed, WWC reviewers check for design violations: selection bias, treatment implementation fidelity, measurement fidelity, and fidelity of implementation. Each violation can be classified as major or minor, and the combination determines whether the study meets standards, meets standards with reservations, or doesn't meet any standards.

Where the Math Breaks Down
Let me be honest about the limitations. WWC's framework struggles with heterogeneous treatments. If your intervention isn't delivered consistently across sites or teachers, the effect size you calculate is really an average across wildly different experiences, and the math treats it as a single clean number. It's not. The confidence interval around that estimate will be wider than WWC's standard formulas suggest. Longitudinal follow-up is another weak spot. WWC prefers short-term posttest measures because they're easier to standardize. Studies that show no immediate effect but significant gains six months later often get rated lower than they deserve because the framework wasn't built for delayed impact patterns. Small sample sizes are where I see the most frustration. WWC requires adequate statistical power, but their power calculations assume certain effect sizes and variance structures that rarely match reality in education research. A well-designed study with 150 participants might get flagged for insufficient power while a sloppy study with 500 participants sails through. The math doesn't always reward good design.
A Workaround I've Used Successfully
When dealing with clustered randomization where schools or classrooms are randomized rather than individual students, WWC's default calculations can inflate Type I error rates. The fix is applying the correct design effect to your standard errors before submitting. Multiply your standard error by the square root of one plus the average cluster size minus one times the intracluster correlation coefficient. I found this through trial and error after a reviewer rejected my initial submission specifically because I hadn't adjusted for clustering. Another thing that helps: submit your analysis code alongside your documentation. WWC reviewers sometimes recalculate your effect sizes independently, and if their numbers differ from yours, it creates unnecessary friction. Providing transparent, executable code that reproduces your results saves hours of back-and-forth. The What Works Clearinghouse Math isn't elegant. It's a practical framework built by people who needed a consistent way to sort through thousands of education studies with limited time and resources. It will miss nuances. It will downgrade solid studies for procedural reasons. But if you understand how it works and plan your research around its actual requirements rather than its stated goals, you can produce work that survives the review process intact.