Effect Size Is The Thing Nobody Checks Until It's Too Late
You run a t-test. The p-value comes back at 0.003. You publish the result. Six months later, someone asks what the actual difference was, and you realize you have no idea. Not because you didn't calculate it, but because nobody ever taught you to look for it. Effect Size Statistics quantify the magnitude of a relationship or difference, independent of sample size. This matters because with enough subjects, even a meaningless difference becomes statistically significant. A study with 50,000 participants might find a "significant" difference of 0.3 IQ points. The p-value says something happened. The effect size says it doesn't matter. These two metrics answer completely different questions, and reporting only one of them gives a misleading picture of what your data actually shows.
Size Of Effect Statistics: A Practical Guide To Calculating And Interpreting Them
The most common effect size measure is Cohen's d, which is simply the difference between two means divided by the pooled standard deviation. Here is the formula as you would type it into R: cohens_d = (mean_group1 - mean_group2) / sd_pooled Pooled standard deviation uses the formula sqrt(((n1-1)*sd1^2 + (n2-1)*sd2^2) / (n1+n2-2)). You can also use the effsize package in R, which handles this in one line with the effsize::cohen.d() function. In Python, the scipy.stats module has a gpower submodule, but many people just use the statsmodels library: statsmodels.stats.effect_size.tt_ind_solve_effsize() for two independent groups.
Cohen's conventions of 0.2, 0.5, and 0.8 for small, medium, and large effects are useful as rough anchors but should not be treated as universal truth. In clinical psychology, a Cohen's d of 0.3 might represent a meaningful treatment outcome. In industrial quality control, you would need a d of 1.5 or higher before the result justifies changing a process. Context determines what "large" actually means for your work.
When The Standard Measures Fail You
I spent two days last year wrestling with a within-subjects design where the variance in the difference scores was essentially zero for one condition. The standard Cohen's d formula broke down because dividing by a near-zero pooled standard deviation produced an effect size of 47.3, which is obviously wrong. What actually happened was that one participant had an extreme outlier in the control condition that inflated the standard deviation massively while doing nothing to the mean difference. The workaround was to use Glass's delta instead, which uses only the control group standard deviation rather than the pooled value. In R, that's straightforward with the effsize package by specifying the control group. If your data has heteroscedasticity, which is extremely common in real-world samples, Glass's delta or Hedges' g (a bias-corrected version of Cohen's d) will give you a more honest number. Hedges' g applies a correction factor that matters most when your total sample is under 20, and it converges toward Cohen's d as N increases. For correlation-based designs, Pearson's r is already an effect size. There is no separate calculation needed. Values range from -1 to 1, and the interpretation is direct. But if your data is ordinal or not linear, Spearman's rho serves the same purpose and is often more appropriate. A common mistake I see is people converting r to Cohen's d using the formula d = 2r / sqrt(r^2) and then treating the result as if it came from a raw mean comparison. It does not. The conversion assumes a bivariate normal distribution, and when that assumption is violated, the converted effect size is garbage.
Confidence Intervals Are Where The Real Information Lives
Most people report a point estimate and stop. An effect size without a confidence interval is like reporting a weather forecast as exactly 72 degrees with no mention of the margin of error. The CI tells you the precision of your estimate, which is usually more useful than the point estimate itself. In R, the effsize package gives you CIs automatically with the conf.level parameter. In Python, statsmodels provides ci_bounds for most effect size functions. A 95% CI that ranges from 0.1 to 1.8 for a Cohen's d is genuinely uninformative. It means your study could support anything from a trivially small effect to a very large one. This is not a failure of the effect size measure. This is your study being underpowered, and the CI makes that visible in a way the p-value never will.
Common Pitfalls That Waste Time
The first trap is using a one-tailed p-value and then reporting a two-tailed effect size. These are not consistent with each other. If your hypothesis was directional, state it upfront and make sure all downstream calculations match. The second trap is applying Cohen's d to binary outcome data. Use odds ratios or risk ratios instead. They communicate the same information but in a scale that clinicians and decision-makers actually understand. A Cohen's d of 0.4 on a binary outcome is harder to interpret than an odds ratio of 1.8, even though they describe the same underlying relationship. A third issue that comes up constantly: reporting standardized effect sizes when your measure has a natural unit. If you measured blood pressure in mmHg and the intervention reduced systolic pressure by 8 points, that is a clearer result than saying the Cohen's d was 0.65. Standardized measures are useful for comparing across studies. They are not universally superior. Raw mean differences with their confidence intervals are often more interpretable for practical decision-making.
Which Tool To Use When
If you are doing a quick analysis in R, install effsize and read effectsize. The effectsize package by Dominique Makowski is the current standard and supports a wide range of designs out of the box. For meta-analysis work, the metafor package handles everything from Cohen's d to eta-squared to Fisher's z transformations. In Python, the statsmodels package covers the basics. For anything beyond a simple two-group comparison, consider using JASP or G*Power if you need pre-study power calculations alongside your effect size estimates. They generate correct output without requiring you to derive formulas from scratch. The bottom line is that effect size statistics exist to tell you whether a finding matters, not just whether it exists. P-values handle existence. Effect sizes handle importance. They are not interchangeable. Anyone who tells you otherwise is either confused or trying to sell you something.
Get the Full Details
