Running a Between Groups Vs Within Groups Analysis
I spent about six months working through this after my first ANOVA implementation came back with results I couldn't explain. The short version is that between-groups variance measures how much your group means differ from each other, while within-groups variance measures how spread out the individual data points are around their own group mean. When you run an F-test, you're dividing the between-groups number by the within-groups number. A big ratio means the groups actually look different. A small ratio means they probably don't. The formulas are straightforward enough, but getting them right in code is where things go sideways. Between-groups sum of squares uses the grand mean as a reference point. You subtract that from each group mean, square it, multiply by the group size, and add them all up. Within-groups sum of squares is just the sum of squared deviations for every single point from its own group mean. I still write this out by hand sometimes when setting up a new dataset just to make sure I haven't mixed up which variance goes where. Here's the part most tutorials skip: your within-groups variance is what you're stuck with no matter what. You can't reduce it by redesigning the experiment. It's your noise floor. The only thing you can really change is whether the signal between groups is large enough to rise above it. That's why sample size matters more for between-groups detection than most people realize. A small n per group means your within-groups variance estimate is unstable, and your F-statistic becomes unreliable even if the between-groups difference is real.
I ran into a specific issue once where I was comparing three treatment groups with roughly 12 participants each. The between-groups effect looked significant on paper, but when I checked the residuals, there was one outlier in Group B that was pulling the within-groups variance way up. That single point dropped the F-ratio from about 4.2 down to 2.8, making everything nonsignificant. The fix was running a robust ANOVA variant using trimmed means instead of raw scores, which is available in the WRS2 package for R. This shifted the p-value from 0.06 to 0.02 without removing any data points arbitrarily. Another thing that trips people up is assuming equal variances across groups by default. The standard F-test in between-groups versus within-groups analysis requires homogeneity of variance. If your groups have wildly different spreads, you should be using Welch's ANOVA instead, which adjusts the degrees of freedom to account for the imbalance. Most statistical packages will flag this, but they won't always tell you to switch methods automatically. When you report your results, you need to include both mean squares and degrees of freedom, not just the F-value and p-value. The mean square between divided by its degrees of freedom gives you the between-groups variance. Same for within-groups. Without those numbers, someone trying to verify your work or compute a post-hoc power analysis has nothing to go on. Effect size measures like eta-squared also come directly from the sum of squares values, so keeping track of them from the start saves you a step later.
The main limitation of this approach is that it tells you whether groups differ, but not which ones. You'll need post-hoc tests like Tukey's HSD or Bonferroni corrections to dig into pairwise comparisons, and each of those carries its own assumptions about variance structure and sample balance. If your design is heavily unbalanced, those post-hoc tests become less trustworthy unless you use a method that accounts for unequal cell sizes. For anyone working with small sample sizes, consider switching to a Bayesian hierarchical model instead. The frequentist between-groups versus within-groups framework can produce wide confidence intervals that are honestly not very informative at n below 20 per group. Bayesian estimation with weakly informative priors tends to give you more stable posterior distributions in those cases, though it does require more setup time upfront.
Get the Full Details
