Planning Statistical Analysis Before You Collect Data

Most people wait until they have data in front of them before figuring out what test to run. That approach usually means a lot of backtracking, missed assumptions checks, or just picking whatever test is easiest rather than whatever is correct. A statistics planner exists to flip that order around. I've built and refined my own Top 10 Statistics Planner spreadsheet years ago because the available tools at the time either required paid software or made you fill out forms that took longer than the analysis itself. The basic concept is simple: you map out your study design upfront, then the planner tells you which test applies, what the assumptions are, how to check those assumptions, and what output to expect.

Why a Top 10 Statistics Planner Makes Sense

The real problem isn't picking the wrong test after you see the data. It's collecting data without knowing whether your design can actually answer the question. Power analysis depends on effect size estimates, and if you haven't thought through your statistical test before running subjects, your sample size calculation is just a guess. Here is how I actually use the planner in practice:

I start by writing down the research question as a one-sentence statement. Then I identify the variables: what is independent, what is dependent, and whether the design is between-subjects or within-subjects. From there I look at the measurement level of each variable. This determines the test family before anything else. The planner then routes you through a decision tree. Continuous dependent variable, one categorical independent variable with two groups? Independent samples t-test. Same setup but the groups are related? Paired samples t-test. More than two groups? One-way ANOVA. The tree catches most of the standard cases and flags when your situation falls outside the top ten list.

Building Your Own Version

You do not need to buy anything for this. I keep mine as a Google Sheet with five main sections.

The first section is a variable declaration table. Every variable gets a row with columns for name, measurement level (nominal, ordinal, interval, ratio), role (independent, dependent, covariate), and expected distribution shape. This sounds tedious but it saves you from realizing after data collection that you treated a Likert scale as interval when it is really ordinal. The second section is the test selection matrix. I list the common tests side by side with their assumption requirements. T-tests require independence, normality of the dependent variable, and homogeneity of variance for the independent version. ANOVA adds the same variance assumption across all groups. Mann-Whitney U and Kruskal-Wallis are the non-parametric alternatives when those assumptions break. The third section is where most people skip ahead and then regret it: the assumption checking protocol. For each test, I list the exact diagnostic steps. Shapiro-Wilk for normality. Levene's test for equal variances. VIF values if you are doing regression. These diagnostics need to happen before you run the main test, not after.

A Specific Problem I Ran Into

I learned through experience that the planner itself needs a version control log. Early on I had a study where I used an older version of the decision tree that did not account for covariates properly. I caught it because the summary output showed the ANCOVA path was missing from that revision. My workaround was to add a revision column to every template and a change log that notes what was fixed in each version. It takes about three extra minutes per update but has saved me from publishing flawed analysis plans twice now.

Common Pitfalls Beginners Miss

The biggest mistake is treating the planner as a one-time document. If your study evolves mid-project and you switch from a pre-registered plan to an exploratory analysis, the planner should reflect that change. Otherwise you are comparing your results against the wrong null distribution. Another issue is the interaction between multiple comparisons and planned tests. When you plan more than one hypothesis test, the planner should include a correction method column. Bonferroni is the default for most people but it is overly conservative when tests are correlated. Holm-Bonferroni gives you more power at almost no extra complexity. Tukey's HSD belongs in the post-hoc column, not the planning column, unless your test is specifically designed as a post-hoc procedure.

What the Top 10 Statistics Planner Cannot Do

It does not run the analysis for you. It does not check your actual data for assumption violations because it has no access to your data until you put it in. It also cannot handle multilevel models, structural equation modeling, or Bayesian approaches unless you expand the template significantly beyond the top ten tests. Those models need specialized planning tools or direct consultation with a statistician. If your study involves repeated measures with missing data patterns or nested clusters, the basic planner will give you the wrong guidance. In that case you should move to a dedicated power analysis tool like G*Power or plan the model structure in R using the simr package before committing to a sample size.

Getting Started Quickly

Open a blank spreadsheet. Create the three sections I described above. List the ten most common tests your field uses and fill in the parameters for each one. Leave empty columns for your specific study variables so you can drop in new data without rebuilding the sheet. Test the planner on a completed study from your past to verify it matches the tests you actually ran. If it does not match, you have found a gap in the decision tree. Fix that gap before using it for a live project. The Top 10 Statistics Planner is not a magic solution. It is a friction reducer that stops you from making decisions under pressure after data collection is complete. That small shift in timing usually makes the difference between a clean analysis and a three-day debugging session.