What most people get wrong about statistics

I spent years cleaning datasets for a living, and the single biggest problem I kept running into was that people treat statistics like a set of commands you run in sequence. It isn't. It's a process of checking assumptions, breaking models, and admitting when the data doesn't want to cooperate. The Top 10 Statistics Tips most guides list are fine as a checklist, but they rarely cover the things that actually go wrong in practice. I'll get to the list after a few notes from real work. 1. Check your distributions before you pick a test. This sounds obvious until you've seen someone run a t-test on income data that looked roughly normal in a histogram but had a log-normal tail that completely violated the assumption. Run Shapiro-Wilk or at least Q-Q plots. If the tails are fat, switch to Welch's t-test or a permutation approach. This alone saved me from publishing a false positive on a project involving survey response times. 2. Normalize or standardize before you cluster. I once ran K-means on a dataset with features ranging from 0-1 to 0-10,000 and wondered why the clustering was garbage. StandardScaler did what it should have done from the start. Without scaling, distance-based methods just follow the highest-variance feature like a dog chasing a car.

3. Don't trust p-values without effect sizes. A p-value of 0.03 means nothing if Cohen's d is 0.08. I worked on a healthcare analytics project where a drug showed statistically significant improvement with a sample of 50,000 patients, but the actual reduction in symptoms was clinically irrelevant. Always report confidence intervals alongside p-values. Your audience won't thank you for giving them a number that sounds impressive and means nothing. 4. Handle missing data by understanding why it's missing, not by just dropping rows. Listwise deletion can delete half your dataset and leave you with selection bias. I've seen MCAR assumptions destroyed because the missingness correlated with the outcome itself. Use multiple imputation with chained equations, or at the very least, flag missingness as its own category if it makes sense for the domain. In one project, missing survey responses correlated with lower income, and dropping them skewed the entire model upward. 5. Outliers are information, not noise. I spent a week debugging a regression model that kept underperforming until I realized the outliers weren't errors—they were a different customer segment entirely. Split them out, model them separately, then decide. Blindly truncating or Winsorizing hides the structure you're actually trying to find.

6. Beware of multiple comparisons inflation. Run ten tests at alpha 0.05 and you'll get one false positive by chance. Run fifty and you'll get roughly two. Use Bonferroni correction, Holm-Bonferroni, or Benjamini-Hochberg depending on how many tests you're running. I learned this the hard way on a marketing attribution project where we tested 30 funnel segments and reported "significant" results that vanished on replication. 7. Cross-validation beats train-test splits for small datasets. A single 80-20 split can give you wildly different results depending on which points land where. Use k-fold or repeated k-fold cross-validation. With datasets under 1,000 samples, I usually go with 10-fold repeated five times. It takes longer but gives you a distribution of performance estimates instead of a single number that might be lucky or unlucky. 8. Collinearity destroys interpretability, not prediction. If your goal is forecasting, VIF scores don't matter much. If you're explaining why something happens to a stakeholder who needs to make decisions, correlations above 0.7 between predictors will make your coefficients flip signs and lose meaning. Check VIF, drop or combine features, or switch to regularization. Ridge regression handles collinearity better than ordinary least squares without requiring you to throw away data.

Get the Full Details

Top 10 Statistics Methodologies for Data Science | Sunil Kumar Yadav ...
Top 10 Statistics Methodologies for Data Science | Sunil Kumar Yadav ...

9. Document your preprocessing pipeline like it's code. I've lost count of the number of times I came back to a project six months later and couldn't reproduce my own results because I'd applied log transformations in one notebook and standardization in another, then merged them without tracking the order. Use sklearn Pipeline objects or equivalent in R. It turns your data cleaning into version-controllable, reproducible steps instead of a memory game. 10. Validate with out-of-sample data, not just in-sample fit. R-squared of 0.95 sounds great until you see the test set performance drop to 0.62. This is the most common failure mode I see, especially from people coming from academic backgrounds where in-sample fit is treated as the final answer. Real-world data shifts. Train on one period, validate on another. If you're working with time series, never shuffle—use temporal splitting.

Where these tips break down

Here's the part most guides don't mention. These tips assume your data is reasonably clean, your sample size is adequate, and your question is well-defined. None of those conditions are guaranteed. When your sample is under 30, even permutation tests struggle. When your features are high-dimensional relative to observations, regularization helps but introduces its own bias. When your missingness mechanism is MAR or MNAR, multiple imputation can still be biased unless you model the missingness mechanism explicitly. I had a case where all ten of these tips were applied correctly and the model still failed in production because the training period didn't capture a seasonal pattern that appeared three months later. The statistical approach was sound. The temporal generalization wasn't. No amount of cross-validation within the training window would have caught that without a proper holdout period that spanned the seasonal boundary. Another edge case: when dealing with count data that has heavy zero inflation, standard negative binomial regression underperforms compared to zero-inflated models. I spent days troubleshooting why residuals showed systematic patterns before someone pointed me at the zero-inflation diagnostic. The tip "check your assumptions" covers this, but the fix isn't always obvious from the assumption check alone.

Practical recommendations based on dataset type

If you're working with small datasets under 500 rows, focus on tips 3, 7, and 9. Validation and reproducibility matter more than complex preprocessing. If you're working with high-dimensional data like text or image features, tips 2, 8, and 10 need extended treatment with dimensionality reduction and careful regularization tuning. For time series data, temporal structure overrides almost everything else—tip 7 needs to become blocked cross-validation and tip 3 needs to account for autocorrelation in significance testing. The short version is that statistics isn't a toolkit you apply in order. It's a workflow where each step informs the next, and skipping ahead because a tip sounds important usually comes back to cost you twice as much time later. I've seen people jump straight to model building on every project and then spend three weeks going back to fix data issues they should have caught in week one. The order of operations matters more than most guides admit.

10 Easy Statistics Tips For Beginners - Graphic Folks
10 Easy Statistics Tips For Beginners - Graphic Folks