Most People Learn Statistics Backwards
They memorize formulas before they understand what they're actually measuring. I watched a junior analyst once spend forty-five minutes computing a standard deviation by hand before realizing her dataset had twelve missing values coded as zeros. The mean was destroyed. The standard deviation was meaningless. This is the first thing you need to understand: the math is the easy part. Knowing when not to run a test is what separates people who can calculate from people who can actually make decisions. The following list isn't ranked by importance because that depends entirely on your domain. A clinical researcher needs power analysis and confidence intervals more than she needs regression diagnostics. A marketing analyst lives in A/B testing and correlation. But these ten concepts appear with enough regularity that ignoring any one of them will cost you eventually. Mean, median, and mode are the foundation, yet most people stop at mean and move on. The median is more robust to outliers than most analysts give it credit for. I was analyzing user session duration last year across a product feature rollout. The mean suggested average sessions jumped from 4.2 minutes to 8.7 minutes after the update. The median only went from 2.1 to 2.8. The mean was being pulled upward by a small number of extremely long sessions from bots. If you only report the mean, you're reporting a distorted picture. Always check both.
Standard deviation and variance measure spread, not accuracy. A low standard deviation doesn't mean your measurements are close to the true value. It means they're clustered tightly, whether that cluster is centered on the truth or not. I saw this play out in an A/B test where variant B showed dramatically less variance in conversion rate than variant A. Everyone celebrated the consistency. What they missed was that variant B's entire distribution was shifted downward. Tight clustering around the wrong number is still the wrong number. Confidence intervals get misinterpreted constantly. A 95% confidence interval does not mean there is a 95% probability that the true parameter lies within your calculated interval. The true parameter is a fixed value. Your interval is the random variable. The correct interpretation is that if you repeated your sampling process infinite times, 95% of your intervals would contain the true parameter. Most introductory textbooks present it differently because the formal definition is unwieldy. That's why you'll see people say "we're 95% confident" as shorthand, but the distinction matters when you're defending your methodology to someone who knows statistics well enough to catch sloppy language. Hypothesis testing and p-values are where the field has the most cultural baggage. A p-value of 0.048 and a p-value of 0.052 are functionally identical. The arbitrary cutoff at 0.05 exists because Ronald Fisher chose it in the 1920s, not because nature draws a line there. I've rejected real results because they sat at 0.051 and accepted clearly spurious correlations at 0.049. The practice is worse than the theory. When I needed to move past this, I started reporting exact p-values alongside confidence intervals and effect sizes. It forces everyone reading the analysis to look at the full picture instead of playing binary pass-fail.
Correlation does not imply causation is the most recycled sentence in statistics. It's also the one people ignore the most because confirmation bias makes it feel like causation anyway. In 2022 I was asked to analyze the relationship between employee training hours and quarterly sales per region. The correlation was 0.71. The natural read was that training drove revenue. The actual story was that high-performing regions got more training budget because they had more revenue to spend. Sales funded training, not the other way around. Controlling for region size and prior revenue eliminated the apparent effect entirely. Regression and regression to the mean are two different things, though beginners conflate them constantly. Regression to the mean means extreme observations tend to be followed by less extreme ones simply due to random variation. This is the reason you can't evaluate a new manager's impact by comparing their first quarter to the previous quarter if that previous quarter was unusually bad. The regression existed independently of any policy change. When building actual regression models, multicollinearity between predictors is the silent killer. Two variables can both be highly correlated with your outcome and with each other, making individual coefficients unstable and direction unreliable. Variance inflation factors above 5 are worth investigating. Above 10, your model is unreliable for inference even if predictions look fine. Sampling methods determine whether your results generalize at all. Convenience sampling is the default in most non-academic environments and it produces results that don't generalize to anything except the specific moment you collected them. A proper simple random sample is rarely practical. Stratified sampling, where you ensure each subgroup is proportionally represented, is usually the minimum standard. I worked on a survey where 73% of respondents were from one demographic because that's who had internet access at the time of the poll. Weighting the data afterward partially corrected it, but you cannot fix selection bias after the fact. The fix has to happen at collection.
Get the Full Details

Chi-square tests check independence between categorical variables. They're straightforward but have a quiet assumption: expected frequencies in each cell should generally be at least 5. When you have sparse contingency tables with many categories and small samples, chi-square breaks down. I hit this when cross-tabulating product return rates across twelve store regions with under 20 returns per region. The test was producing nonsensical significance. Fisher's exact test handled it, though it gets computationally expensive with tables larger than about 5 by 5. There's no universal shortcut for large sparse tables. Bayesian statistics offer an alternative framework that's often more intuitive for decision-making but require priors, which introduces subjectivity at the start. The Bayesian approach updates beliefs with data rather than testing whether data contradicts a null hypothesis. In practice, this means you can directly calculate the probability that a treatment effect is positive, which is the actual question people want answered. The trade-off is that your prior choice can influence results when sample sizes are small. I used a weakly informative normal prior centered at zero with a standard deviation of 2.5 on the log-odds scale for a medical study, and the results were nearly identical to the frequentist analysis because the data dominated. That might not hold in smaller studies. Simpson's paradox occurs when a trend appears in several subgroups but reverses when the subgroups are combined. It's not a bug. It's a feature of aggregated data masking underlying structure. A hospital reported that surgery A had a higher overall mortality rate than surgery B. When broken down by patient severity, surgery A had lower mortality in both mild and severe categories. The confounding variable was that surgery A treated more severely ill patients, dragging its aggregate rate up. This is why you should never trust an aggregate statistic without checking whether subgroup breakdowns tell a different story.
The biggest limitation across all of these examples is that statistics describes patterns in observed data. It cannot tell you whether your data is representative, whether your measurements are valid, or whether the relationships you find are causal rather than coincidental. No statistical method fixes bad data collection. The tools above are designed to extract signal from noise, not to generate signal from nothing. If you skip the design phase and jump straight to analysis, you're just automating the production of confidently wrong answers.