What Actually Comes Up When They Ask You Statistics Questions
The first thing you need to understand is that most data science statistics interview questions are designed to test whether you can reason through uncertainty, not whether you can recite formulas from memory. I've sat on both sides of that table. The candidates who do well are the ones who think out loud and admit what they don't know rather than bluffing through a derivation they half-remember. Here is how this actually plays out in practice. You will get asked about p-values, confidence intervals, Bayesian versus frequentist thinking, A/B testing design, and hypothesis testing frameworks. Sometimes they ask the textbook version. More often they ask a messy version where the data has missing values, the groups are unbalanced, or the assumption of normality is clearly violated. That is the point. They want to see if you can adapt. I remember one interview where they asked me to design a test for a feature that changed the user signup flow. The catch was that the treatment and control groups had vastly different baseline conversion rates because of how users were assigned. Most candidates immediately started talking about a standard two-proportion z-test. That was the wrong answer. What actually worked was a logistic regression with group-level random effects and a check for covariate balance first. I walked them through that reasoning step by step. They offered me the job after that question. The rest of the interview was just formality.
Data Science Statistics Interview Questions
Let me break down the categories and what they are really looking for. Hypothesis testing and p-values come up constantly. The basic question is straightforward: explain what a p-value is and what it is not. A p-value is the probability of observing data at least as extreme as what you got, assuming the null hypothesis is true. It is not the probability that the null hypothesis is true. It is not the probability that your alternative is false. These distinctions matter because people who confuse them will make terrible decisions in production. Here is a specific edge case that catches people out. You run an A/B test and get a p-value of 0.04. You declare significance and ship the feature. Two weeks later the effect disappears. What happened? The most common reason is that the test ran long enough for seasonal effects to creep in, or the sample size grew so large that trivial differences became statistically significant while the actual business impact was negligible. The workaround I use is to define an effect size threshold before the test starts and treat anything below it as noise, regardless of the p-value. Statistical significance and practical significance are not the same thing. They rarely are.
Confidence intervals are another bread-and-butter topic. A 95% confidence interval does not mean there is a 95% chance the true parameter is in your interval. It means that if you repeated the experiment an infinite number of times, 95% of the intervals you construct would contain the true parameter. This is a subtle distinction but one that separates people who understand statistics from people who just memorized definitions. In practice, I have seen teams misinterpret narrow confidence intervals as proof of precision when the real issue was a biased sampling method. If your survey only reaches email subscribers, your confidence interval might be 0.3 ± 0.01, but that is meaningless if your population is all users. Wide intervals from good data beat narrow intervals from bad data every time. Bayesian versus frequentist approaches will show up, usually in the form of "which do you prefer and why." The honest answer is that it depends on the problem. Frequentist methods are simpler to implement and easier to explain to stakeholders. Bayesian methods handle small sample sizes better and give you a natural way to incorporate prior knowledge. If you are building a recommendation system with sparse user data, Bayesian smoothing can make or break your model. If you are running a simple control versus treatment experiment with thousands of users per group, a t-test will get you the same answer faster.
Get the Full Details

I once had to debug a Bayesian model that was producing wildly incorrect posterior distributions. The issue turned out to be an improperly specified prior that was too informative relative to the actual signal in the data. The fix was running a prior predictive check before fitting the model. That is something most people never learn until they break something in production. A/B testing design is where theory meets reality and things get messy. The canonical questions involve sample size calculation, randomization strategy, and multiple testing correction. You need to know how to compute power, how to use the formula n = 2²(z_/2 + z_)²/², and when that formula is wrong. It assumes normality, equal variance, and independent observations. Real data violates all three sometimes. One thing beginners consistently miss is the peeking problem. If you check your results after 100 users per group, see a promising p-value, and then keep going, you have inflated your false positive rate. The fix is either a fixed sample size determined before the test or a sequential testing method like the alpha spending function. I use the latter when I cannot predict how long data collection will take, which is most of the time in practice.
Regression assumptions are fair game too. You should be able to look at residual plots and tell whether the assumptions hold. Homoscedasticity, linearity, independence, normality of errors. If the residuals fan out, you have heteroscedasticity. If they curve, you have a non-linear relationship. If they are autocorrelated, your observations are not independent. Each of these problems has a standard remedy: weighted least squares, polynomial terms, or mixed effects models with clustering. I worked on a project where the dependent variable was a count with heavy skew. Ordinary least squares gave terrible predictions because the variance was proportional to the mean, not constant. Switching to a Poisson generalized linear model fixed it in one line of code. The interview question that led there was just "what do you do when OLS residuals show a pattern?" The answer is not "fix the model" in the vague sense. It is specifying the right conditional distribution and link function. Multivariate statistics shows up less often but when it does, it tends to be about PCA, dimensionality reduction, or handling correlated features. The key insight here is that PCA is not a feature selection method. It creates new features that are linear combinations of the old ones. If your goal is interpretability, PCA makes things worse, not better. If your goal is prediction and you have multicollinearity, it helps a lot.
There is also a common trap with p-values in high-dimensional settings. When you run 100 tests at = 0.05, you expect about five false positives by chance alone. The Bonferroni correction is conservative. The Benjamini-Hochberg procedure controls the false discovery rate and is usually more appropriate for data science work. Knowing the difference between family-wise error rate and false discovery rate is a sign that you have actually done this work. When preparing for these interviews, practice explaining concepts without using jargon first. If you can describe a confidence interval to someone who has never taken statistics, you understand it. Then add the technical details. Most candidates jump straight to the technical details and reveal that their understanding is surface-level. Another practical tip: bring a concrete example from your own work when they ask open-ended questions. Not a tutorial project. A real one where something went wrong and you fixed it. I once talked through a time when my randomization was broken because of a timezone bug that assigned users to treatment based on their local time rather than UTC. The result was a systematic difference in user behavior between groups that looked like a treatment effect until someone noticed the pattern. That story came up in three separate interviews and it always impressed people because it showed I had shipped code, not just run notebooks.

The statistics portion of a data science interview is not about getting every answer right. It is about demonstrating that you can think rigorously under uncertainty, recognize when a standard method is inappropriate, and communicate your reasoning clearly. The candidates who struggle are the ones who memorize procedures without understanding when those procedures fail.