Getting Your Data Checked Before Running a Regression

Most people skip assumption checking because they are worried about the p-values on their main effects. That is backwards. The p-values are meaningless if your data violates core assumptions hard enough to warp the model estimates. Tabachnick And Fidell Using Multivariate Statistics remains one of the most practical references for this, even though the 7th edition came out years ago and has not been updated. I still reach for their checklist when setting up a multiple regression or MANOVA. Their approach is not theoretical. It is a series of mechanical diagnostics you run in order, document, and then decide whether to proceed or transform. Here is how the process actually plays out in practice.

Checking Assumptions Without Losing Your Mind

Start with the data file open in whatever software you use. SPSS, R, Python, it does not matter. Generate a scatterplot matrix first. This takes thirty seconds and usually reveals nonlinearity, heteroscedasticity, or obvious multivariate outliers before you do anything else. You skip past it at your own risk. Next, look at pairwise correlations and VIFs. If any predictor pair exceeds 0.70, you are already in multicollinearity territory. VIF above 10 is the hard cutoff most people use. I flag anything above 5 and think about whether those variables are measuring the same construct poorly. Splitting them apart or combining them into a single index often fixes the problem faster than hoping the regression software will handle it gracefully. Then check residuals. Standardized residuals beyond plus or minus 3 indicate a likely outlier. studentized deleted residuals are more reliable because they account for the fact that a single extreme point can pull the regression line toward itself. Cook's distance above 1 is the traditional warning flag, though I treat anything above 0.5 as worth investigating. The rule of thumb is N divided by 4 for Cook's D, which gives a tighter threshold for smaller samples.

For multivariate outliers specifically, calculate Mahalanobis distance. In SPSS you get this from the regression diagnostics menu. Compare your values to the chi-square critical value with degrees of freedom equal to the number of predictors. For five predictors at alpha 0.001, the critical value is 20.515. Any case exceeding that is a multivariate outlier. I usually remove these one at a time and rerun the model to see if coefficients shift dramatically. If they do, you keep the case flagged and report both models. Normality of residuals matters less than people think with large samples. With fewer than fifty cases, run the Shapiro-Wilk test and look at Q-Q plots. A few deviations in the tails rarely matter. Systematic curvature in the Q-Q plot does. That means your model specification is wrong, not just that your errors are slightly skewed.

Get the Full Details

Using Multivariate Statistics : Tabachnick, Barbara G., Fidell, Linda S ...
Using Multivariate Statistics : Tabachnick, Barbara G., Fidell, Linda S ...

The Problem I Ran Into With Tabachnick And Fidell Using Multivariate Statistics

My dataset had 142 cases and eleven predictors. Mahalanobis distance flagged twelve cases above the chi-square cutoff. Standardized residuals flagged fourteen. The overlap was messy. I followed the prescribed process of removing flagged cases one by one, but each removal changed the regression solution enough that the next round of diagnostics looked completely different. This went on for three cycles and the model never stabilized. The workaround was straightforward once I accepted that the data violated multivariate normality in a way that no amount of deletion would fix. I switched to bootstrapped regression estimates. In R, the quantreg package handles this cleanly. I ran 5,000 bootstrap samples, pulled the bias-corrected confidence intervals, and compared them to the original OLS coefficients. The point estimates barely shifted. The standard errors widened slightly, which made two previously significant predictors drop below 0.05. I reported the bootstrap results and noted the violation in the limitations section. That was faster and more honest than continuing the deletion cycle. If you are using SPSS, go to Analyze > Regression > Linear, then check the box for Mahalanobis under Distance. It outputs the distances as a new variable you can sort and compare to the chi-square table. In R, you can compute it with mahalanobis() and the colvars argument, then use qchisq(0.999, df = n_predictors) to get your threshold.

What Nobody Tells You About These Diagnostics

High leverage does not always mean a bad data point. A case can have high leverage and still be perfectly valid. Leverage measures how far a case is from the centroid of your predictor space, not whether it is wrong. A legitimate participant with extreme but accurate scores will have high leverage. The influence statistic combines leverage and residual size. That is what matters. DFBETAS above 2 over the square root of N tells you which coefficient is being distorted by which case. Another counter-intuitive point: transforming variables to fix normality of residuals is almost never the right move. Transformations change the interpretation of your coefficients in ways that make your model harder to communicate. Better to address non-normality by adding missing predictors, checking for curvilinear relationships, or switching to robust standard errors. A square root transformation will not rescue a misspecified model. Heteroscedasticity is easier to fix. The Breusch-Pagan test tells you whether variance changes with predicted values. If it does, weighted least squares or heteroscedasticity-consistent standard errors do the job. HC3 standard errors in R's car package are my default choice now. They adjust the covariance matrix without requiring you to restructure the model.

The big limitation of the Tabachnick and Fidell framework is that it assumes you are running regression or MANOVA type analyses. If you are doing structural equation modeling, factor analysis, or mixed effects modeling, the diagnostic rules shift considerably. Mahalanobis distance still applies, but fit indices and modification indices replace Cook's D as your primary concern. The book covers some of these in later chapters, but it was written before many modern extensions became standard. For SEM, I recommend looking at Kline's Principles and Practice of Structural Equation Modeling instead. For mixed models, Pinheiro and Bates or the nlme documentation in R are more current. Also, this approach does not handle missing data well. The diagnostics assume complete cases. If you have more than five percent missingness, listwise deletion shrinks your sample enough to destabilize everything else. Multiple imputation before running diagnostics is the safer route. Run the assumption checks on each imputed dataset separately, then pool the results. It adds about twenty minutes to the workflow but prevents you from making decisions based on a biased subset of your data. I usually spend between forty-five minutes and two hours running through the full diagnostic sequence on a new dataset, depending on how many transformations or case deletions are required. If the model passes cleanly, that time pays for itself. If it fails, you avoid publishing results that will get torn apart in review.

Using Multivariate Statistics : Tabachnick, Barbara G., Fidell, Linda S ...
Using Multivariate Statistics : Tabachnick, Barbara G., Fidell, Linda S ...