Continuous Data in Biostatistics: What Actually Works
I used to teach biostatistics to public health students, and the thing that trips people up every semester isn't the math itself. It's the assumption-checking. Most students skip straight to running a t-test on their data without verifying whether the data actually meets the conditions required for that test. This happens because the procedures feel mechanical — plug numbers into a calculator, get a p-value, done — but the moment you look at the data closely, the whole framework starts to wobble. The core idea with continuous data is straightforward. You are dealing with measurements that can take any value within a range — things like blood pressure, weight, cholesterol levels, reaction times. These are not counts or categories. They are measured on a continuous scale. The statistical tools built for this kind of data assume certain things about how those measurements are distributed, and ignoring those assumptions is the fastest way to get wrong answers in any analysis. Here is the thing nobody warns beginners about: the Central Limit Theorem does not save you from everything. Yes, with large enough sample sizes, the sampling distribution of the mean approaches normality regardless of the underlying distribution. But "large enough" depends entirely on how skewed your data is. In practice, I have seen students use n = 30 as a hard cutoff and run parametric tests on heavily right-skewed data like income or length of hospital stay, where the mean is meaningless and the median tells a different story. The CLT does not magically make skewed continuous data behave nicely with small samples. If your data has extreme outliers or a long tail, increasing the sample size from 30 to 50 will not fix the problem. You either transform the data, use a non-parametric method, or model it with something like a generalized linear model.
Let me walk through how I approach this when I actually have data in front of me, not from a textbook. First, you describe the data before you test anything. Plot it. A histogram or a boxplot tells you more in ten seconds than a dozen summary statistics ever will. If you are comparing two groups, you need to check two things: normality of the residuals and homogeneity of variances. For normality, do not rely solely on the Shapiro-Wilk test when your sample is large. With n > 100, that test becomes extremely sensitive and will flag trivial deviations as significant, which drives students to over-correct. Visual inspection of a Q-Q plot is more useful here. For variance homogeneity, Levene's test is the standard, but again, with very large or very small groups the test loses power or becomes overly sensitive. A rule of thumb that works: if the ratio of the larger variance to the smaller variance exceeds 4, treat the assumption as violated. When the variances are unequal, Welch's t-test is the default adjustment. It modifies the degrees of freedom to account for the imbalance. Most statistical software runs this automatically now, but it is worth knowing why the output looks different from a standard Student's t-test. The p-value can shift enough to change your conclusion, especially with small samples.
For three or more groups, one-way ANOVA follows the same logic. The F-test assumes normally distributed residuals and equal variances across groups. If either assumption is violated, the Type I error rate can be inflated or deflated depending on the direction of the violation. I ran into this last year with a dataset where a treatment group had substantially higher variance than the control — not because the treatment was heterogeneous in effect, but because the measurement instrument was less reliable at higher values. Standard ANOVA gave a significant result, but the effect was mostly driven by variance differences, not mean differences. I switched to Welch's ANOVA, which handles unequal variances directly, and the result became non-significant. That is a case where running the textbook procedure would have led to a false conclusion. Non-parametric alternatives exist. The Mann-Whitney U test for two groups and the Kruskal-Wallis test for multiple groups do not assume normality. But they come with their own hidden assumption: the distributions in each group must have the same shape. If one group is skewed and another is symmetric, the test is not just comparing medians — it is testing stochastic dominance, which is harder to interpret in a clinical or public health context. I prefer transforming the data when possible. A logarithmic transformation often stabilizes variance and makes skewed continuous data much more tractable. It also produces results that are easier to communicate: a unit change on the log scale corresponds to a multiplicative change on the original scale. Regression with continuous outcomes, ordinary least squares, is where these issues compound. Every predictor in the model contributes to the residual structure, and violations of assumptions become harder to diagnose. Always check residuals against fitted values, not just the outcome distribution. A plot of residuals versus each predictor will reveal non-linearity that you would miss looking at univariate histograms alone. I once spent two weeks troubleshooting a model where the p-value for a key predictor was borderline, only to discover a clear U-shaped relationship with a continuous exposure variable. The linear term was masking a real effect. Adding a quadratic term resolved it. This is the kind of thing that never shows up in the worked examples in textbooks.
Get the Full Details

Missing data in continuous measures is another practical concern that gets short shrift in most courses. Listwise deletion sounds efficient but can introduce bias if the data are not missing completely at random. If patients with more severe symptoms are more likely to drop out of a study, your remaining continuous measurements will be systematically higher than the true population value. Multiple imputation is the standard workaround, and modern software makes it reasonably straightforward. But it is not a magic fix — the imputation model needs to include all variables related to the missingness mechanism, and if you do not have enough information to build that model, you are better off acknowledging the limitation than pretending the analysis is valid. Power analysis for continuous data is another area where assumptions matter more than people realize. The standard formula depends on the effect size, sample size, alpha level, and standard deviation. But that standard deviation comes from your data, and if your pilot data come from a highly selected population, your power estimate will be optimistic for the broader population you eventually study. I always add a buffer of at least 20 percent to whatever sample size a power calculation suggests, and I recalculate once I know something about the actual variance in the target population. The bottom line is that continuous data in biostatistics is not just about picking the right test. It is about understanding what your data are telling you before you commit to a procedure. Most errors in applied biostatistics come from skipping that step and treating the statistical method as something that exists independently of the data it is being applied to. The numbers do not care about your hypothesis. They care about their own structure. Work with that structure, not against it, and the analysis usually becomes straightforward.