Why Most Public Health Programs Fail at Data Analysis

I keep seeing the same mistakes in grant proposals and outbreak reports. People treat statistics like something you bolt onto the end of a project instead of building the study around from the start. It does not work that way. I spent about nine years doing surveillance and evaluation work for county health departments before moving into consulting, and the projects that fell apart always had one thing in common: someone ran a t-test on data they never thought about collecting properly. Biostatistics in public health is not about crunching numbers after the fact. It is about deciding what question you can actually answer with the data you can realistically get, then structuring everything around that constraint. The essentials break down into three areas that everyone glosses over: study design alignment, proper handling of clustering, and knowing when your sample size is lying to you. The first one comes up constantly. You pick a design first. If you are looking at a vaccination campaign rollout across twelve rural clinics, a simple random sample does not exist. Those patients are nested inside clinics, and clinics are nested inside health districts. Ignoring that hierarchy gives you confidence intervals that are far too narrow. I had a colleague who published a paper on maternal care access in a specific region and got torn apart in peer review because the intracluster correlation coefficient was 0.18 and the author had run a standard regression treating every patient as independent. The p-values were basically fake. The fix was a mixed-effects model with clinic-level random intercepts. It changed the conclusions entirely.

What Actually Matters in Practice

Most introductory courses teach you the formulas. They do not teach you what happens when your data looks nothing like the textbook examples. Here are the things that will bite you if you are not paying attention. Survival analysis is one of the most underused tools in public health, and not because it is too hard. It is underused because people do not understand censoring. Right-censored data means you know a subject survived at least until a certain point, but you do not know when the event happened after that. Kaplan-Meier curves handle this fine. Cox proportional hazards models handle it too, but only if the proportionality assumption holds. I learned that the hard way during a cohort study tracking time to disease onset among workers exposed to a certain chemical. The hazard ratio shifted over time in a way the model could not capture. We ended up splitting the follow-up into two periods and running separate models for each. It was messy but accurate. A standard Cox model would have given a single hazard ratio that meant almost nothing. Another thing nobody warns you about: ecological fallacy. Aggregating data to the neighborhood or district level and then making claims about individuals is a cardinal sin. I saw a study that claimed a certain intervention reduced disease incidence based on district-level vaccination rates and hospital admission numbers. The correlation was strong. The inference was worthless. Individual-level exposure and outcome data are required for that kind of claim. Ecological studies can generate hypotheses. That is it.

Software and What Actually Works

R is the standard. Stata is still widely used in epidemiology. SAS hangs around in pharmaceutical and regulatory settings. For pure biostatistics work in public health, R with the survival, lme4, and survey packages covers most needs. The survey package is critical if you are dealing with complex survey designs, which is almost always the case in public health data. It handles weighting, stratification, and clustering in one pass. Without it, your standard errors are wrong. If you are working with small datasets or rare outcomes, penalized logistic regression (Firth correction) prevents separation issues that would otherwise make coefficients blow up to infinity. This comes up more often than you would expect in outbreak investigations where the number of cases is small. Power analysis is another area where people make careless mistakes. Post-hoc power calculations are essentially meaningless. If your study finds no significant result, calculating "observed power" tells you nothing useful. Do your power calculations before the study starts, using realistic effect sizes from prior literature, not from your own unpublished results. G*Power handles basic designs. For cluster-randomized trials, you need to account for the design effect, which inflates the required sample size by a factor of 1 plus the intracluster correlation times the average cluster size. That number can double or triple your recruitment target without you realizing it.

Get the Full Details

Essentials of Biostatistics in Public Health by Lisa Sullivan (2017 ...
Essentials of Biostatistics in Public Health by Lisa Sullivan (2017 ...

Common Pitfalls and Where Methods Break Down

Multiple testing is the most common issue in large-scale public health datasets. If you run fifty comparisons at alpha 0.05, you expect two or three false positives by chance alone. Bonferroni correction is conservative but simple. Benjamini-Hochberg controls the false discovery rate and is usually more appropriate for exploratory public health work. Pick one and state it clearly in your methods section. Not doing so is a red flag for reviewers. Missing data is rarely missing completely at random. If you are analyzing survey data and certain demographic groups systematically skip questions, listwise deletion biases your results. Multiple imputation is the standard approach. R's mice package handles this well. But imputation assumes the data is missing at random conditional on the variables you include. If there is an unmeasured reason for the missingness, your imputed values are still biased. There is no clean fix for that except acknowledging it as a limitation. Multilevel models are powerful but fragile. Convergence failures are common when you have few clusters or sparse data at higher levels. I once had a three-level model fail to converge because there were only six schools in the top level. Switching to a two-level model with school-level clustering was the practical solution, even though it meant losing some granularity. Reporting that limitation is better than reporting a model that never converged.

How to Actually Learn This

Read Applied Logistic Regression by Hosmer, Lemeshow, and Sturdivant. It is dense but covers the ground you need. For survival analysis, Kleinbaum and Klein's Survival Analysis is the reference. Online, the Johns Hopkins biostatistics courses on Coursera are solid for basics, but they skip the messy real-world details. You learn those from doing the work. Work through actual datasets. The NHANES data from the CDC is freely available and has complex survey weights built in. Running the survey package on it forces you to confront weighting and clustering immediately. CDC WONDER provides aggregate mortality and morbidity data that is useful for ecological analysis practice. The key is getting your hands on data that does not behave nicely, because that is what you will actually encounter. The fundamental skill is not running the right test. It is recognizing when the question you are asking cannot be answered by the data you have, and adjusting accordingly. That is the part that takes experience. Everything else is just mechanics.