Field Epidemiology Is Mostly Data Cleanup Before It Is Analysis

I have been working in outbreak investigation for more than fifteen years, and if there is one thing that becomes obvious fast, it is that the textbook version of epidemiology looks nothing like what actually happens when a cluster surfaces at 2 AM. The Principles Of Epidemiology In Public Health Practice framework gives you the architecture, but the work lives in the gaps between the boxes. The CDC curriculum is divided into three courses, and each one maps to a phase of public health work. Course 1 handles foundational concepts: disease frequency, measures of association, the epidemiologic triangle, and study designs. Course 2 moves into surveillance systems and outbreak investigation steps. Course 3 is biostatistics applied to public health, covering confidence intervals, hypothesis testing, and regression basics. The material is accurate. It is also deliberately simplified because the audience includes nurses, health department staff, and graduate students who need working knowledge, not a thesis. Here is what the course does not tell you, probably by design: real surveillance data is rarely clean enough to plug directly into any formula. I ran a norovirus investigation out of a suburban county health department once where the line listing had forty-seven cases and twelve of them had conflicting exposure dates. The date of symptom onset was listed as November 3, the date of lab confirmation was November 17, and the date of the suspected party was November 1. Two people had the same symptom onset date written as both 11/03/2023 and 03/11/2023. You cannot run a proper attack rate calculation on that without making decisions first, and those decisions are where the actual epidemiology lives.

Study Design Decisions That Actually Matter

Beginners tend to treat study design selection as a classification problem. Pick the label that fits. In practice it is a trade-off problem under uncertainty. A case-control study is faster and cheaper but introduces recall bias when exposure history depends on memory. A cohort study gives you incidence and relative risk directly but requires a defined population and time, which public health departments rarely have when an outbreak hits. Cross-sectional surveys are easy to administer and good for prevalence estimates, but they cannot establish temporality, which means they are almost useless for outbreak investigation unless you are mapping the damage after the fact. The trade-off I learned to stop ignoring is that every design choice commits you to a specific type of error. Case-control studies with poorly matched controls produce odds ratios that look precise but are wrong in the direction that matters least. I once saw a foodborne investigation where the control group was selected from hospital outpatients, which meant they were systematically different from the source population in ways that biased the exposure odds ratio toward one pathogen and away from another. The final conclusion was overturned when the data was reanalyzed with community-based controls recruited through random digit dialing. It added three weeks to the timeline and doubled the cost, but it was the only way to salvage the findings.

Measures of Association Beyond Relative Risk and Odds Ratios

The curriculum covers relative risk and odds ratios thoroughly. What it glosses over is attributable risk and population attributable fraction, which are the measures that actually drive resource allocation. A high relative risk does not automatically mean a high public health impact. If an exposure is rare, even a large relative risk may account for very few cases in the population. I remember reviewing a study where a novel occupational exposure showed a relative risk of 4.2 for a respiratory condition, but because fewer than one percent of the workforce was exposed, the population attributable fraction was under two percent. The right response was targeted workplace screening, not a public awareness campaign. Confusing the two leads to misallocated funding every single time. Attributable risk among the exposed tells you how much disease burden would disappear if you removed the exposure from the exposed group. Population attributable risk tells you the same thing for the entire population. These numbers matter when you are sitting across from a health officer who has to decide whether to issue a boil-water advisory or a shelter-in-place order. Relative risk is academic. Attributable risk is operational.

Get the Full Details

Principles of Epidemiology in Public Health Practice, 3rd Edition | PHF
Principles of Epidemiology in Public Health Practice, 3rd Edition | PHF

Surveillance Systems and the Problem of Completeness

Passive surveillance is the default in almost every jurisdiction because active surveillance is expensive and labor-intensive. The consequence is that every case count you see in a dashboard is an underestimate, and the degree of underestimation varies by disease, by demographic group, and by geographic area. I worked a measles investigation where the passive surveillance system flagged fourteen cases over six weeks. Retrospective chart review and provider outreach found thirty-one total cases. The difference was not random. Undocumented households and uninsured patients were underrepresented in the passive system by an estimated forty percent. Attack rate calculations based on the passive data alone would have understated the true spread by a factor closer to two than one. This is not a flaw in the epidemiologic method. It is a structural feature of public health infrastructure. The workaround is to treat every surveillance-derived estimate as a lower bound and adjust your interpretation accordingly. When you present findings to decision-makers, state the completeness limitation explicitly rather than burying it in a methods section where nobody reads it.

Outbreak Investigation Steps in Practice

The standard sequence runs like this: prepare for field work, confirm the outbreak exists, define and identify cases, describe data by time place and person, develop hypotheses, evaluate hypotheses analytically, conduct additional studies if needed, implement control and prevention measures, and communicate findings. Each step has decision points that the textbook presents as linear but are actually iterative and often contradictory. Confirming an outbreak sounds straightforward. You compare observed case counts to expected baseline and check whether the excess exceeds a predefined threshold. The threshold is usually a two standard deviation bump above the moving average for that week. But thresholds vary by disease and by season. Influenza has a predictable seasonal curve. Rotavirus does not. Norovirus spikes in winter but also appears in summer at summer camps. I learned to stop relying on automated alert thresholds entirely and to build a manual review step where an epidemiologist looks at the raw numbers before escalating. Automated alerts produced more false positives than true signals in my experience, which desensitized the team and delayed response to real events. Case definitions are where most investigations either succeed or fail early. A loose case definition captures too many people who do not have the disease and dilutes the association. A strict case definition misses cases and reduces statistical power. The balance shifts depending on the phase of the investigation. During the early descriptive phase you want sensitivity. During the analytical phase you want specificity. I have seen investigators stick with a sensitive case definition throughout an entire outbreak because changing it felt arbitrary, and the resulting odds ratios were pulled toward the null because the outcome group contained many non-cases. The fix was to run the analysis twice, once with the broad definition and once with the narrow one, and report both.

Bias and Confounding Are Not Abstractions

Selection bias occurs when the probability of being included in the study differs between exposed and unexposed groups in a way that distorts the measure of association. I encountered this in a point-prevalence survey of antibiotic resistance in a long-term care facility. The sample was drawn from residents who had culture results on file, which meant sicker residents were overrepresented. The resistance rate in the sample was 38 percent. The true resident population rate was likely closer to 22 percent. The bias came from the sampling frame, not from the lab methods. The fix was to weight the results by length of stay and acuity level, which brought the estimate much closer to the confirmed prevalence from a subsequent active survey. Confounding is more common and more dangerous because it is harder to detect. A confounder is associated with both the exposure and the outcome but is not on the causal pathway. Age is the classic example in nearly every public health analysis. I worked a cluster of Legionnaires disease where the initial unadjusted analysis suggested a strong association with hotel water consumption. The confounder was underlying pulmonary disease. Hotel guests in that particular building were older on average than the comparison group, and older adults with COPD or emphysema were both more likely to stay in that hotel and more likely to develop severe legionellosis. After age and comorbidity adjustment, the association weakened substantially. The final conclusion still pointed to the hotel water system, but the magnitude of risk changed from alarming to moderate, which affected the scope of the remediation order.

Principles of Epidemiology in Public Health Practice 3rd Edition - Principles of Epidemiology ...
Principles of Epidemiology in Public Health Practice 3rd Edition - Principles of Epidemiology ...

When Epidemiologic Methods Fail and What to Do Instead

Epidemiology assumes that patterns in populations reflect underlying causal structures. This assumption breaks down in small populations, in highly mobile populations, and in situations where exposure measurement is impossible. I have seen investigators try to apply standard outbreak methods to a population of fifty people in a remote community and get results so unstable that the confidence intervals spanned the entire range from protection to harm. In those cases the method is not wrong. The sample is just too small to support inference. The honest answer is to report descriptive findings, acknowledge the uncertainty, and recommend enhanced surveillance rather than forced analysis. Another failure mode is when the exposure is common and the outcome is rare. Case-control studies handle this fine. Cohort studies become impractical because you would need tens of thousands of person-years to observe enough outcomes. In those situations the alternative is a case-cohort design or a nested case-control study within an existing cohort. These designs are underutilized in public health practice because they require pre-existing cohort infrastructure, which most local health departments do not have. The workaround is to partner with academic medical centers or state registries that already maintain longitudinal data.

Communication Is Part of the Method

Epidemiologic findings are useless if the audience cannot act on them. I have watched carefully conducted investigations get ignored because the final report was written for epidemiologists instead of for the people making decisions. Health officers need to know what the numbers mean for their budget and their timeline, not what the p-value indicates about the null hypothesis. I learned to lead with the bottom line: how many cases are likely, what the probable source is, what intervention would reduce cases fastest, and what the uncertainty is. The methods section comes after, and it is kept short enough that someone reading on a phone between meetings can still follow it. Uncertainty communication is the hardest part of this job. People want certainty. The data rarely provides it. The most effective approach I have found is to state the confidence interval alongside the point estimate and then translate it into plain language. "The relative risk is 3.1 with a 95 percent confidence interval from 1.4 to 6.8, which means the true effect could be as low as a doubling of risk or as high as a sixfold increase." That sentence carries more information than any narrative summary and forces the reader to confront the range of possible outcomes rather than latching onto a single number.

Practical Resources Beyond the CDC Curriculum

The CDC course is free and available through the CDC website. It is the starting point, not the endpoint. For study design fundamentals, Rothman Greenland Lash remains the reference most practicing epidemiologists keep on their desks even though it is dense. For applied outbreak work, the WHO Field Epidemiology Manual and the APHL practical guides are more relevant because they are written for the context you actually work in. The key message from all of these sources is the same one I have learned the hard way: the principles are stable, the practice is messy, and the gap between them is where the work happens.

Principles of Epidemiology in Public Health Practice Third Edition | ScholarFriends
Principles of Epidemiology in Public Health Practice Third Edition | ScholarFriends