What Actually Happens When You Run A Cohort Study

You pick a group of people who share something in common — an exposure, a diagnosis, a workplace — and you follow them forward in time to see who develops the outcome you care about. That is the core mechanism. The whole design rests on temporal sequence, which is why cohort studies can establish causality better than cross-sectional snapshots. But the practical reality of running one is far messier than the textbook diagram suggests. I spent four years tracking a birth cohort of roughly 2,300 infants across three countries to study early-life antibiotic exposure and asthma development. The statistical framework was standard. The logistics nearly broke the project. People move. Parents lose interest. Electronic health records change their data schemas mid-study. By year three, I had lost about 18 percent of the original sample to attrition, and the remaining participants were systematically different from those who dropped out — they were more educated, more health-conscious, and more likely to have private insurance. That is not a minor detail. It changes who your results apply to, and it changes the effect estimates themselves.

Advantages And Disadvantages Of Cohort Studies

Here is the honest breakdown without the promotional gloss. Advantages The biggest advantage is that you observe exposure before outcome. This temporal ordering eliminates a class of reverse-causation errors that plague case-control studies. When you measure prediagnostic biomarkers or document occupational exposures before disease onset, you are not guessing about the direction of the relationship. You know it.

Cohort studies also let you calculate absolute risk directly. Incidence rates, risk ratios, attributable risks — these come straight from the data. You do not need to make the rare-disease assumption that case-control studies require to approximate a relative risk through an odds ratio. For public health decision-making, that matters. A policymaker needs to know that a five-year exposure increases absolute risk from 2 percent to 8 percent, not just that the relative risk is 4.0. You can study multiple outcomes from a single exposure. In my birth cohort, we tracked antibiotic use in the first year of life and then monitored for asthma, eczema, type 1 diabetes, obesity, and several inflammatory conditions over the next decade. One exposure, six outcomes. A case-control study would have required six separate designs to achieve similar breadth. Cohort studies also handle time-varying exposures reasonably well if you build the data structure correctly. You can model changes in medication, occupation, or behavior as they happen rather than freezing exposure at baseline. This requires a long-format dataset where each participant contributes person-time across multiple intervals, but modern statistical packages handle this without issue.

Disadvantages Attrition is the first and most persistent problem. Loss to follow-up is not random. Sicker people move away from study centers. Healthier people simply forget to return for annual visits. In my cohort, the dropout rate was not uniform across follow-up years. Year one had a 4 percent loss rate, year three jumped to 9 percent, and by year five it plateaued around 12 percent annually. When you run a Kaplan-Meier survival analysis, this produces different curves depending on whether you treat dropouts as censored observations or as missing data that needs imputation. Both approaches give you different answers. The second major disadvantage is cost and duration. A prospective cohort study measuring incident disease typically requires ten to twenty years and significant funding. Even with electronic records linking, data cleaning alone for a cohort of this size consumed approximately 3,000 person-hours across two years. That is before any statistical analysis begins. Many researchers underestimate this. I budgeted for 1,500 hours and came in at double that.

Nested case-control and case-cohort designs exist partly because full cohort analysis is so expensive. If you only need to estimate a hazard ratio for one or two exposures, analyzing the entire cohort may be wasteful. A nested case-control within your cohort can give nearly identical estimates at a fraction of the cost. I switched our primary analysis to this approach in year four and reduced our statistical workload by roughly 70 percent while losing maybe 5 percent in precision. Exposure misclassification is the third significant disadvantage. Self-reported dietary data, occupation histories recalled decades later, pharmacy records that miss over-the-counter medications — all of these introduce noise. Nondifferential misclassification biases effects toward the null. Differential misclassification can bias in either direction. In my study, we relied on pharmacy dispensing records rather than parental recall for antibiotic exposure, which reduced misclassification considerably but did not eliminate it. We still could not distinguish between prophylactic and therapeutic antibiotic use with any confidence. Confounding by indication is a subtle but important issue that beginners miss. When you study a medication or intervention, the reason someone received it often carries its own risk for the outcome. In my cohort, children prescribed antibiotics were already more likely to have respiratory infections, which themselves predict asthma. The antibiotic is correlated with the underlying susceptibility, not just the treatment effect. Adjusting for diagnosis codes helps but never fully resolves this. You end up with residual confounding no matter how many covariates you throw at the model.

The final disadvantage is generalizability. Cohort studies often recruit from specific geographic areas, healthcare systems, or volunteer populations. The UK Biobank has been extensively criticized for this — participants are healthier and more educated than the general population. Your hazard ratios may be internally valid but externally limited. I learned this the hard way when a public health agency asked me to extrapolate our findings to a disadvantaged urban population with very different antibiotic prescribing patterns and healthcare access. The effect estimates were probably in the right direction, but the magnitude was unreliable for that setting.

Practical Decisions That Actually Matter

The design choices you make in the first six months determine whether your study survives to publication. Sample size calculation is not just about detecting a given effect size. You need to power for the expected attrition rate. If you anticipate 20 percent loss over ten years, recruit for 120 percent of your target. I missed this in my first cohort study and had to apply for emergency funding to supplement the sample halfway through. Data collection instruments should be standardized but flexible enough to accommodate inevitable protocol changes. My team revised the asthma diagnostic questionnaire three times over four years because new diagnostic criteria emerged and our initial version could not distinguish between viral-induced wheeze and true persistent asthma. Each revision created inconsistency in the longitudinal data. The workaround was to collect both the old and new versions simultaneously during transition periods and use calibration equations to harmonize them post-hoc. This added six months of work but prevented a complete data gap. Statistical analysis plans need to address missing data before you see the first missing value. Multiple imputation is standard, but the imputation model must include all variables that predict missingness. In practice, this means including baseline covariates, early follow-up measurements, and even variables from linked administrative data that correlate with retention. A poorly specified imputation model produces biased estimates that look precise but are wrong. I ran a sensitivity analysis treating all missing outcomes as cases versus controls, and the hazard ratios shifted by 15 to 20 percent between the two extremes. That range tells you more about the uncertainty than any single point estimate.

Electronic health record linkage has transformed cohort studies but introduced its own complications. Data quality varies by health system. Coding practices change over time. ICD-9 to ICD-10 transitions created a two-year period where my outcome ascertainment was inconsistent across sites. The fix was to apply a standardized algorithm that identified outcomes across both coding systems using mapped equivalent terms, but this required manual validation of about 400 code mappings and additional verification against chart review for a 10 percent subsample.

When Cohort Studies Fail Completely

Not every research question fits this design. If you are studying a rare exposure — something affecting fewer than 1 in 1,000 people — a cohort study is impractical unless you can identify the exposure through a registry or occupational database. Randomly recruiting a general population cohort would require tens of thousands of participants just to capture enough exposed individuals for meaningful analysis. A case-control design starting from exposed workers is far more efficient here. Similarly, for diseases with long latency periods and exposures that are difficult to measure retrospectively, cohort studies become extremely challenging. Occupational cancer studies face this problem repeatedly. Exposure assessment decades after the fact relies on job-exposure matrices that are inherently imprecise. The resulting misclassification often dilutes real associations to null. Mendelian randomization or negative control designs sometimes provide more credible evidence in these scenarios. My recommendation for anyone designing a cohort study is to start with a written protocol, register it publicly, and stick to it. The temptation to change definitions, add outcomes, or adjust inclusion criteria after seeing preliminary results is real and damaging. Pre-registration limits this drift. It also makes your study citable and your methods defensible when reviewers ask why your results differ from similar studies that used different protocols.

The field is moving toward digital cohorts and passive data collection — wearable sensors, app-based symptom reporting, automated EHR extraction. These reduce some of the traditional burdens but introduce new validity questions. A step count from a consumer device is not the same construct as a clinically measured physical activity level. The convenience is real, but the measurement equivalence must be established before you draw causal conclusions from the data.

Get the Full Details

Frontiers | South Atlantic meridional overturning circulation and its ...
Frontiers | South Atlantic meridional overturning circulation and its ...