Polygenic characteristics and why they make prediction messy

A polygenic characteristic is a trait influenced by many genes across the genome rather than a single gene or a handful of genes. Height, body mass index, educational attainment, type 2 diabetes risk, schizophrenia liability — these are the usual suspects. Each variant contributes a small amount, often less than half a percent of variance, and the combined effect is what we model as a polygenic risk score or a polygenic index. That is the basic definition. The reality is messier. When people ask what are polygenic characteristics, they usually want a clean answer. The clean answer is a clean lie. These traits do not behave like Mendelian diseases where one gene flips a switch. They behave like weather: lots of interacting factors, weak individual signals, and distributions that shift rather than snap.

What Are Polygenic Characteristics in Practice

I spent three years working with GWAS summary statistics and polygenic scoring for a insurance risk modeling project. We were building predictive scores for metabolic disease incidence. The first mistake everyone makes is treating the polygenic risk score like a diagnosis. It is not. It is a ranked estimate with confidence intervals that widen dramatically at the tails where you actually care about them. The pipeline looks straightforward on paper. You take a discovery GWAS, extract effect sizes for variants in your target cohort, weigh each genotype by its effect size, and sum across variants. That gives you a raw polygenic score. You standardize it. You test its predictive power. You publish an R-squared value and move on. Usually around step four you realize something is wrong. The edge case that broke our model was population stratification that looked clean on the surface. Our discovery dataset was mostly Northern European. Our target validation set had a visible PCA cluster that still fell within what most people would call "European." We ran the standard PCA correction with ten principal components. The score looked fine. Then we tested it against a geographic variable and found that the score tracked north-to-south gradients within Europe at roughly the same magnitude as the trait itself. We were not measuring genetics. We were measuring ancestry in a way that correlated with environment, diet, healthcare access. The score inflated heritability estimates by about thirty percent once we accounted for this gradient properly. We added more PCs, ran LD score regression to get the intercept, and applied LDpred2 with a population-matched reference panel. The final cross-validated AUC dropped from 0.71 to 0.63. Still useful. Just not as useful as the first run suggested.

The mechanics behind polygenic prediction

Here is what actually matters when you are working with these characteristics and nobody tells you. Linkage disequilibrium is the first thing that gets glossed over. Variants are not independent. They travel in blocks because recombination does not randomly shuffle the genome every generation. When a GWAS reports a variant as associated with a trait, that variant is often just tagging a causal variant somewhere nearby in the same haplotype block. If you build a polygenic score naively from GWAS hits, you are double counting correlated signals. Methods like LDpred2, PRS-CS, and SBayesR exist to approximate the true underlying effect sizes by modeling the LD structure. Using a naive clumping and thresholding approach with PRSice can cut your predictive accuracy in half compared to a Bayesian shrinkage method on the right dataset. That is not theoretical. That happened to me with a cardiovascular disease score built from UK Biobank summary statistics applied to an independent cohort. Predictive portability across ancestries is the second thing that gets glossed over. A polygenic score derived from a European GWAS typically loses between forty and eighty percent of its predictive power when applied to African ancestry populations. The reasons are structural: different LD patterns, different allele frequencies, different causal variant sets, and the historical underrepresentation of non-European cohorts in GWAS. This is not a fixable bug with better software. It is a data problem. If you are building scores for clinical use and your model only generalizes to one ancestry group, you are building a tool that works for some people and fails silently for others. The failure is silent because the score still produces a number. It just produces the wrong number with the wrong uncertainty.

Get the Full Details

Polygenic Traits - Biology Simple
Polygenic Traits - Biology Simple

Heterogeneity of effect sizes across environments is the third thing. Gene-environment interaction is real and it matters more than most papers admit. A polygenic score for educational attainment predicts differently in high-poverty versus low-poverty environments because the environmental ceiling or floor constrains how much genetic potential can express itself. The same genotype does not mean the same phenotype across different socioeconomic contexts. This is sometimes called gene-environment correlation, sometimes GxE, and most often it is ignored because adjusting for it makes the model less clean and the story less publishable.

When polygenic scores actually work

They work best for traits with high SNP heritability estimated from methods like GREML or LD score regression. Height, BMI, schizophrenia, coronary artery disease — these have solid SNP heritability estimates and reasonably large GWAS discovery samples. For these traits, polygenic scores can explain anywhere from five to twenty-five percent of phenotypic variance depending on the trait, the sample size of the discovery GWAS, and how well matched the target population is. They work worse for traits that are heavily influenced by rare variants, structural variation, or epigenetic regulation. Complex traits with significant de novo mutation contributions or strong parental imprinting effects do not yield clean polygenic architectures. The infinite sites model that most polygenic scoring methods implicitly assume breaks down when a meaningful fraction of causal variation comes from rare structural events that are poorly tagged by common SNPs on genotyping arrays. They work even worse for binary traits when applied to populations with substantially different base rates. A schizophrenia polygenic score calibrated on a population with two percent prevalence will produce systematically biased absolute risk estimates if applied to a population where the prevalence is one percent. The relative ranking might still be valid. The absolute risk translation requires careful calibration that most clinical implementations skip entirely.

A practical workflow that does not waste your time

Start with a high-quality GWAS summary dataset. Check the effective sample size, the lambda GC, and the LD score regression intercept. If the intercept is above 1.1 for a traits like height where true polygenicity should not inflate it that much, you have residual stratification and your effect sizes are contaminated. Do not proceed until you fix that. Use a reference panel that matches your target population for LD modeling. 1000 Genomes phase 3 is free but it is not ideal. The TOPMed panel or a population-specific imputed panel gives you better LD estimates and meaningfully improves score accuracy. One paper showed a seven percent increase in R-squared for height scores when switching from 1000 Genomes to a population-matched reference. Seven percent sounds small until you are working at the tail of a distribution where every point of discrimination matters. Apply a Bayesian shrinks method rather than naive clumping and thresholding unless you have a computational constraint that makes Bayesian methods infeasible. LDpred2, PRS-CS, and SBayesR all outperform manual clumping on held-out data across most traits. The run time is longer — minutes to hours instead of seconds — but the accuracy gain is usually worth it.

Polygenic — Definition & Examples - Expii
Polygenic — Definition & Examples - Expii

Always validate in an independent sample from the same ancestry background before you ever think about cross-ancestry transfer. Report both R-squared on the liability scale for binary traits and AUC or C-statistic. Do not report one and call it proof. Report both. Report the calibration slope too. A score with good discrimination but terrible calibration is worse than useless in a clinical setting because it gives you the wrong confidence in the right direction. If you need to transfer across ancestries, use multi-ancestry GWAS summary statistics when available and methods like CT-SLEB or MR-MEGA that explicitly model ancestry-specific effects. Do not just run the European score on African samples and rationalize it away. The performance gap is not a nuance. It is a hard boundary imposed by population genetic structure.

The uncomfortable things nobody says about polygenic characteristics

Polygenic scores are increasingly being commercialized for direct-to-consumer health reports. The science behind many of those products is thin. They use outdated methods, small discovery samples, and zero ancestry matching. A person receives a "high genetic risk for Alzheimer's" score and has no idea that the underlying GWAS had only ten thousand cases, that the score explains four percent of variance in Europeans but two percent in their actual ancestry, and that the absolute risk conversion is essentially a guess. This is not a hypothetical concern. I reviewed several of these products for a consulting job and the gap between what they sell and what the metrics actually support was large enough that I recommended against using any of them for decision making. The other uncomfortable truth is that polygenic scores are already being used in ways that most researchers would consider ethically questionable. Employer screening, insurance adjustments, credit scoring — the technical infrastructure exists. The legal framework in most jurisdictions has not caught up. Just because you can build a score that predicts a health outcome does not mean you should deploy it in a context where the consequences of error are asymmetric. A false positive in a research setting means a data point is re-examined. A false positive in a hiring or lending context means a person loses opportunity. The statistical properties are identical. The human costs are not. There is also the quiet problem of winner's curse in early-stage GWAS. Effect sizes from discovery samples are systematically inflated, especially for variants near significance thresholds. If you build a polygenic score using uncorrected effect sizes from a modest sample, your score will look impressive in the discovery cohort and underperform in any validation set. The magnitude of this inflation shrinks as sample size grows, which is why the recent mega-GWAS with hundreds of thousands of participants have produced scores that actually generalize. Before those larger studies, a lot of published polygenic scores were overfitted by design.

Where the field is actually heading

Multi-trait methods that leverage genetic correlations between traits are starting to improve prediction for traits with modest GWAS sample sizes. Summarized-data-based MR uses genetic correlations to borrow strength across related phenotypes. This can boost predictive accuracy by five to twelve percent for traits like depression or ADHD where discovery samples are smaller than for height or schizophrenia. Fine-mapping integration is the next incremental step. Instead of treating every GWAS variant as equally likely causal, methods like SuSiE and FINEMAP identify credible sets of causal variants and weight polygenic scores by posterior inclusion probability rather than raw effect size. This is computationally expensive and requires individual-level genotype data or very dense summary statistics, but it reduces noise in the score and improves portability slightly. The biggest upcoming shift is the move toward whole-genome sequencing based polygenic scoring rather than imputed SNP arrays. Rare variants, structural variants, and non-additive effects are largely invisible to current approaches. As sequencing costs drop and methods for aggregating rare variant effects improve, polygenic scores will likely become more accurate, especially for traits where rare variants contribute meaningfully to risk. This is years away from routine clinical use but the first papers are already appearing.

Polygenic Chart
Polygenic Chart

For now, the honest answer about what polygenic characteristics are and what polygenic scores can do is narrower than the hype and wider than the criticism. They capture real signal. That signal is small, population-specific, environmentally contingent, and easily misused. The people who understand this tend to be the ones who build scores that actually work. The people who ignore these constraints tend to be the ones who publish results that do not replicate.