Population genetics isn't what most people think it is
You pick up a textbook and it opens with Hardy-Weinberg equilibrium like it's some sacred law of nature. It's not. It's a null model, and knowing when to use it and when it's actively misleading is what separates people who can actually run an analysis from people who can just recite equations. I've spent more years than I care to admit watching researchers apply H-W equilibrium to datasets where it was completely inappropriate and then wondering why their results made no sense. The basic idea is straightforward enough. In an idealized population, allele and genotype frequencies stay constant across generations if there's no selection, no drift, no migration, no mutation, and random mating. The equation p^2 + 2pq + q^2 = 1 is just the binomial expansion applied to allele frequencies. Most introductions treat this as the foundation of everything, but in practice you'll rarely encounter a real population that even approximately satisfies those assumptions. That's fine. The model is still useful as a baseline. What matters is understanding what deviations from it actually tell you.
A Primer Of Population Genetics: What Actually Comes Up in Practice
I remember working with a cohort dataset a few years back where we were doing basic genotype quality checks. Something looked off with the heterozygosity call rate on a particular SNP chip batch, but the deviation wasn't random. It was concentrated in a specific geographic subgroup, and the pattern of excess homozygosity suggested inbreeding rather than genotyping error. The automated QC pipeline flagged it as a batch effect and wanted to drop the whole variant. Instead of accepting that, I recalculated the inbreeding coefficient F using the method-of-moments approach and compared it against the genome-wide distribution. The variant in question had an F value that was three standard deviations above the mean for that population, which pointed toward recent consanguinity in that subgroup, not a technical artifact. Dropping the variant would have thrown away a real biological signal. We ended up running a separate analysis stratified by ancestry and the signal held up. This kind of thing comes up constantly. The difference between a real population structure signal and a technical artifact is rarely obvious from the summary statistics alone. You have to understand what each metric is actually measuring before you can make that call. Let me talk about F-statistics, because this is where most people get tripped up. Wright's F-statistics measure deviations from Hardy-Weinberg expectations at different hierarchical levels. F_IS measures inbreeding within individuals relative to their subpopulation. F_ST measures differentiation between subpopulations relative to the total. F_IT measures inbreeding of individuals relative to the total population. They're not independent of each other. The relationship is (1 - F_IT) = (1 - F_IS)(1 - F_ST). Beginners often treat F_ST as if it's a fixed property of a pair of populations, but it depends heavily on the marker set you use and the Allele frequency spectrum you're sampling from.
Here's a counter-intuitive point that doesn't get emphasized enough: F_ST can be high even when the absolute difference in allele frequencies between populations is small, provided the alleles are at low frequency. Conversely, F_ST can be deceptively low when you're looking at common variants that have drifted to similar frequencies by chance. This is why some people prefer measures like Jost's D or the G' statistic when you're dealing with highly polymorphic markers like microsatellites or SNP arrays with dense coverage. Each has assumptions that matter. Using the wrong one for your data type will give you numbers that look reasonable but don't actually mean what you think they mean. Genetic drift is another concept that's taught as this clean mathematical process, but in real data it interacts with everything else in messy ways. The effective population size N_e is not the same as the census population size, and it's rarely constant over time. A population that went through a severe bottleneck ten thousand years ago can still show signatures of that event in contemporary genetic data, while a population that's been declining steadily for the last few generations might not show strong drift signals yet because the allele frequency changes haven't had time to accumulate. The coalescent framework handles this, but you need to think about time scales. When you're actually running analyses, the practical workflow usually goes something like this. You start with raw genotype data, which means dealing with call rate filters, minor allele frequency thresholds, and Hardy-Weinberg equilibrium p-value checks. You're not checking H-W to confirm the population is in equilibrium. You're checking it to flag potential genotyping errors, population stratification, or selection at a locus. A SNP that deviates from H-W in your control group might just be poorly clustered on the array, or it might be under balancing selection, or there might be a hidden sample contamination issue. The p-value alone won't tell you which.
Get the Full Details

After basic QC, you'll typically want to assess population structure. Principal component analysis is the standard approach, and it works well most of the time, but there are edge cases. If your sampling scheme is uneven, with one population having a hundred samples and another having ten, the PCA axes can be dominated by the larger group and the structure in the smaller group gets compressed. I've seen this enough that I now always check sample sizes per group before interpreting PCA plots at face value. ADMIXTURE or STRUCTURE-style analyses have similar issues with unbalanced sampling and with choosing the right K value. The cross-validation error curve from ADMIXTURE can be flat, meaning there's no clear optimal K, and picking one is mostly a judgment call at that point. Linkage disequilibrium is where things get computationally expensive and conceptually tricky. LD decays with distance due to recombination, but the rate of decay depends on recombination rate, population history, and selection. When you're doing genome-wide association studies, LD is both your friend and your problem. It's what allows you to use tag SNPs instead of sequencing everything, but it's also what makes fine-mapping difficult because you can't easily distinguish the causal variant from variants in LD with it. The r^2 metric tells you about correlation between alleles at two loci, while D' tells you about the history of recombination between them. They often disagree, especially with rare variants, and using the wrong one for your purpose leads to wrong conclusions about whether two variants are in LD. Selection signals are notoriously hard to detect reliably. The classic approach looks for extended haplotype homozygosity or elevated population differentiation at specific loci. But demographic events like bottlenecks and expansions produce genome-wide patterns that can mimic selection signals. I worked on a project where we found a very strong selection candidate in a population that had experienced a well-documented bottleneck. The signal looked compelling in isolation, but when we compared it against simulations under the inferred demographic model, the observed pattern was entirely consistent with drift alone. We dropped it from the candidate list. This is why simulation-based null models are essential, not optional. You can't just look at a site and say "that's selection" without accounting for the demographic history.
For people getting into this, I'd recommend starting with practical software before diving too deep into the theory. PLINK is the workhorse for basic QC and association testing. VCFtools handles variant filtering well. For population structure, admixture and PCA tools are standard. For coalescent simulations, msprime is fast and well-maintained. For actual analysis work beyond basics, R packages like dartR and hierfstat cover a lot of ground, though they have quirks you'll learn to work around. One thing that genuinely frustrates me is how many people treat these tools as black boxes. You can run a PCA in three commands and have a pretty plot in twenty minutes, but if you don't understand what each axis represents and what artifacts might be present, that plot is just decoration. Similarly, running F_ST calculations across thousands of SNPs and picking the highest values as "selection candidates" without correcting for demographic structure is one of the most common mistakes I see in the literature. The top F_ST outliers are often just regions of low recombination or areas near centromeres where drift has stronger effects. The field moves fast. New methods for detecting selection, inferring demography, and handling ancient DNA keep coming out. But the fundamentals haven't changed much in thirty years. Understanding allele frequency distributions, knowing what your statistics are actually measuring, and being honest about the limitations of your data will take you further than chasing the latest software release.
If you want a solid entry point that covers the mathematical foundations without getting lost in abstractions, Hill and Weir's work on F-statistics remains the reference I come back to. For practical applications, McVean's lectures on population genetics are freely available online and worth the time. The field has plenty of open problems, and the gap between textbook examples and real data is where most of the interesting work happens.
