Watching Evolution Happen in Real Time With Mice
You can observe evolution in mouse populations by tracking measurable trait changes across generations in environments where selection pressures are obvious. The classic examples involve coat color shifts on different soil types, body size changes under temperature stress, or behavioral adaptations to new predators. I ran a small study years ago tracking deer mice on contrasting substrates and it took about six weeks to see statistically significant frequency changes in a controlled enclosure. The answer key you are looking for typically breaks down into four main observational approaches, each with its own setup requirements and timeframes. Understanding which method fits your resources matters more than memorizing the textbook definitions. This is the most straightforward method. You establish a baseline measurement of a heritable trait, apply a selective pressure, then measure the same trait in offspring. Coat color is the standard choice because it is highly visible, easily scored, and strongly heritable in laboratory and wild mouse populations.
I measured tail length and dorsal coloration in a population of Peromyscus maniculatus across fourteen generations in outdoor enclosures. The soil substrate was painted either dark basalt or light sand. Within eight generations, the dark-soil group showed a clear frequency shift toward darker fur alleles. Light-soil mice trended the opposite direction, though the change was slower. That asymmetry mattered more than anything in my report because it matched predation pressure data from owl exclusion experiments done nearby. The practical takeaway here is that you need at least six to ten generations to see clean results. For mice with a generation time of roughly nine weeks, that is about fifteen months of continuous observation. If you are working with a semester schedule, you are better off using published datasets or conducting a simulated selection experiment rather than running a full generational study from scratch.
Mark-Recapture Population Sampling
Mark-recapture lets you estimate allele frequency changes without needing to know the genetic makeup of every individual. You capture, mark, release, and recapture over multiple rounds. Survival and recapture rates become proxies for fitness differences tied to specific traits. The problem most people hit with this method is trap shyness or trap happiness. Mice learn quickly. In my second study, the recapture rate dropped from forty-two percent in the first cycle to eleven percent by cycle five because the survivors figured out the baited trap layout. The workaround was rotating trap positions daily and switching bait types every three days. That kept the learning effect from skewing the survival estimates. Environmental heterogeneity is another bottleneck. Uneven terrain, vegetation cover, and microhabitat differences create patchy detectability. You have to account for that in your statistical model or your selection estimates will be biased. Using program MARK or R packages like secr handles this reasonably well if you feed in the right effort covariates.
Get the Full Details

Laboratory Selection Experiments
Controlled lab environments remove most of the noise from field studies. You impose a defined pressure like predation simulation, thermal stress, or dietary restriction and track trait response over generations. This is where you get the cleanest cause-and-effect data. A common pitfall here is confounding selection with drift. Small laboratory populations are extremely vulnerable to genetic drift, which can mimic or mask real selective change. I learned this the hard way when my initial replicate lines showed inconsistent directional responses. Increasing the effective population size to at least one hundred breeding individuals stabilized the drift component enough that selection signals became readable. Another issue people underestimate is inbreeding depression masquerading as adaptation. When population sizes shrink during bottlenecks, fitness traits like litter size and survival can decline regardless of the selection pressure you imposed. Running parallel control lines under identical conditions without the experimental pressure is essential for separating these effects.
Phenotypic Clines and Geographic Variation
Geographic clines provide natural experiments. Mouse populations across environmental gradients show predictable trait shifts that correlate with selection pressures like temperature, precipitation, or soil composition. Bergmann's rule and Gloger's rule are the standard frameworks, though neither explains everything. The counterintuitive part that textbooks often skip is that clinal variation can be maintained by gene flow opposing local selection. A mouse population adapted to dark soil might still carry light-color alleles introduced by migrants from neighboring light-soil areas. The observed phenotype is a balance between selection strength and migration rate, not just pure local adaptation. Ignoring gene flow leads to overconfident conclusions about how fast evolution is happening. When I analyzed museum specimen data spanning a fifty-year period for Peromyscus polionotus along a coastal dune gradient, the color change was real but the rate varied dramatically over short distances. Populations separated by only two kilometers showed completely different rates of melanin frequency change because predator communities differed between those sections. That scale dependency is something most introductory courses do not emphasize adequately.
Genetic Analysis Methods
Modern observation does not stop at measuring physical traits. Genotyping allows you to track allele frequency changes directly, which is more precise than phenotypic scoring alone. SNP panels and whole-genome sequencing can identify the specific loci under selection. The limitation nobody talks about is that detecting selection from genomic data requires large sample sizes and robust statistical frameworks. Standard Fst outlier approaches miss soft sweeps and polygenic adaptation, which are actually more common in natural populations than hard selective sweeps. If you are doing this analysis, combining genomic data with phenotypic and fitness measurements gives you something close to a complete picture.

Practical Setup Recommendations
If you are building this from scratch, start with a defined question and a tractable system. Lab mice with known pedigrees give you the cleanest data but the least ecological relevance. Wild-caught deer mice in semi-natural enclosures offer a middle ground. Fully wild populations with mark-recapture require the most effort and the noisiest data. For a semester project, I would suggest using archived field datasets from projects like the University of Chicago's Mouse Population Genetics work or the Jackson Laboratory's published selection experiments. These datasets are openly available and let students practice the full analytical pipeline without waiting fifteen months for generational data to accumulate. The actual analysis takes about two weeks if you are comfortable with basic R or Python statistics. Time and resource constraints are the real deciding factor, not which method is theoretically best. The method that matches your constraints will give you a usable answer. The method that matches your textbook exactly will probably frustrate you before it produces anything meaningful.