Working With Ancient DNA From The Americas — What The Papers Actually Say

The field shifted fast around 2015 when Patterson and others published the first broad ancient DNA analyses covering pre-Columbian populations across North and South America. Before that, you were mostly reading linguistic reconstructions, skeletal measurements, and occasionally a tiny mtDNA study with twelve samples. After that, every major journal had a paper on the peopling of the Americas. I started digging into this properly because I needed population-level data for a project on Admixture timing in Mestizo communities, and the published figures kept contradicting each other by two to three hundred years depending on which lab processed the sample. The core dataset you'll run into comes from the Reich lab, the Max Planck Institute, and several South American groups at USP in São Paulo and CONICET in Argentina. They published genome-wide data from roughly forty ancient individuals spanning from around 9,000 years before present to contact-era remains. The key finding that matters for anyone actually using this data is that there was not one clean migration event. There was a basal Northern Native American lineage that split from East Asian populations roughly 20 to 25 thousand years ago, then a later divergence that produced Southern Native American clusters, and after that a series of regional bottlenecks and drift events that are still being sorted out. The Southern Brazilian and Peruvian ancient samples cluster differently from the Great Basin and Pacific Northwest ones, and that split is real, not an artifact of modern admixture.

A Genetic History Of The Americas

If you are searching for a single reference that covers the broad sweep, the 2018 paper by Fu and others in Nature, plus the later work from Lampert and the Reich group in Science around 2021, will get you most of the way there. There is no definitive book-length synthesis yet because the field is still overturning its own conclusions every eighteen months. What exists right now is a collection of supplementary tables, fastq files on SRA, and a few interactive portals. The SRA does have the raw reads. The 1000 Genomes Project has modern population coverage but almost none of it is from Indigenous American groups, which is a problem if you are trying to distinguish ancient structure from colonial-era mixing. The HGDP and the Simons Genome Diversity Project fill part of that gap, but their sampling of the Amazon and the Andes is thin. Here is the practical issue I ran into that took me three weeks to work around. I downloaded the ancient genome calls from a published supplement and tried to run qpAdm to model a particular Andean sample against a set of reference sources. The model came back with a p-value below 0.05 every time, which meant the assumed source populations were wrong, but the literature said they were right. The problem turned out to be post-mortem contamination and damage patterns that the authors had partially corrected but not fully. The ancient DNA community uses a tool called schmutzi to estimate contamination, but many of the older papers I was pulling data from only reported mitochondrial contamination estimates, not nuclear. I had to re-run the damage profiles through mapDamage2, filter for cytosine-to-thymine misincorporations at read ends, and then discard any sites where the damage pattern suggested exogenous DNA. That added about a day of processing per sample and removed roughly fifteen percent of my call sites, but the qpAdm models started working after that. The takeaway is that if you are building your own analysis on top of published ancient genomes, check the supplement for the actual damage metrics, not just the summary numbers in the main text. There is also a structural limitation that most people new to this area miss. The term "Ancient DNA" in the Americas context mostly refers to samples preserved in cold or dry environments. Cave sites in Mexico and the American Southwest give you decent coverage. Peruvian coastal sites with arid conditions work. But the tropical regions that cover most of the continent — the Amazon basin, the Chocó, much of Central America — yield almost nothing reliable past about two thousand years before present because the DNA degrades too fast in heat and humidity. This means the genetic history we currently have is skewed toward higher latitudes and arid zones. You cannot currently make strong claims about pre-Columbian population structure in the Amazon because there are almost no ancient genomes from there. Modern population studies infer structure there using unsupervised clustering on contemporary samples, but that conflates ancient patterns with centuries of post-1492 mixing. If someone tells you they have a clear picture of Amazonian pre-contact genetics, they are likely working from modern data and modeling assumptions, not direct ancient evidence.

Another thing that comes up when you actually try to use this material: the reference branch naming conventions are not stable. One paper calls a cluster "Northern Native American," the next calls it "Ancient Beringian," and a third calls it something else entirely. These are not always the same thing. Ancient Beringian specifically refers to the ANA-1 lineage found in the Anzic-1 child from Alaska, which sits basal to both Northern and Southern Native American branches. But some studies use "Ancient Beringian" more loosely to mean any pre-Beringian-strait population. I learned this the hard way when I tried to merge dataset labels across two published studies and ended up comparing apples to oranges in a PCA. My workaround was to look at the F3 statistics and the outgroup f-stats rather than trusting the labels, because the numeric values don't lie even when the naming does. For people who want to dig into the actual data without building a full pipeline from scratch, there are a few portals worth knowing about. The DDBJ/ENA/SRA trio holds the raw sequencing reads. Thereich lab posts some processed genotype files on their website, though they change the URLs often. The Human Origins dataset from the Reich group is downloadable if you request access through their form, and it includes many of the ancient American samples alongside the modern global panels. The Allen Ancient DNA Resource, or aadDNA, is another aggregation point that indexes a lot of the published datasets with consistent naming. I use that one most often because it saves me from chasing down individual SRA accession numbers. The computational side runs on standard population genetics tooling. ADMIXTOOLS 2 is the current version and it handles qpGraph, qpAdm, and the newer D-statistics better than the old ADMIXTOOLS 1. For quality control, pamDamage or mapDamage2 for ancient damage, schmutzi for contamination, andANGSD if you are working with low-coverage data where you cannot call genotypes confidently. I usually stick to ANGSD for anything under fivex coverage because hard genotype calls introduce bias in those ranges. For visualization, PCA is done with EIGENSOFT, and TreeMix or qpGraph handles the admixture graph fitting.

Get the Full Details

Pdf⚡️(read ️online) Origin: A Genetic History of the Americas
Pdf⚡️(read ️online) Origin: A Genetic History of the Americas

If you are approaching this from a genealogy or direct-to-consumer angle, the commercial testing companies do not currently offer much useful ancient American breakdown. 23andMe and Ancestry give you broad continental categories and some Indigenous North American signals, but those are modeled against modern reference panels and they conflate pre-Columbian and post-contact structure. If you have Indigenous ancestry and want something closer to the ancient signal, the only real path is to participate in a research study thatsequences your genome and then compares you against published ancient samples, or to wait for the field to mature enough that commercial companies incorporate better reference panels. Neither happens quickly. The state of the field right now is that we have good coverage for the northern parts of the continent and the Andean corridor, rough coverage for Mesoamerica, and almost nothing for the tropics. The dating is improving but still has error bars of plus or minus a few hundred years on individual samples. The population splits are broadly resolved but the finer structure, especially around the Southern Beringian standstill hypothesis and the possible later coastal migrations, is still contested. The data is real and it is growing, but it is not complete and it will not be complete for a while because preservation bias is a hard constraint, not something you can solve with better lab techniques.