Working with Deep Time Datasets in Paleoanthropology
The Last Two Million Years is a reference framework and accompanying dataset that researchers use when mapping hominin fossil distributions against paleoclimatic data. It compiles dated occurrences, stratigraphic columns, and environmental proxies spanning roughly 2 million years ago to the present. If you are new to this, the first thing you will notice is that the documentation assumes you already know how to handle radiometric date ranges and how to merge heterogeneous source datasets. The dataset is available for download from the supplementary materials section of the associated publication. The primary file is a CSV containing occurrence records with latitude, longitude, age range, formation, and bibliographic references. There is also a metadata file that explains the column naming conventions, which are not intuitive if you are used to working with GBIF exports. I spent about a week before realizing the age ranges were given in cal kyr BP rather than raw radiocarbon years. That mismatch cost me two days of debugging on my first attempt to cross-reference it with a Paleoclimate database query. The download itself is straightforward. Go to the project page linked in the publication, navigate to the Data Availability section, and grab the zip archive. Extract it to a dedicated working directory. The archive contains the main occurrence file, the metadata document, a README with installation notes for the optional R package, and a few Python scripts for common transformations. Nothing fancy. It works on Linux, macOS, and Windows, though the scripts assume you have Python 3.9 or later and a few standard packages installed.
What You Actually Need to Know Before Using It
The biggest gap in the documentation is the treatment of temporal uncertainty. Every fossil occurrence has an age range, but those ranges are not uniform. Some come from single radiometric dates with error bars. Others are stratigraphically interpolated. A few are entirely contextual, pulled from associated fauna rather than direct dating. The dataset flags these types, but the flag column is labeled type_code and the legend is buried in the metadata file rather than prominently displayed. If you run queries without accounting for type, your results will look clean and be completely unreliable. I learned this the hard way when I published a range estimate that turned out to be based on three contextual dates masquerading as direct dates. Another thing that trips people up is the geographic coordinate system. The latitudes and longitudes are in WGS84, but some of the older entries in the dataset have coordinates that are clearly misaligned with their reported formations. This happens because the original publications used different datums. You should always visually verify coordinates against a map before running spatial analyses. A quick check in QGIS takes thirty seconds and will save you from publishing locations that are hundreds of kilometers off.
Common Workflow
Here is how most people actually use this. Load the CSV into your analysis environment. Filter to the species or genus you care about. Handle the temporal uncertainty by either collapsing age ranges to midpoints for rough work or propagating the full ranges for anything publication-grade. Merge with whatever environmental reconstruction you are using, keeping in mind that the resolution of paleoclimate data varies significantly across time and space. African sites tend to have better coverage than sites in eastern Asia or the Americas for the earlier parts of the timeline. The optional R package provides functions for some of the common merges and visualizations. It is functional but not polished. The documentation for the more advanced functions is thin. I ended up reading the source code to figure out how the weighting worked when combining multiple date estimates for a single occurrence. The package also lacks support for certain date ranges that extend beyond its internal calendar assumptions. If your data includes events near the 2 million year boundary, you may need to adjust the calendar handling manually.
Get the Full Details

Where This Falls Apart
The dataset is strongest for East African Plio-Pleistocene sites and weakens considerably elsewhere. Sites from Europe, South Asia, and the Americas are sparser and have less consistent dating. If your research focus is outside of well-sampled regions, you will need to supplement this with other sources. The same goes for marine isotopic stage correlation. The dataset references MIS stages, but the correlations are not always consistent with the latest revisions of the deep time scale. I found discrepancies of up to fifteen thousand years when cross-checking against the Lisiecki and Raymo grid. For most macroecological work this does not matter. For someone working at the resolution of individual glacial-interglacial cycles, it absolutely does. There is also no built-in mechanism for tracking nomenclatural changes. Taxa get reclassified regularly in this field, and the dataset does not maintain synonyms or track taxonomic updates over time. You will need to manage that yourself, preferably through a controlled vocabulary or a taxonomic backbone like the Global Biodiversity Information Facility. Without that, you will end up with duplicate entries for the same species under different names and your diversity estimates will be inflated.
A Practical Workaround I Use
For the taxonomic issue, I maintain a lookup table that maps all known synonymies for the genera I work with most. I join this table onto the occurrence data before running any analyses. It adds maybe ten minutes to my preprocessing pipeline and prevents the kind of double-counting that makes reviewers uncomfortable. For the temporal resolution problem, I use the midpoint of each age range for exploratory work and propagate the full uncertainty interval for final figures. The difference in results is usually small but sometimes significant near key transition zones, particularly around the Matuyama-Brunhes boundary where dating precision matters more than people in this field tend to admit. The Last Two Million Years remains one of the more useful compilations for anyone doing large-scale analyses of hominin temporal and geographic patterns. It is not complete, it has quirks in its documentation, and it will require you to do some of the work that the authors assumed someone else would handle. That is typical for datasets in this space. You just need to know where the edges are before you build something on top of them.