Working With American History Primary Sources: What Actually Goes Wrong
I spent about four years digging through Federal Avenue Pension files for a family history project, and the most valuable thing I learned wasn't how to find them but how to stop trusting what I found on the first read. You pick up a document that looks definitive and move on. Two months later you realize the handwriting on page three contradicts the typed index on page one, and you've been citing the wrong date for your whole argument. The method most people get wrong is starting with the document rather than the context. Before you open any scanned letter or government record, you need to know what institutional framework created it and why. A Confederate soldier's diary entry means something entirely different than a Union quartermaster's supply ledger, even when they describe the same battle. The lens changes the facts. Start by identifying the record type and the agency or individual that produced it. Then check the processing history. Digitization projects sometimes renumber pages, drop entire folders from the microfilm, or merge separate documents that were originally filed independently. The National Archives maintains finding aids for most major collections, but those finding aids are themselves human-curated documents that contain errors. I found a Revolutionary War pension file where the folder label said "Johnson, William" but the documents inside belonged to a completely different William Johnson who had died thirty years earlier. The index card at the front of the file was misfiled somewhere in the 1970s and nobody caught it during the microfilming pass.
What I did was trace the box number and roll number back to the original manifest. The archive's online catalog lists those identifiers next to each record entry. Once I had the physical box location, I could request the unprocessed folder from the reference room instead of relying on the digitized version, and the misattribution became obvious immediately because the handwriting matched a different person entirely. This workaround takes about twenty minutes longer than ordering the digital copy, but it saves you from building an argument on a false premise. Here is something counter-intuitive that beginners consistently miss: a primary source is not automatically more authoritative than a secondary one simply because it is older. A soldier's letter written three days after a skirmish is full of confusion, misdirection, and honest uncertainty. A historian's monograph published forty years later may have corrected factual errors the soldier himself didn't know he'd made. The primary source gives you the raw material. It does not give you the truth. Use both, but weight them differently. Another common pitfall involves dates. Before September 1752, England and its colonies used the Julian calendar, which means March 25 was New Year's Day, not January 1. Documents from the 1600s and early 1700s often carry dual dates like "15 February 1684/5" to resolve the ambiguity. If you search a database using only the modern year, you will miss roughly half the relevant records. This matters enormously for colonial-era research and for any work involving probate records, land deeds, or court proceedings from that period.
I ran into this exact problem when tracking an ancestor's land transaction in Virginia. The deed was dated 12 March 1719 in the catalog entry, but the actual document text read "12 March 1718." The catalog had applied the modern calendar conversion without noting it. I caught the discrepancy because the neighboring entry in the same bound volume used the old-style date format, and the land description referenced a tree marking that had been established in 1717 by someone who was already dead by 1719 under the modern reckoning. The double-date convention would have made this transparent if the catalog had included it. Handwriting is another area where assumptions cause real damage. Secretary hand, running hand, and various regional cursive styles evolved across centuries, and automated transcription tools still struggle with pre-1800 documents at rates that vary wildly by scribe. I once spent three weeks trying to read a Federal Period tax list that turned out to be in a cramped, rushed hand where every "t" crossed like an "e" and every "a" looked like an "o." The transcription software output was almost entirely wrong. Going back to the original scan and reading each character slowly against the image, rather than relying on the generated text, cut my error rate from roughly forty percent down to somewhere under five percent. It takes longer, obviously, but the alternative is building a dataset of garbage. Provenance matters more than most people realize. A document sitting in a university archive usually has a cleaner chain of custody than one pulled from a private estate sale or a genealogy website. That doesn't mean university-held documents are trustworthy, but it does mean the risk of someone inserting, removing, or swapping pages is lower. When I worked with the Freedmen's Bureau records, I noticed that some files in the digital collection had page numbers that didn't match the corresponding microfilm roll. A few pages were missing entirely, and in one case a completely unrelated letter had been stapled into the folder. The digitization vendor had assembled composite files from multiple source materials without documenting the merges.
Get the Full Details

The workaround here is straightforward: always cross-reference the digital version against the microfilm or the original physical file when possible. The National Archives provides roll numbers for every collection. If a digital image seems off, mismatched, or oddly placed, pull up the roll and verify the sequence. This adds maybe fifteen minutes of work per document, but it prevents the kind of error that makes your entire research look careless. One more thing that isn't obvious: not everything labeled as a primary source actually is. Collections frequently include facsimiles, transcripts, and abridged versions marked with the original author's name but produced decades or centuries later. You'll find these especially in older published compilations and in some family history self-publications. The text might be accurate to the best of the compiler's ability, but it has been filtered through another person's interpretation, and sometimes their interpretation is wrong. Always go to the original image whenever the archive provides one. There are also sources that are primary in one context and derivative in another. A census transcript is a primary source for the act of census-taking itself, but it is a secondary source for anything about the people listed in it, since the enumerator's understanding of names, ages, and relationships may be flawed. Treat each layer separately and cite accordingly.
The biggest bottleneck in primary source research is usually time, not access. Most major collections are digitized to some degree, but the quality varies from excellent to practically unusable. The Library of Congress, the National Archives, state archive systems, and university special collections all have different scanning standards. Some provide high-resolution TIFFs with page-level metadata. Others provide low-resolution JPEGs with no structured data. I've spent hours trying to work from poor scans only to discover that a single good look at the original would have resolved every question in twenty minutes. When possible, prioritize collections with IIIF support or at minimum downloadable high-resolution images. The difference between a screen-sized preview and a proper TIFF can be the difference between reading a faded signature and guessing at it. This is especially critical for documents with water damage, foxing, or ink fade, all of which are common in eighteenth and nineteenth century paper. If you are doing serious work in this area, learning to read the physical characteristics of the document itself will serve you better than any research guide. Paper type, watermark, ruling, binding method, ink composition, and seal wax all carry information that text alone cannot provide. A document printed on machine-made paper with a post-1840 watermark cannot genuinely be from the colonial period, no matter what the content says. I encountered a forged Loyalist petition that looked authentic at first glance because the language and format were correct, but the paper date and the seal wax composition were both anachronistic. The archive had accepted it into the collection because no one had examined the physical medium.
These kinds of problems are rare but they happen. Being able to spot them requires practice and a willingness to look past the text and at the object. That means spending time in reading rooms, handling documents under supervision, and paying attention to the material details rather than just extracting information and moving on. For practical next steps, start with the records closest to your question rather than browsing broadly. Define what you need before you start looking, because the volume of available primary source material is large enough to swallow a project if you let it. The Library of Congress Chronicling America database is useful for newspapers, but its coverage is uneven and biased toward urban centers and English-language publications. State archives hold probate, land, and court records that are often more valuable for social history than federal collections. County clerks' offices retain original records that were never sent to the state or federal level. Use the finding aids as starting points, not destinations. Every major archive publishes some form of inventory or guide, and those guides are imperfect. They omit subseries, they misdate files, and they sometimes describe records that were weeded out during accession processing. Verify what you find against the actual catalog entries and the container lists, and be prepared to adjust your expectations when the guide doesn't match the collection.

The work is slow, the errors are real, and the material degrades regardless of how carefully you handle it. But the documents are there, and most of them have not been thoroughly examined by anyone outside a small circle of specialists. If you approach them with skepticism toward your own initial readings and a habit of verifying against the original image, you will produce work that is more reliable than what most people turn out. That is the practical value of primary source research, and it is the only measure that matters.