Born-Digital Archiving Is a Mess. Here Is How You Actually Deal With It.
Most people think archiving digital material is just copying files into a folder and calling it a day. That works fine until you have a dataset that contains active software, nested formats, broken file associations, and no documentation. Then it does not work at all. Stephen Musacco has been writing about this problem for years, and his work on Going Postal Stephen Musacco Phd touches on the same reality: the moment you try to preserve born-digital records, you are dealing with obsolescence, format decay, and metadata gaps that never existed in the paper world. I spent about three years managing a small institutional collection of born-digital records, mostly from the early-to-mid 2000s. The kind of stuff that came in on DVDs, Zip drives, and occasionally hard drives that were so full you could barely read the labels. The problem was never simple format migration. It was figuring out what the files actually were, whether they still opened, and whether they still meant anything. You cannot just run a conversion tool and move on.
Going Postal Stephen Musacco Phd
Musacco's work focuses heavily on the practical side of digital preservation, particularly around born-digital collections and the institutional challenges of acquiring and maintaining them. The term "going postal" in this context is not about mailing anything. It is about recognizing that digital materials, when treated the same way we treat physical archives, fall apart. The infrastructure breaks, the formats rot, and the context disappears. His PhD-level research and subsequent publications make the case that archival institutions need a fundamentally different approach for digital-born materials than for digitized surrogates. Here is a practical breakdown of how this actually plays out in a real workflow. Step one: intake and documentation. When a donor or department submits digital material, you need to know what arrived, in what condition, and with what surrounding metadata. This means checking file counts, formats, checksums, and any accompanying documentation. I learned this the hard way. I once processed a collection that appeared to be 400 PowerPoint presentations. They were all corrupted beyond recognition because they had been saved from an older version of PowerPoint without the proper embedded fonts and macros. The metadata file, which was a text document left on the desktop, was the only thing that told us what each file was supposed to contain. Without that, we had 400 empty presentations and no idea how to reconstruct them.
Step two: format identification and risk assessment. You need to determine what formats you are working with and prioritize which ones are at highest risk of becoming unreadable. Common at-risk formats include older Adobe Flash files, proprietary CAD files, and legacy database formats. Stable formats like PDF/A, TIFF, and plain text should not get the same level of urgency. This is not rocket science, but it does take time. In my experience, a modest collection of 5,000 to 10,000 files typically requires about two weeks of full-time format identification and risk prioritization, depending on how messy the incoming material is. Step three: preservation copies and migration. For at-risk formats, you create preservation copies in stable formats when possible. For some formats, especially complex or interactive materials, true migration is impossible and you need to consider emulation or virtualization instead. This is where most institutional workflows break down. People assume migration solves everything. It does not. A migrated spreadsheet might open fine, but the formulas, macros, and formatting can change behavior in ways that alter the record. I encountered this with a series of environmental monitoring spreadsheets from a county agency. The original files had VBA macros that calculated pollution thresholds. After migration to a modern Excel version, those macros ran differently and produced slightly different results. The data was still there, but the institutional logic behind it was altered. Step four: metadata capture and provenance tracking. You need to document the chain of custody, the transformations applied, and the conditions under which the material was acquired. PREMIS metadata is the standard for this, though implementing it correctly requires effort and staff time that many institutions do not budget for. A common mistake is treating metadata as an afterthought. It is not. If you acquire born-digital material without capturing provenance metadata at the point of intake, you will spend three times as long trying to reconstruct that information later, and you will never succeed completely.
Get the Full Details

Musacco's contribution to this field is that he pushes back against the assumption that digital preservation is simply a technical problem waiting for the right tool. It is an institutional problem. Funding models, staffing structures, and acquisition policies all shape whether born-digital materials survive or vanish. His research demonstrates that well-funded institutions with dedicated digital preservation staff outperform underfunded ones not because they have better software, but because they have time to handle the messiness that comes with digital materials. There are limitations to the current state of digital preservation that nobody likes to discuss publicly. First, interactive and software-dependent materials remain largely unsolved. Emulation is possible for some applications, but it is expensive and slow. Second, there is no universal standard for born-digital acquisition that all institutions follow. Third, the cost of long-term storage is often underestimated. Cloud storage sounds cheap until you are paying egress fees on millions of files and realizing you cannot easily audit what you have without specialized tools. If you are working with a small collection and do not have a full digital preservation team, the best approach is to start small. Focus on the most at-risk formats first. Document everything. Build a simple but functional metadata framework. Do not wait until you have a crisis to figure out what you are doing. The people who struggle most with born-digital materials are the ones who assumed they would figure it out as they went along. They did not.
The field is moving slowly toward better tools and standards, but the gap between what the literature says should happen and what actually happens in practice remains wide. Musacco's work helps narrow that gap by keeping the focus on the institutional realities rather than the technical ideals.