Setting Up Your African American History Project

I spent three weeks last year trying to get the African American History Project working properly for a community archive we were building. Most tutorials online skip over the messy parts—the encoding issues, the metadata mismatches, the way the default configuration assumes you already have a clean Unix environment. Here's what actually works. It's an open-source archival framework designed specifically for organizing and publishing African American historical records. The repo is structured around TEI-compliant text encoding, relational metadata schemas, and a Flask-based web interface. You pull transcripts, photographs, oral histories, and institutional records into a unified searchable repository. That's the pitch. The reality involves more wrestling with Python dependencies than most people expect. Clone the repository from GitHub. The latest release requires Python 3.9 minimum. Earlier versions claimed 3.8 compatibility but broke on import when you tried to run the metadata migration scripts. Don't bother with 3.8. Use a virtual environment—this thing installs six incompatible versions of lxml in its dependency tree and it will quietly wreck your system Python if you let it.

After cloning, run pip install -r requirements.txt. Then copy config.example.yaml to config.yaml and edit it. The config file controls your database connection, archive root path, and authentication backend. The default SQLite setup works fine for local testing, but if you're doing more than twenty simultaneous users, switch to PostgreSQL before deployment. The project's ORM layer handles both equally well, and migrating later requires rebuilding your entire search index from scratch.

Structuring Your Archive Data

The project expects a specific directory layout. Records go under /records, metadata under /metadata, and media files under /assets. I learned this the hard way after spending a day troubleshooting why the OAI-PMH harvest endpoint kept returning empty results—it turned out I'd put my XML metadata files in the wrong subdirectory by one folder level. The error logs didn't mention this at all. Each record needs an associated TEI XML file. The schema files are in the docs/ directory of the repo. They're thorough but not beginner-friendly. If you're importing from an existing collection, plan on spending real time mapping your current metadata fields to the project's schema. The included import script (scripts/import_records.py) handles CSV and simple XML, but it doesn't do field translation. You either write your own mapping layer or spend hours in the database admin interface fixing mismatches manually.

Get the Full Details

Digital Black History Month Project- African American Scientist and Inventors
Digital Black History Month Project- African American Scientist and Inventors

Search and Indexing

The search backend is Solr. The project ships with a Solr configuration in solr/config/. It uses custom analyzers tuned for the kinds of spelling variations you find in historical documents—archaic spellings, phonetic transcriptions of names, OCR errors from scanned newspapers. This is where the project actually earns its keep. Standard search setups miss half the records in a well-sourced African American history collection because they don't account for the naming conventions in Freedmen's Bureau records or the OCR drift in period newspapers. Running the indexer is straightforward once Solr is up. python manage.py index rebuild. This scans your records directory and rebuilds the full-text index. For a medium-sized collection—maybe five thousand records with associated transcripts—it takes roughly twelve minutes on a modern machine. Don't interrupt it. I've corrupted the index twice this way and both times had to start from zero.

A Real Problem I Hit

One thing nobody mentions: the date range filtering breaks on records without complete dates. The project stores dates as ISO 8601 strings, and when a historical document only has an approximate year—say, "circa 1865"—the filter throws a type error and silently drops that record from search results. The UI doesn't even show an error. It just won't appear. I found out after a community member reported that a specific document she knew existed was "missing from the archive." The workaround is to create a preprocessing step that normalizes approximate dates into range bounds before they hit the indexer. I wrote a small utility function that converts "circa" and "between" style dates into start-end pairs the search backend can handle. It's not part of the main repo. If anyone maintains a patch for this, I'd recommend pulling it in early rather than dealing with it after your collection is already indexed.

Export and Distribution

The project supports OAI-PMH out of the box, which matters if you're sharing with other archives or institutions. Set up the endpoint in your config and point it at your hosted instance. Most digital collections platforms can harvest from it. You also get static site generation if you want a flat HTML export for backup or offline distribution. The authentication system supports LDAP and OAuth2. If you're running this for a library or university, LDAP integration is the cleanest path. I went with OAuth2 for a community-run archive because it let people use their Google accounts to submit transcriptions and corrections without managing separate credentials. Budget about an afternoon for the OAuth2 setup. The docs cover it but the callback URL configuration tripped me up until I realized the nginx reverse proxy was stripping the port from the redirect.

Black History Month Research Project Famous African American Study Resource
Black History Month Research Project Famous African American Study Resource

Common Pitfalls

The memory consumption on larger collections is the biggest issue. Each record gets loaded into memory during search operations. A collection of ten thousand records with full transcripts runs the Flask worker at around 400MB per process. Without proper worker limits in your deployment config, a single busy search query can push the container to swap. Set your Gunicorn workers explicitly and keep the per-worker memory cap at 512MB. This isn't optional if you're running on anything less than 8GB of RAM. Another thing: the backup script (scripts/backup.sh) doesn't include the Solr index. It backs up the database and the file assets, but your search index is separate. I set up a cron job that runs the index export alongside the standard backup. Without it, a disaster recovery scenario would leave you with all your data but no way to search it until you rebuilt the index from scratch.

Where This Falls Short

The African American History Project isn't a turnkey solution. It's a framework, and like most frameworks, it assumes you can handle the assumptions it makes. The documentation covers the happy path pretty well. The edge cases—partial dates, bulk imports from non-standard sources, multilingual metadata—are where you'll spend your time. It also hasn't seen a major release in over a year as of my last check, so the Python version compatibility drifts over time. If you need something fully managed with less customization, there are commercial alternatives like PastPerfect or Tropic. But those cost money and lock you into their data models. If you're building a permanent archive that belongs to the community it serves, this project gives you ownership of the stack and the data. The tradeoff is the work required to make it yours.

Next Steps

Start with a small test collection—maybe fifty records from a single source like a church register or a newspaper issue. Get the full pipeline working end to end before you commit to a larger archive. Once you've done that once, the second time is significantly faster. The hardest part is always the first mapping between your existing materials and the project's schema assumptions. The repo is on GitHub under the Sapiens AI organization. The documentation lives in the docs/ folder, and the examples/ directory has a sample collection you can load to verify your installation before touching real materials.

African American History Research Project | Black History Month | TPT
African American History Research Project | Black History Month | TPT