Getting Started With Data Science For Librarians
You walk into a library, pull out your laptop, and realize you have three thousand circulation records that haven't been cleaned since 2019. The Excel file has missing values scattered across columns, ISBNs in seven different formats, and at least one row where someone typed "I love reading!!!" into a date field. This is where Data Science For Librarians actually begins. Not with a workshop or a fancy dashboard, but with a messy spreadsheet that refuses to cooperate. The most practical starting point is Python with pandas. You install Anaconda, open a Jupyter notebook, and load your data. A typical workflow for cleaning library records looks something like this: import pandas, read the CSV, check for null values with df.isnull().sum(), then start fixing things column by column. The first column you usually encounter is patron IDs. They come in different lengths depending on whether the system was migrated from an older ILS. Some branches use nine-digit numeric codes, others use alphanumeric strings with a leading zero. Normalizing these takes about twenty lines of code and five minutes to run, which is considerably faster than whatever manual process your staff has been using. For bibliographic data, the real challenge is deduplication. Two copies of the same book get cataloged under slightly different records because the OCLC numbers don't match or the author name has an extra space. I spent a whole Tuesday in 2023 trying to merge two datasets from different branch systems where one used "King, Stephen" and the other used "Stephen King" as the author field. A simple string comparison failed because the fields were structured differently. The workaround was extracting the author's surname, lowercasing everything, and comparing on a normalized composite key of normalized author plus title minus punctuation. It caught about ninety-four percent of duplicates on the first pass. The remaining six percent required manual review, which took another three hours.
Visualization comes after the cleaning stage. Matplotlib and seaborn handle most standard charts fine. If you need to show circulation trends over time, a basic line plot with df.plot(x='date', y='circulations') gets you there in thirty seconds. For heatmaps showing peak hours by branch location, you aggregate the data with a pivot table first, then pass it to seaborn's heatmap function. The output is crude but functional. Library directors don't need publication-quality graphics. They need something they can glance at during a meeting and immediately understand which branch is underperforming. More advanced work involves text mining from patrons' reading lists or complaint forms. The library I worked with received about four hundred feedback entries each semester that were stored as free-text PDFs. Extracting the useful parts required reading the PDFs with PyPDF2, feeding the text into spaCy for named entity recognition, and tagging mentions of specific authors, genres, or equipment problems. This pipeline took about a week to build properly, but once it was running it processed a semester's worth of feedback in roughly forty minutes instead of the twenty-five hours it would have taken by hand. One thing nobody tells you about applying data science methods to library collections is how much domain knowledge matters. A machine learning model trained on circulation data will flag any book with low checkouts as "low demand." That sounds correct until you consider that some materials simply serve a different purpose. Reference works, archival documents, and reserve materials are meant to be consulted on-site, not checked out repeatedly. Without understanding that distinction, your analysis will recommend weeding perfectly valuable materials. I learned this the hard way when a colleague's automated recommendation engine suggested removing half the microfilm collection because it had zero circulations in five years. The microfilm housed local genealogical records that patrons consulted exclusively in the reading room. Those records were never going to circulate. The workaround was adding a flag for on-site-only materials and excluding them from any weeding algorithm.
SQL is still necessary if your library's ILS exposes a database layer. Most modern systems like Sierra or Alma have SQL access. Writing queries to pull patron demographics, material usage patterns, and interlibrary loan statistics is faster than exporting to CSV and processing externally. A well-written SQL query can join circulation, holdings, and patron tables in seconds and return the exact subset you need without loading your entire database into memory. There are limitations worth acknowledging upfront. The biggest one is data quality. If your catalog records are incomplete or your circulation logs have gaps due to system outages, any analysis built on top of that data is unreliable. Garbage in, garbage out applies here with particular viciousness. Another limitation is that many library databases don't export cleanly into standard formats. MARC21 data in particular is painful to parse programmatically unless you use a dedicated library library like pymarc. Python libraries for MARC exist but the documentation is sparse and the learning curve is steep. A common pitfall is over-engineering. Beginners often write elaborate pipelines for tasks that could be solved with a simple pivot table in Excel. Before you build a full automated reporting system, verify that a manual approach gives you the same answer. If it does, you've just saved yourself three days of development time.
Get the Full Details
Building a Working Pipeline
Start small. Pick one dataset that causes you repeated pain. Maybe it's the patron attendance logs from your children's programs. Maybe it's the vendor invoice spreadsheet that your acquisitions department maintains. Clean it. Document what you did. Save the script. Repeat. After three or four iterations you'll have a collection of reusable functions and a much clearer sense of what your data actually looks like. For installation, Anaconda is the most straightforward path. It bundles Python, pandas, numpy, matplotlib, and Jupyter. Download it from anaconda.com and install the standard package. If you're working on a shared lab computer where you can't install software, Google Colab runs in the browser and gives you access to the same libraries with no setup required. Libraries like pynetmz and pymarc are worth looking into if you deal with networked metadata or MARC records respectively. pymarc handles MARC21 parsing and writing and is the standard tool for anyone working with bibliographic data at scale. You can install it via pip with pip install pymarc.
The reality is that most library data projects don't need machine learning. They need better data hygiene and clearer reporting. Focus on cleaning, structuring, and presenting what you already have before reaching for clustering algorithms or predictive models. The results will be more useful and significantly less work.