A Practical Guide to Working With Coronavirus Sequence Data

I've spent the better part of three years parsing SARS-CoV-2 genomic datasets, and the first time I tried to work with raw GISAID exports, I nearly quit. The files are messy, the metadata is inconsistent, and the documentation is worse than most bioinformatics projects I've seen. But once you get past the initial friction, the workflow settles into something repeatable. The phrase "courting the coronavirus" isn't standard terminology, but it describes exactly what people end up doing: you invest significant time, you learn the quirks of the data, you develop relationships with the various databases and pipelines, and eventually you stop fighting the format and start working with it. That shift is the entire project. Here's how the actual process works when you're pulling and analyzing SARS-CoV-2 sequences from scratch.

Start with Nextstrain'saugur pipeline. It's the most straightforward entry point, and the Docker setup alone saves approximately six hours of dependency troubleshooting compared to a manual conda install. Pull the global dataset using their build script, feed it through the align step, and you'll get a phylogenetic tree with mutation annotations in roughly forty minutes on a standard laptop. That's the baseline. The problem most people hit is strain designation. GISAID uses their own internal strain naming convention that frequently conflicts with lineages from Pango or Nextclade. I spent two weeks debugging a color-mapping issue only to realize the metadata file I was feeding into augur had overlapping strain IDs from a previous export that hadn't been refreshed. The workaround is simple but easy to miss: before every analysis run, verify your strain IDs against the most recent Nextstrain clade assignment file, not the raw GISAID download. I set up a small script that cross-references both and flags mismatches automatically. Runs about 12 seconds per dataset. For mutation calling, nextclade is the standard. The QC thresholds are reasonable — they flag sequences with too many missing positions, excessive mutations, or suspicious deletions. But here's what the docs don't make clear: nextclade's ambiguous site handling can produce misleading results when you're working with early-pandemic sequences from lower-coverage genomes. Those older sequences often have N's in the spike gene region, and nextclade will still assign them to a lineage based on whatever partial data exists. You get a clean-looking result that's actually built on sparse evidence. I learned this when I was building a regional transmission model and the tree kept grouping a handful of samples into improbable clusters. The fix was applying a minimum quality score filter before running the alignment, which I'd set at Q20 across at least 90 percent of the genome.

Metadata cleanup is the part nobody enjoys. GISAID submissions come with fields like "division" that mean different things in different countries, collection dates that are sometimes just year-month with no day, and host information that occasionally reads as blank or populated with random text. For any serious analysis, you need a normalization step. I use a combination of pandas to standardize the date fields into YYYY-MM-DD format and a lookup table that maps irregular administrative divisions to consistent regions. Takes about twenty minutes for a full global dataset. If you're working with variant frequency over time, the simplest approach is to bin the nextclade lineage assignments by collection date and calculate proportions per week. The edge case here is when you're tracking a rapidly replacing variant — like what happened with Omicron's initial wave. The frequency curves can look artificially stepped because of sampling bias, not actual epidemiological dynamics. More sequences from certain regions in certain weeks will skew the numbers. I've found that smoothing with a seven-day moving average and cross-checking against CEPI's sequence counts helps separate signal from sampling noise. For visualization, Auspice (Nextstrain's web interface) handles most needs. Export the JSON and serve it locally or push it to a static host. If you need publication-quality figures, export the Newick tree and color it in iTOL or ggtree in R. The latter gives you far more control over layout and annotation, though it requires learning the tidytree package structure.

Get the Full Details

EU revises China coronavirus report, courting controversy
EU revises China coronavirus report, courting controversy

Hardware requirements are modest. A machine with 8GB RAM and a four-core processor handles datasets up to about 50,000 sequences comfortably. Beyond that, you'll want to subsample or move to a cloud instance. The compute cost for a full global run on AWS is roughly $2 to $4 depending on instance type and whether you use spot pricing. Nothing dramatic. There are also limitations you should acknowledge upfront. The primary one is that sequence data is a convenience sample, not a representative one. Countries with robust sequencing infrastructure will dominate any frequency analysis, which means your conclusions about global variant spread are really conclusions about global sequencing capacity. Secondary, the mutation rate assumptions built into most phylogenetic tools don't account for within-host diversity. A single consensus sequence from a patient masks the quasispecies structure, which matters if you're studying transmission bottlenecks or drug resistance. The alternative if you don't want to manage the full pipeline is to use the precomputed Nextstrain datasets directly. They update continuously and cover the vast majority of use cases without any local setup. The tradeoff is that you're dependent on their curation choices and you can't modify the alignment or substitution model parameters.

I keep a small library of helper scripts for common tasks — metadata normalization, subsampling by date and geography, exporting lineage tables for external analysis. They're not published anywhere formal, just organized in a personal repo. If you're starting out, cloning and adapting someone else's workflow is faster than building from scratch, but don't skip reading through what they actually do. I've seen too many people copy a pipeline without understanding which steps are essential and which are optimizations for a specific dataset. The field moves fast enough that tutorials are rarely current more than six months after publication. The core tools don't change drastically, but the expected practices around metadata standards and quality filtering do. Staying close to the Nextstrain documentation and the GISAID data use policy pages is probably the most reliable way to keep up.

Tools and Resources

Essential Software

augur — Phylogenetic analysis and visualization pipeline from the Nextstrain team. Handles alignment, tree building, and clade assignment. nextclade — Quality control, mutation calling, and Pango lineage assignment for SARS-CoV-2 sequences. Auspice — Web-based visualization interface for Nextstrain trees. No compilation required if you use the Docker image.

Courting innovation: Australia’s judicial system in the context of COVID-19 Webinar - YouTube
Courting innovation: Australia’s judicial system in the context of COVID-19 Webinar - YouTube

ggtree / tidytree — R packages for custom tree visualization and annotation. Necessary if you need figures for papers or reports.

Data Sources

GISAID — Primary repository for submitted SARS-CoV-2 sequences and metadata. Requires registration and acceptance of their data use terms. NCBI Virus — Alternative or supplementary archive, often easier to query programmatically via Entrez. CoV-Seq — UK-specific sequence data hosted by COG-UK. Useful for regional analysis focused on British sample collections.

Helper Scripts and Templates

The community maintains several starter repositories on GitHub. Look for ones that explicitly mention Nextstrain version compatibility, since interface changes between major releases can break older scripts. The most reliable ones include a requirements.txt or environment.yml file pinned to specific versions rather than allowing flexible dependency resolution.

Coronavirus Proposals: Why These Couples Got Engaged During COVID
Coronavirus Proposals: Why These Couples Got Engaged During COVID