Running RNA-seq on low-input samples is where most people hit their first wall

I set up my first RNA extraction pipeline back when we were working with 10,000 cells, not single cells. The library prep took three days and gave us barely readable reads. We wasted about six months troubleshooting before realizing the issue wasn't the sequencer or the bioinformatics pipeline. It was the RNA integrity number dropping below 7 during extraction, and nobody had been measuring it properly between steps. That's the kind of thing that doesn't show up in any textbook.

Getting Started With Genomics And Molecular Biology Workflows

You need to pick a starting point and stick with it long enough to actually understand the failure modes. Most beginners jump straight into sequencing without understanding what happens before the sample hits the machine. Let's start there. The first decision is whether you're working with DNA or RNA. They require completely different handling. RNA degrades in minutes if you let it. DNA is stable for years at minus twenty Celsius. If you're extracting RNA from tissue, you need RNase-free everything. Gloves, tips, tubes, water. Not because it's annoying but because RNases are everywhere. They're in your skin. They're on surfaces. They don't die when you autoclave them at standard temperatures. You need to bake glassware at two hundred fifty degrees for two hours or use DEPC-treated water. For DNA extraction, the standard phenol-chloroform method still produces the highest quality yield when you're working with fresh or frozen tissue. The spin-column kits are convenient but they cut longer fragments. If you need fragments above fifty kilobases for anything like long-read sequencing or certain cloning applications, spin columns will frustrate you. I learned this the hard way when I tried to assemble a large plasmid construct using kit-purified DNA and kept getting Shearing artifacts around the thirty thousand base pair mark.

Setting up your first sequencing library

I'll walk through Illumina short-read library prep since that's what most people end up using. The general workflow is fragmentation, end repair, A-tailing, adapter ligation, and PCR amplification. Each step has a time window and a temperature. Deviate from the recommended conditions and you'll see drop-offs in library complexity. Fragmentation can be done mechanically with sonication or enzymatically with transposases. The transposase method, sometimes called tagmentation, is faster but introduces sequence bias. Certain regions get over-represented because the transposase prefers open chromatin or specific sequence motifs. If you're doing ChIP-seq or ATAC-seq this matters a lot. For standard RNA-seq it usually doesn't make a noticeable difference. PCR amplification is where most people introduce duplicates. Every PCR cycle doubles your material but also doubles the chance that identical molecules get amplified together. After twelve cycles you're looking at maybe two percent duplicates. After eighteen cycles it jumps to fifteen or twenty percent. If your input material is limited and you need more cycles, consider using unique molecular identifiers. These are short random barcodes ligated before PCR that let you collapse reads from the same original molecule later.

Pipeline choices and what actually works

There are dozens of bioinformatics pipelines for processing sequencing data. The ones I recommend depend entirely on your experimental design. For RNA-seq quantification, I use Salmon or Kallisto for transcript-level quantification followed by tximport for gene-level aggregation. This approach is alignment-free and usually takes about ten minutes for a human sample on a standard laptop. The traditional method of aligning with STAR then counting with featureCounts takes roughly forty minutes and produces nearly identical results for most differential expression analyses. The speed difference only matters when you're processing hundreds of samples. For variant calling from DNA sequencing data, GATK's best practices pipeline is the standard but it's also extremely verbose and slow. The HaplotypeCaller does local de novo assembly of haplotypes which is accurate but computationally expensive. For a typical whole-exome sample it takes about two hours on eight cores. If you're working with whole-genome data at thirtyX coverage it can take a day or more on the same hardware. I switched to DeepVariant for our lab about two years ago. It uses a convolutional neural network to call variants from pileup images and produces comparable accuracy with significantly less hands-on time configuring the pipeline. A whole-genome sample at thirtyX runs in roughly six hours on a single GPU.

Get the Full Details

Molecular Biology, Genetics and Medical Concept. 3D Rendering Stock Illustration - Illustration ...
Molecular Biology, Genetics and Medical Concept. 3D Rendering Stock Illustration - Illustration ...

A specific problem that almost cost me three weeks

I ran a custom panel sequencing experiment targeting oncogenic mutations in circulating tumor DNA. The expected variant allele frequencies were below one percent. My initial analysis using GATK's default parameters called roughly four hundred variants per sample. When I checked these against known databases, about sixty percent were artifacts. The issue was that ctDNA libraries have extremely low complexity because you're sequencing fragmented cell-free DNA from plasma. Standard duplicate marking removed legitimate variants that shared the same start position simply because the fragments were short. The workaround was to switch to UMI-based error correction. I reprocessed the data using fgbio to group reads by UMI and build consensus sequences before variant calling. This reduced the artifact rate to below five percent and cut the variant call list from four hundred down to about forty real calls per sample. The downside is that fgbio requires careful UMI handling and you need enough sequencing depth to build reliable consensus. If your UMI diversity is low or your reads per UMI drop below five, the error correction becomes unreliable.

Common pitfalls that nobody warns you about

Batch effects will ruin your experiment if you don't plan for them. I once processed samples across three separate library prep days without randomizing between groups. The downstream clustering in PCA showed that samples grouped by prep day rather than by biological condition. We couldn't separate the signal from the noise without reprocessing half the samples. The fix is to randomize your sample processing order and include batch as a covariate in your statistical model. Contamination is another silent killer. Human DNA contamination in microbiome samples is almost unavoidable. Your reagents contain trace amounts of bacterial DNA. The solution isn't to eliminate it but to measure it. Run extraction blanks alongside your samples and subtract the background. Tools like decontam in R can automate this using prevalence-based or frequency-based methods. Reference genome choice matters more than most people realize. If you're working with a non-human organism, check whether the reference has recent updates. GRCh38 replaced GRCh37 several years ago and the coordinate differences cause problems when mixing data from different builds. I've seen people compare variants called against different genome builds and report positions that don't actually overlap. Always note your reference version and liftOver coordinates if you need to convert between versions.

When to skip next-generation sequencing entirely

Not every question needs NGS. If you're validating a known mutation in a small number of samples, Sanger sequencing is faster and cheaper. A single Sanger read costs about five dollars and takes two hours from PCR product to result. NGS overkill for that use case unless you're working with mixed populations or low-frequency variants. PCR followed by gel electrophoresis still has its place. Checking a cloning construct, verifying a gene knockout, or confirming primer specificity doesn't require a flow cell. I see too many people sequence a colony PCR product when a gel would have told them everything they needed to know in twenty minutes. The field moves fast and the tools keep changing. What I described here reflects the current state of practice but some of these recommendations will age out within a few years. The principles of good experimental design and critical evaluation of your data don't change though. Pay attention to your controls. Question your results before you publish them. And keep your raw data organized because you will need it when something goes wrong, which it eventually will.

Abstract DNA Helix Structure and Cells, Genetics and Molecular Biology Concept Stock ...
Abstract DNA Helix Structure and Cells, Genetics and Molecular Biology Concept Stock ...