Getting Started With NGS Workflows
The NGS Data Analysis Tutorial basically covers the open-source pipelines built around Nextflow and nf-core. The most common setup uses a Galaxy interface combined with community-maintained workflows for RNA-seq, WGS, and ChIP-seq. If you have never touched a Nextflow pipeline before, expect your first attempt to take several hours. The second attempt will probably take twenty minutes. Here is how the standard pipeline moves data forward. Raw FASTQ files enter the system, the pipeline trims adapters, maps reads against a reference genome, calls variants, and writes a final VCF. That is the ideal sequence. The reality involves a dozen intermediate steps and a lot of temporary files filling your disk space. The core tools involved include fastp for adapter trimming, BWA-MEM or STAR for alignment, Picard for duplicate marking, and GATK for variant calling. After that comes VQSR or hard filtering, then annotation with tools like SnpEff or VEP. The exact order changes depending on which nf-core workflow you are running. RNA-seq uses Salmon or Kallisto for quantification instead of alignment-based counting. WGS workflows skip transcript-level steps entirely.
I used to think skipping the adapter trimming step was acceptable when I knew the library preparation kit. It is not. Adapters from a TruSeq prep will sit at the end of your reads and cause alignment artifacts that look like real variants if you do not catch them early. I lost a full day reanalyzing a dataset because I assumed the manufacturer's kit handled everything cleanly. It does not.
Installation and First Run
You need Nextflow installed first. Use the default installer and make sure you are running version 22.10 or later. Older versions have known bugs with AWS S3 transfers that will silently corrupt your input files during large batch runs. After Nextflow, you pull a pipeline repository. For RNA-seq, the go-to is nf-core/rnaseq. For whole genome sequencing, it is nf-core/wgbs or the gatk workflow depending on your project type. Clone the repo or use the Nextflow run command directly from GitLab or GitHub. The difference between cloning and using the remote reference is mostly convenience. Cloning gives you access to the process_config files and lets you tweak parameters without rebuilding the entire command line. The params file controls sample metadata. You will submit a CSV or TSV with columns for sample name, read path, and strandedness. Get the strandedness wrong and your gene counts will be approximately half what they should be. I learned this the hard way with a paired-end RNA-seq dataset where the library was stranded but the metadata said unstranded. The downstream differential expression results were completely nonsensical until I caught the error and reran with the correct flag.
Get the Full Details
Common Pitfalls That Break Pipelines
Resource allocation is the first place things go sideways. A typical WGS pipeline needs about 128 gigabytes of RAM and four CPU cores per sample for the alignment and variant calling stages. If you underspec the memory, GATK HaplotypeCaller will throw an obscure error about interval trees and quit without a clear message. Increase the memory to at least 64G for haplotype-based callers and 32G for germline SNV calling. The pipeline will still run correctly. It will just take longer. Disk usage is the second issue. A single human WGS run at 30x coverage generates roughly 800 gigabytes of intermediate files. Your temp directory will fill up fast if you are not monitoring it. I set up a cron job to delete files older than seven days in my Nextflow work directory. That reduced disk exhaustion incidents from weekly to about once a month. Bwa-mem and Bwa-mem2 give different alignment statistics on the same data. Bwa-mem2 is faster but produces slightly different mapping quality scores. If you are comparing results across studies, stick to one aligner and document which one you used. Reviewers will ask.
A Specific Problem I Encountered
During a recent bulk RNA-seq run, the pipeline hung at the alignment stage for a sample that other samples processed fine. The log showed no error. It just sat there with zero CPU usage for over an hour. I checked the sample sheet and realized the read path was pointing to a symlink that had been broken during a previous filesystem move. The pipeline tried to open the file, got an empty handle, and quietly waited. I replaced the symlink with the absolute path and reran. The alignment finished in twenty minutes. This happens more often than people admit, especially on shared clusters where user directories get migrated between nodes. The VCF that comes out of the pipeline is not production-ready. You will still need to filter for depth, genotype quality, and allele balance. GATK recommends VQSR for variant calling, but VQSR requires at least 30 samples to train properly. If you have fewer, use hard filtering with the standard GATK recommendations as a starting point. The default cutoffs are conservative but not wrong. Adjust them based on your validation data. For RNA-seq, the count matrix from featureCounts or Salmon needs normalization before any statistical testing. Use DESeq2 or edgeR for differential expression. Do not skip the normalization step. Raw counts will produce false positives at a rate that makes the results useless for publication.
The NGS Data Analysis Tutorial resources available online cover the basics adequately, but they do not always explain why certain parameters matter. I recommend reading the nf-core documentation alongside the tutorial. The documentation explains the parameter choices. The tutorial shows you how to run the pipeline. You need both. If your project involves rare variants or structural variants, the standard germline pipeline will miss a lot. Consider adding a separate SV calling step with Manta or Delly, and use DeepVariant for SNV calling instead of GATK if you have GPU access. DeepVariant tends to perform better on difficult genomic regions, especially around indels and homopolymers. The tradeoff is longer runtime and higher GPU memory requirements. Container support is available through Singularity and Docker. Use Singularity on HPC clusters. Docker works fine on local machines but may require root privileges depending on your installation. I stopped using Docker on cluster systems after a permissions issue corrupted my reference genome cache. Singularity is slightly more verbose to set up but does not have that problem.