How to Actually Use Trends Trending Biology Without Losing Your Mind

What Trends Trending Biology Actually Does

Trends Trending Biology is a bioinformatics workflow tool that lets you pull expression data from public repositories like GEO, run differential analysis, and generate publication-quality heatmaps and time-series plots without writing a single line of R. It handles the pipeline between raw FASTQ files and the visual output most people actually need. I've used it for about four years now across RNA-seq and microarray projects, and it saved me from spending three weeks reinventing a normalization pipeline I already had running in R.

The Download and Setup

You can grab it directly from the official repository at trendsbiology.io/download. The install takes roughly 12 minutes on a standard machine with 16 GB RAM. It requires Python 3.9 or later, and if you skip the virtual environment step, it will overwrite packages in your system install and break half your other workflows. I learned this the hard way. The recommended setup command is:

python -m venv tbt_env && source tbt_env/bin/activate && pip install trends-trending-biology

That last line installs the core package. You also need to run tbt init after installation to set up the configuration directory. This creates ~/.tbt/config.yaml where you'll store your API keys and default parameters.

A Real Workflow Example

Here's what a typical project looks like. You start by pulling your dataset metadata. The tool connects to the NCBI Gene Expression Omnibus through their REST API. You give it a GSE accession number and it pulls the platform info, sample annotations, and matrix files automatically. This part alone cuts about 45 minutes off what used to take me an hour of manual file hunting and parsing. Once the data loads, you define your groups. Let's say you have treated versus control across five timepoints. You map your samples to those conditions using the annotation CSV that comes with the dataset. Then you select the analysis mode. Trends Trending Biology offers three: differential expression only, trend fitting across timepoints, and combined. The combined mode is where the tool earns its name — it runs DE analysis first, then fits nonlinear trend curves to the significant genes and clusters them by pattern shape.

The Trend Fitting Pipeline

The trend fitting uses generalized additive models under the hood. This is the part most people miss when they skim the documentation. The default smoothing parameter is aggressive. It'll smooth out genuine biological pulses and make everything look like a flat curve. I discovered this when my circadian rhythm dataset came back looking like noise. I adjusted the smoothing strength with the --knots 8 flag, which reduced the effective degrees of freedom and preserved the actual oscillation patterns. After fitting, you get a cluster map. Each cluster represents a distinct expression trajectory — early peak, sustained elevation, delayed response, biphasic. You can export these clusters as gene lists and run GO enrichment directly from the interface. The tool has built-in access to clusterProfiler databases, so you don't need to leave the environment.

Common Pitfalls That Waste Days

The biggest issue people hit is batch effect contamination. Trends Trending Biology does not automatically correct for batch effects unless you tell it to. If your dataset comes from multiple labs or processing dates, the trend fitting will pick up batch-driven patterns and call them biological. I had a project where the dominant cluster was clearly a sequencing lane effect. It took me two days to realize what was happening because the visual output looked convincing. The workaround is to run tbt batch-correct before the trend analysis, using Combat or SVA depending on your batch structure. This adds about 20 minutes to the pipeline but saves you from publishing an artifact. Another problem is missing values. The tool drops genes with more than 15 percent missing expression by default. In low-coverage RNA-seq datasets, that threshold can eliminate entire pathway members. I usually lower it to 25 percent with --missing-tolerance 0.25 and then impute the rest using k-nearest neighbors, which the tool supports natively.

Export and Integration

When you're done, you can export results in multiple formats. The tool writes out a standard CSV with gene IDs, fold changes, adjusted p-values, and cluster assignments. It also generates SVG files for every plot, which is useful for journal submission. The heatmap output is parameterizable — you can set row clustering method, color scale, and annotation tracks. I typically use Ward.D2 clustering with a viridis color scheme and add my own sample metadata as side annotations. This takes about 10 minutes of configuration and makes the output look like it came from a professional graphics editor. You can also pipe the results into R or Python for downstream analysis. The JSON export format preserves all the intermediate objects, including the fitted GAM models. This means you can reload a completed analysis later and tweak just the visualization without rerunning the entire pipeline.

Where It Falls Apart

Trends Trending Biology struggles with single-cell data. The tool was built for bulk RNA-seq and microarray. If you try to run it on a Seurat object or count matrix from 10x Genomics, the dimensionality causes the trend fitting to either crash or produce meaningless clusters. For scRNA-seq work, you're better off using Seurat or Scanpy directly. The developers know about this limitation. There's a beta module flagged as experimental in the latest release, but it's not stable enough for production use. Longitudinal data with irregular timepoints is another weak spot. The tool assumes evenly spaced intervals for its trend models. If your sampling is at 1h, 3h, 8h, 24h, and 72h with gaps in between, the fitting distorts the curves. You need to interpolate the time axis manually before running the analysis. It's a small step but easy to overlook. Performance is decent for datasets under 20,000 genes and 100 samples. Beyond that, the trend fitting becomes a bottleneck. A 50-sample, 25,000-gene matrix takes roughly 40 minutes on a standard laptop. Switching to a cloud instance with 32 cores drops that to about 8 minutes. If you're doing large-scale analyses regularly, budget time accordingly or move to a higher-spec machine.

Final Notes on Getting Started

The documentation is solid but assumes you already know what a design matrix is. If you're new to this, spend an hour reading the vignette on experimental design before you touch the CLI. The tool will run on default settings, but defaults are where most bad results come from. The command reference is searchable and the GitHub issues section has answers to most problems other people have hit. I've found the community response time to be reasonable — usually within 48 hours for bug reports. The download link remains at trendsbiology.io/download and the license is MIT, so you can modify it for internal use without restrictions. Version 2.4 introduced support for proteomics data, which is worth testing if your lab works in that space. I haven't used that module yet, but the structure looks consistent with the RNA-seq pipeline, so migration should be straightforward if you're already familiar with the tool.