Getting started with Simons Lab

Simons Lab is a bioinformatics and computational biology research platform operated under the Simons Foundation umbrella. It provides tools for single-cell genomics, spatial transcriptomics, and large-scale omics data analysis. If you're coming from a wet-lab background, the first thing you need to adjust to is that this isn't a point-and-click wizard — it's a suite of programs and notebooks that expect you to know your way around Python and R. I spent about three weeks last year working through their spatial transcriptomics pipeline on a batch of mouse brain sections, and honestly, the documentation assumes you've already read half the papers it references. You'll save yourself a lot of pain by skimming the methods sections of the key publications before you touch a single command.

Simons Lab download and setup

The primary codebase lives on GitHub under the Simons Foundation organization. You'll want to clone the relevant repository for your use case — most people start with the scRNA-seq preprocessing workflows. The installation isn't trivial. Their environment relies on a mix of Conda packages and some dependencies that don't play nicely on older systems. I'd recommend spinning up a fresh Conda environment with Python 3.10 or later and letting it sort the dependency tree rather than trying to merge it into an existing setup. The GitHub README gives you the basic install commands, but they skip a few steps that matter if you're working on a macOS machine. Specifically, the HDF5 dependency needs to be installed via Homebrew first, otherwise the h5py build will fail silently and waste an afternoon. After you get past that, run the test suite they provide. If it passes, you're ready to load data. If it fails, check your C compiler version — that's usually the culprit.

How the workflow actually works

The core pipeline takes raw sequencing reads or already-processed count matrices and runs them through quality control, normalization, dimensionality reduction, and clustering. On paper that sounds simple. In practice the QC step is where most projects stall out because the default parameters are tuned for specific datasets and don't generalize well. I hit a real problem with a batch of human peripheral blood mononuclear cells where the mitochondrial gene filtering was far too aggressive. The default threshold flagged over sixty percent of cells as low-quality, which made the dataset look unusable. I ended up manually inspecting the scatter plots of total UMI counts versus percent mitochondrial reads for each sample, then set per-sample thresholds based on the bimodal distribution I saw rather than using the global cutoff. That brought the retention rate up to about eighty-two percent and the downstream clustering became interpretable. The pipeline lets you override the QC filters by passing a dictionary of custom thresholds to the preprocess function, but finding that in the docs took me longer than it should have. After QC you normalize the counts and run the PCA or integration step depending on whether you're working with a single sample or multiple batches. The integration module handles batch correction across samples, which is critical if you're combining data from different donors or sequencing runs. It's based on reciprocal PCA alignment, and it works reasonably well when your batches aren't wildly different in composition. If you're mixing tissue types that have fundamentally different cell populations, the integration will force similarities that don't actually exist and you'll get misleading cluster assignments.

Get the Full Details

Eurobike 2025: Демпферы от Simons.Lab / Новое железо / Twentysix
Eurobike 2025: Демпферы от Simons.Lab / Новое железо / Twentysix

Things the documentation won't tell you

Most beginners miss how important the selection of highly variable genes is before you run dimensionality reduction. The default method picks genes based on dispersion, which works fine for heterogeneous tissue but systematically underperforms when you're studying a relatively homogeneous population like a cell line or a purified immune subset. In those cases the gene selection step ends up picking housekeeping genes as the top variables because the biological signal is subtle. I switched to using the model-based method for those scenarios and it consistently produced cleaner separation in the UMAP plots. Another thing nobody mentions upfront: memory usage scales badly with cell count. The pipeline can handle a few hundred thousand cells comfortably on a standard workstation, but once you push past a million, the clustering step becomes the bottleneck. I ran a dataset of roughly 1.4 million cells and the Leiden clustering step consumed over ninety gigabytes of RAM and took about four hours on a machine that handled the earlier steps without breaking a sweat. Splitting the dataset into smaller chunks, running clustering on each chunk, and then merging the labels at the end cut the runtime down to roughly forty-five minutes and kept memory under twenty gigabytes. It's not the intended workflow but it's the one that actually works at that scale.

When Simons Lab is the wrong tool

This pipeline is built for exploratory analysis of scRNA-seq and spatial data. If you're doing bulk RNA-seq differential expression, you're better off with DESeq2 or edgeR. If you're working with targeted panels or CITE-seq data where you have protein marker counts alongside transcript counts, the integration isn't seamless and you'll spend more time writing custom code than the tool saves you. For those cases dedicated pipelines like Seurat or Scanpy with their multimodal extensions handle the data structures more cleanly. There's also the question of reproducibility. Because the code lives on GitHub and gets updated frequently, a notebook that worked last month might break this month if a dependency changed. I keep a pinned version of the repository for each project and document exactly which commit I used. It adds a step to the workflow but it saves you from wondering why results changed after a routine system update. If you want to get started, head to the Simons Foundation website and look for the Lab section. The code repositories are linked from there along with a few case study notebooks that walk through complete analyses. Working through one of those end-to-end before applying it to your own data is probably the fastest way to figure out what's going on.