Working with Run2 Data from the Vera Rubin Observatory

Run2 is the second operational phase of the Vera Rubin Observatory's Legacy Survey of Space and Time. It's where the telescope moved past initial commissioning and started producing real scientific-grade imaging data at scale. If you're trying to figure out how to actually use that data, here's what you need to know. The survey is designed to image the southern sky repeatedly over ten years. Run2 data includes the first-years of that public survey, captured through the camera's thirty-four individual CCDs across multiple filters (g, r, i, z, y). Each visit produces two co-added images — a single-visit coadd and a stacking of multiple visits. The raw data is huge. A single visit can generate terabytes of uncalibrated data before reduction, and the processed products are similarly large when you pull in the full catalog. One thing people don't always grasp is the data structure. The Rubin Science Platform organizes Run2 collections with a specific schema — collections for each filter, each visit type, and separate branches for raw, calibrate, and coadd pipelines. If you try to query without understanding the naming convention, you'll spend more time debugging than actually analyzing. The collection names follow a pattern like rubin/S24A/run2.7w/y1_run2.7 or similar strings that encode the year, season, and run designation.

Accessing the Data

You can access Run2 data through the Rubin Science Platform at platform.sdp.slac.stanford.edu. They host both interactive notebooks and command-line tools. The primary interface is via their hosted notebooks in Google Colab, which come pre-loaded with the necessary dependencies. You don't need to install anything locally unless you want to work offline or process data at a scale the hosted environment can't handle. For local work, they provide a Python package called lsst_stack, which wraps all the underlying pipelines. Getting lsst_stack set up is genuinely painful unless you're already comfortable with the DM stack environment. I ran into a specific problem last year where my colima virtual machine couldn't handle the memory footprint of loading even a modest Run2 coadd. The default VM settings cap out around 8 gigabytes, and a single tract's worth of processed data plus catalogs will eat that in seconds. I switched to working entirely on the hosted platform after that, using their notebooks with larger compute instances. It's slower for interactive exploration but doesn't crash.

The Query Problem Nobody Warns About

Here's the thing that catches most people off guard: querying Run2 data is not fast. The data volumes are so large that a naive cone search or a poorly constrained table query can hang for minutes or even timeout entirely. The trick is to always constrain your queries by collection name, date range, and sky footprint. Never fire off a broad cross-schema query. I learned that the hard way when a script I wrote tried to join three different catalog tables across the full Run2 dataset and consumed forty-five minutes of compute time for results I could have gotten in thirty seconds with proper filtering. Use the database's built-in indices. The Rubin Science Platform's SQL engine supports filtering on coordinates, magnitude limits, and observation dates. Build your query with those filters first, then join tables on the constrained result set. The difference in execution time is usually an order of magnitude or more.

Common Pitfalls

Coadd variants matter. There are different types of coadds — object-based, forced-photometry, and raw single-visit stacks. They serve different purposes. If you're doing photometric variability analysis, you want single-visit data. If you're building a deep source catalog, you want the object coadd. Pulling the wrong variant and not realizing it will mess up your entire analysis pipeline. Bias correction is baked in, but not always where you expect. The raw images go through a pretty aggressive calibration pipeline before they reach you. Flux offsets, flat-fielding, and background subtraction are all applied. This is good — it means less work on your end — but it also means you need to understand what the pipeline did to your data before you trust any measurement. I once ran a color-magnitude diagram and got results that looked plausible until I checked the filter transmission curves against what the pipeline had applied. A small but significant shift in the effective wavelength meant my colors were systematically off by about 0.02 magnitudes. Fixing that required re-applying the correct synthetic photometry corrections on top of the pipeline output. Data completeness varies by filter and sky region. Run2 hasn't covered the entire survey footprint equally. Some areas have deep multi-visit coadds while adjacent regions are still shallow. Before you trust any null detection or limit calculation, verify your sky patch has sufficient coverage. The platform provides visit count maps you can overlay on your target region.

When Run2 Data Isn't the Right Tool

There are cases where Run2 data simply won't work for you. If you need extremely high-cadence monitoring — say, tracking variable stars on timescales of hours over weeks — this survey isn't built for that. It revisits the same patch roughly every three nights, which is fine for longer-period variables but useless for short-timescale phenomena. In those cases, you'd be better off combining it with data from TESS or ZTF, or looking at other facilities entirely. Similarly, if you're working in the northern sky, Run2 coverage is basically nonexistent. The observatory is located in Chile and only sees the southern hemisphere. You'd be looking at Northern Sky surveys for anything above declination zero.