Getting Gene Expression Data Out of Raw Scan Files
Most people hit a wall within the first hour of processing microarray data. They download Affymetrix CEL files or Agilent feature files from a public repository, fire up whatever pipeline they found on GitHub, and wonder why their results look like noise. I've done this with my own datasets, usually after spending a day prepping RNA and running the hybridization, so you'd think the bioinformatics part would be straightforward. It isn't. The gap between a raw scan and a publishable expression matrix is where most projects die. Here's what the actual workflow looks like, not the textbook version. You start with image files from the scanner. Those get processed into raw intensity values through the manufacturer's algorithm. For Affymetrix, that means RMA or GCRMA normalization. For Agilent two-color arrays, you're doing within-array loess normalization followed by between-array quantile normalization. The output is a matrix of expression values, one row per probe set, one column per sample. Then you do quality control. Not the kind of QC where you glance at a boxplot and move on. I mean actual diagnostic plots: MA plots, array quality heatmaps, principal component analysis to check for batch effects, and probe-level diagnostics. You need to see whether a sample is an outlier before you commit to downstream analysis. I once spent three days troubleshooting weird clustering results, only to discover that one array had a bad wash step during processing. The PCA plot showed it immediately as a clear outlier, but I had ignored it initially because the raw intensities looked reasonable. That cost me a week of lost time.
After QC, you filter out low-expressed probes. A common mistake here is being too aggressive. If you filter out anything below a certain percentile across all samples, you might remove genuinely interesting low-abundance transcripts. I usually set the threshold based on the interquartile range of all probe sets, removing only those consistently at background levels across every single sample. This typically removes about 20 to 30 percent of probes depending on the platform. Next is differential expression. For simple two-group comparisons, limma is still the standard workhorse. It uses empirical Bayes moderation of the standard errors, which stabilizes variance estimates even with small sample sizes. The voom function in limma is useful when you want to incorporate mean-variance relationships properly. For more complex designs with multiple factors or batch variables, you encode them directly in the model matrix. Just make sure your design is full rank or you'll get singular fit warnings that silently break your analysis. Common software choices: Bioconductor packages in R remain the most widely used environment. affy and oligo handle Affymetrix data. limma handles the statistical modeling. For Agilent arrays, the limma package with marray or just reading the feature extraction output directly works fine. If you're working with older platforms like cDNA arrays, you'll likely need to use custom scripts or older Bioconductor versions.
Here's a practical tip that isn't obvious from the manuals. When you're dealing with batch effects, never ignore them in the normalization step. ComBat from the sva package is the standard correction method, but the key is to apply it correctly. You need to include the biological variable of interest in the mod argument, otherwise you'll accidentally remove the signal you're actually trying to detect. I've seen this mistake repeatedly in published papers where the batch-corrected data completely washed out the experimental condition because the correction model was misspecified. Another thing people get wrong is the interpretation of fold change values from microarrays. These are approximate. The dynamic range is limited compared to RNA-seq, usually around three orders of magnitude versus five or six. A reported two-fold change might actually be anywhere from 1.5 to 2.5 fold when you account for platform-specific bias. If you need precise quantification, especially for lowly expressed genes, consider validating with qPCR. I typically validate at least 10 to 15 percent of my differentially expressed genes this way, and the concordance rate is usually around 80 to 85 percent, which is acceptable for most studies but not perfect. The main limitation of Gene Expression Analysis Microarray is that you can only measure what you've pre-designed probes for. If your transcript isn't on the chip, you won't see it. There's no discovery component. RNA-seq has largely replaced microarrays for de novo studies, but microarrays still have a place in large cohort studies where cost per sample matters and you only need to measure known transcripts. A typical Affymetrix GeneChip costs roughly 200 to 400 dollars per sample for processing and scanning, compared to 50 to 150 dollars per sample for RNA-seq depending on depth. The throughput advantage of microarrays is real though. You can process 96 samples in a single run on most scanners, which would be impractical with RNA-seq library preparation.
Get the Full Details

For those looking for downloadable tools and pipelines, the Bioconductor project at bioconductor.org is the primary resource. The affy, oligo, limma, and sva packages cover the vast majority of standard analysis needs. There are also several Shiny apps and standalone GUI tools like Partek Flow and GeneSpring if you prefer point-and-click interfaces over coding. The open-source route through R gives you more flexibility and reproducibility, which matters when you need to revisit your analysis months later or share it with collaborators.
Practical Edge Case: Cross-Hybridization Artifacts
I ran into a specific problem last year where several probe sets showed extremely high variance across replicate samples, even though the technical replicates were processed identically. The coefficients of variation were in the 30 to 40 percent range, which is way outside normal expectations for well-prepared RNA. After digging into the probe sequences, I found that many of these problematic probe sets shared significant homology with pseudogenes or paralogous gene families. The probes were binding to multiple transcript variants, creating inconsistent signal intensities that varied between samples depending on their individual isoform expression profiles. The workaround was to switch from probe-set-level analysis to gene-level summarization using the latest annotation packages. Instead of relying on individual probe sets, I aggregated signals across all probes targeting a given gene symbol. This reduced the noise significantly, though it also meant losing some resolution on isoform-specific effects. For most downstream applications like pathway enrichment analysis, the gene-level approach produced cleaner and more biologically interpretable results. It's a reminder that the annotation data you're using matters as much as the raw analysis. Using outdated chip annotations can introduce systematic errors that no amount of statistical correction will fix.