Working Through Correspondence Analysis Without Losing Your Mind

I spent way too many hours debugging row profiles that refused to behave themselves before I actually understood why. Correspondence analysis is one of those techniques that looks straightforward when you read the textbook but turns into a mess the moment you try to apply it to real data. The Greenacre book from the Wiley series is the standard reference, sure, but it assumes you already know where the bodies are buried. Here is how I actually use it. The basic idea is simple enough. You take a contingency table or any two-way frequency matrix, decompose it using singular value decomposition, and plot the row and column points in a low-dimensional space. Distances between points approximate the chi-squared distances from the original table. That is the theory. The practice involves a dozen decisions before you even get to the plot. The first thing I do is check whether my data is actually suitable. CA works on counts or frequencies. If you give it proportions without converting them back to the underlying mass structure, the inertia calculations will be wrong and your plot will look reasonable but mean nothing. I once ran a correspondence analysis on a market segmentation table that was already weighted by population density. The row points were all clustered near the origin and the columns were spread out weirdly. Took me three days to realize the weighting had distorted the row masses. I had to un-weight the table, run the analysis, then map the original weights onto the plot as point sizes afterward. That workaround usually saves me from having to redo the whole thing.

Another thing nobody tells you: the choice between row canonical and column canonical representation matters more than the eigenvalues. Most software defaults to the principal form, which is fine for visualization, but if you are doing downstream regression on the factor scores, you want the factorial form. I learned this the hard way when a client asked me to use the CA scores as predictors in a logistic model and the results were nonsensical. The principal and factorial forms give different coordinate scales. Switching to factorial form fixed it immediately.

What Actually Goes Wrong

Sparse tables are the biggest problem. When you have lots of zeros in your contingency table, the chi-squared distances become unstable and the singular value decomposition can produce garbage components that look meaningful on the plot but are just noise. I typically apply a small continuity correction or merge rare categories before running the analysis. How much you merge depends on your table. A rule of thumb is that any cell with an expected frequency below five is worth investigating. If you have more than twenty percent of cells in that range, consider whether CA is the right tool at all. Interpreting the dimensions is another minefield. The first dimension usually captures the strongest association, which is almost always the main effect. The second dimension picks up the next independent pattern. But when your table has three or four strong categorical effects interacting, the dimensions can become mixed and hard to label. I stop trying to name every dimension and instead focus on which points are far apart along each axis. The labels come later if anyone asks for them.

Get the Full Details

Correspondence Analysis: Theory, Practice and New Strategies. By E. Beh ...
Correspondence Analysis: Theory, Practice and New Strategies. By E. Beh ...

Practical Workflow

I usually work in R using the ca package or FactoMineR. The pipeline is: load your contingency table, verify the row and column masses add up to one, run the decomposition, check the explained inertia for each dimension, plot with both rows and columns, and then validate by looking at the contribution of each point to each dimension. The contribution metric tells you which cells are driving the structure, which is often more useful than the coordinates themselves. One practical detail: the scree plot of eigenvalues is not always decisive. In correspondence analysis the eigenvalues represent inertia, not variance in the traditional sense, so they do not follow the same rules as PCA. I look at the cumulative inertia and stop when adding another dimension contributes less than two percent. Usually two dimensions give you about sixty to seventy-five percent of the total inertia for a well-structured table. If you are getting below fifty percent in two dimensions, the table is probably too noisy or too high-dimensional for CA to handle cleanly.

When It Fails Completely

Correspondence analysis breaks down when your data is not a proper contingency table. I see people feed it similarity matrices, correlation matrices, or standardized residuals and wonder why the results look arbitrary. CA requires non-negative data with a clear row-column frequency structure. If your table has negative values from any kind of adjustment, the whole decomposition becomes invalid. There is no fix other than going back to the raw counts. It also does not handle panel data or repeated measures well. If you have the same categories observed across multiple time periods and you want to track movement, standard CA will mix the time dimension into the spatial structure in ways that are hard to separate. You need multiple correspondence analysis or a dynamic variant for that, and neither is as clean as the standard approach. If your table has more than about fifty rows or columns, the plot becomes unreadable regardless of what you do. I down-select to the most informative subset based on point contributions before plotting. Running the full analysis on all points, then filtering the plot to show only those contributing more than a threshold amount to at least one dimension, usually produces something that is actually legible within ten minutes of work.

The Greenacre reference remains useful for the theoretical grounding, but the practical details about data preparation, mass handling, and interpretation pitfalls are mostly learned through trial and error. The field has not changed dramatically since that book came out, which means the core method is stable but also means the common mistakes are the same ones people keep making.

Categorical Data Analysis (Wiley Series in Probability and Statistics ...
Categorical Data Analysis (Wiley Series in Probability and Statistics ...