What This Actually Measures
When demographers talk about unions between spouses from different social categories, they are looking at dissimilarity in education, income bracket, parental class, religion, or ethnicity. The standard metric is the index of dissimilarity, often paired with a cross-tabulation of one partner's attribute against the other's. If you are running this analysis yourself, the first thing you need is a clean couple-matched dataset. Most people skip that step and end up with noisy results. The core workflow is straightforward but easily broken. You take two variables, one for each spouse, collapse them into ordered categories, and then build a matrix. The diagonal shows like-with-like pairings. Everything off-diagonal is cross-category. The ratio of off-diagonal to total gives you your mixing index. In practice I use a 5 by 5 table for education and a 4 by 4 table for occupational class. Smaller tables stabilize the numbers. Larger tables produce empty cells and meaningless fractions. I worked on a project a few years ago where we matched respondents from two different census microdata files using names and dates of birth. About twelve percent of records failed to match because one spouse was listed under a maiden name in one file and a married name in the other. The workaround was simple but tedious. I ran a phonetic name hash, then allowed a single character edit distance, and finally cross-referenced the last four digits of their social security numbers. That reduced the mismatch rate to under two percent. Without that step, the off-diagonal counts would have been inflated by random pairings.
Here is what most people miss. Homogamy is not the default. Even in societies with high educational attainment, college graduates still marry each other at rates well above random chance. But the pattern flips for religion in secular populations. When both partners are non-religious, the cross-category signal disappears because the category itself has dissolved. You will see a spike in apparent mixing, but it is an artifact of a collapsed baseline. Always check your base rates before interpreting a high dissimilarity score. Another nuance that trips people up is the difference between absolute and relative mobility. A union where one partner moved up two education levels and the other moved down one is not the same as one where both moved sideways but started from different points. Most published studies treat all off-diagonal cells equally. They are not. The social meaning of a nurse marrying a plumber is not the same as a CEO marrying a social worker, even if the statistical distance is identical. I add a directional weight to the matrix when the research question demands it. It takes about twenty minutes to code and saves you from misleading claims later. If you want to replicate this, the data sources are the American Community Survey for the United States, the British Household Panel Survey for the UK, and the EU-SILC for Europe. Each has spousal linkage built in, but the variable definitions differ. ACS calls it "marital status" with a spouse subcode. BHPS uses the family module. EU-SILC uses the HH090 variable. Align them manually. Do not assume the labels match across datasets.
The main bottleneck is sample size. Once you split by education and occupation together, many cells drop below fifty observations. That is too small for stable estimates. My fix is to pool adjacent categories rather than drop them. Combine upper intermediate with routine occupations. Combine bachelor and master degrees into a single tertiary bucket. You lose some granularity, but the estimates become usable. I typically report both the pooled and unpooled versions so readers can judge the sensitivity. There is also a gender asymmetry that standard models ignore. Women still tend to marry up or stay level in education more often than men do. Men show a stronger pattern of marrying down. If you analyze the data without separating by gender, these opposite forces cancel out and produce a deceptively flat result. Run the analysis separately for male and female respondents, then compare. The difference is usually significant and often larger than the raw mixing index itself. The biggest limitation of this approach is that it captures structure, not meaning. A high cross-category rate does not tell you whether these unions face friction or whether they function normally. For that you need survey data on relationship satisfaction, household decision making, or subjective class identity. The matrix alone cannot answer those questions. Combine it with qualitative responses or at least a basic satisfaction scale from the same survey wave. Otherwise you are describing patterns without explaining anything.
Get the Full Details
