Getting CFA Working in Stata Without Losing Your Mind

Stata has a semimetric command called sem that can run confirmatory factor analysis, but it doesn't have a single "run CFA" button like some other packages. You build the model by writing out the measurement equations. It's not hard, but the syntax will bite you if you're not paying attention. Here's how I actually set one up. You start with your latent variable names and observed indicators, then you specify the factor loadings and error variances as parameters. The basic structure looks like this: That tells Stata that F1 is a latent factor measured by four observed variables. By default sem estimates factor loadings and error variances. You get fit indices automatically - chi-square, CFI, TLI, RMSEA. The output is straightforward enough once you know what you're looking at.

I should mention something most tutorials skip. Stata's sem command uses full information maximum likelihood by default, which handles missing data better than listwise deletion. But if you have a lot of missingness - I'm talking above 20 percent - your standard errors get wobbly. I learned this the hard way on a survey dataset where about 28 percent of responses were missing on two key indicators. The model converged but the parameter estimates were drifting depending on how many cases dropped out. Switching to ml estimation with the technique(bfgs) option stabilized things, but honestly the better move was just recoding the missing values as a separate category and moving on. Your call depending on what you're publishing. Another thing nobody warns you about: identification. With a single-factor model, Stata automatically fixes the first loading to 1 for scale identification. That's fine. But if you try to estimate a model with correlated factors and only one indicator per factor, you will get an identification error every single time. You need at least two indicators per latent variable unless you fix something else. I ran into this when trying to squeeze a theoretical model down to one observed variable per construct. Had to go back and add a second indicator or restructure the whole thing. For model fit interpretation, the standard cutoffs still apply. RMSEA below 0.06, CFI and TLI above 0.95. But with large samples - and I mean N over 500 - the chi-square test will reject your model even when the other indices look fine. Don't treat a significant chi-square as a death sentence. Check the close-fit indices and move on. That's just how asymptotic tests work with big data.

If you want to constrain loadings to be equal across groups - say testing measurement invariance between two demographics - you use the group() option. Set the first group's parameters free, then constrain the second group. It's a bit clunky compared to running separate models and comparing them manually, but it works and saves you from exporting to another program. The command reference is in Stata's documentation under [SEM] sem. It's thorough but dense. I keep it open while I'm building models and search for the specific syntax I need rather than reading it cover to cover. Takes some getting used to, but once you have a template you can adapt, fitting CFA models in Stata becomes a matter of minutes instead of hours.

Get the Full Details

Confirmatory factor analysis demo using STATA SEM builder (2018) - YouTube
Confirmatory factor analysis demo using STATA SEM builder (2018) - YouTube