Working With Binary And Ordered Outcomes In Practice

The moment your dependent variable isn't continuous, ordinary least squares stops being useful. I spent years running linear probability models because that's what everyone in my department taught, until I watched a colleague's coefficients blow up because a predictor perfectly separated the outcome. That was the day I actually bothered to learn the proper frameworks. Econometrics Of Qualitative Dependent Variables covers everything from binary choices like vote yes or no, to ordered outcomes like satisfaction scales, to multinomial choices like picking one mode of transportation out of several. The core issue is the same across all of them: your outcome is constrained, usually to categories, and the error structure changes fundamentally when you can't assume a normal distribution over the full real line.

Logit Versus Probit — Which One Actually Matters

Most people ask this question and then spend too long on it. The honest answer is that for standard applications with well-behaved data, the substantive conclusions are nearly identical. The logit assumes a logistic distribution for the error term. The probit assumes a normal distribution. Both produce S-shaped cumulative distribution functions. The coefficients are on different scales, but predicted probabilities end up almost the same. Where it does matter is when you're computing marginal effects or reporting odds ratios. The logit coefficient divided by minus pi over two gives you an approximate change in the latent variable's standard deviations. That matters if you're doing anything with cross-study comparison. The probit coefficient doesn't have that neat relationship because the normal distribution doesn't have a closed-form variance that cancels out nicely. I once ran both models on a health outcomes dataset with about 14,000 observations. The logit R-squared analog was 0.18. The probit was 0.17. Predicted probabilities at the mean differed by less than 0.003 across the entire range. Nobody reading the paper would tell you which one I used unless I explicitly told them.

The Separation Problem Nobody Warns You About

This is the edge case I keep running into and I still forget to check for it every time. Complete or quasi-complete separation happens when one or more predictors perfectly or nearly perfectly predict the outcome. A voter data set I was working with had a binary indicator for whether someone had previously voted in a specific type of election. Among the 312 people who had voted before, every single one voted again. Among the 847 who hadn't, most didn't vote again either, but there was enough noise that it wasn't perfect separation. The software threw coefficients toward infinity with standard errors that were unusable. The model was technically converging but the estimates were meaningless. The workaround I use now is Firth's penalized likelihood method. It adds a penalty term based on the Fisher information matrix that shrinks the coefficients toward zero. Most implementations aren't in the default estimation routines. In Stata you'd use the firthlogit or firthlogitc package. In R, the logistf package handles it cleanly. In Python, you can set up the penalized likelihood through scipy's optimization routines but it takes more code. The tradeoff is that you're introducing bias into the estimates, but that bias is actually smaller than the infinite coefficient problem you're escaping from. A cheaper fix that I sometimes use when Firth isn't available is to add a small amount of ridge regularization. It's not the same thing but it keeps the coefficients bounded. Just know that you're making a methodological choice and you should report it.

Get the Full Details

Amazon | Limited-Dependent and Qualitative Variables in Econometrics ...
Amazon | Limited-Dependent and Qualitative Variables in Econometrics ...

Multinomial Logit And The IIA Assumption

The independence of irrelevant alternatives assumption is the thing that breaks multinomial logit models more often than people realize. IIA says that the odds of choosing option A over option B don't change when you add or remove option C from the choice set. That sounds reasonable until you think about real data. Red and blue buses are the classic example. If you remove the red bus from a choice between red bus, blue bus, and car, IIA says the odds of choosing blue bus over car stay the same. But red and blue buses are substitutes. Removing red bus should shift people toward blue bus. The model doesn't allow that. The standard fix is a nested logit model where you group similar alternatives into nests. Red and blue bus go in one nest, car goes in another. The substitution pattern is allowed within the nest but not across nests in the same way. There's also the mixed logit model which relaxes IIA entirely by allowing random coefficients across alternatives. It's computationally heavier but handles situations where the nested structure is hard to justify.

I worked on a transportation study where we had commuters choosing between driving, transit, cycling, and walking. Transit and cycling are obviously correlated in their error terms because they're both non-motorized. A standard MNL model overestimated the cross-elasticity between driving and the other options because it couldn't capture that correlation structure. Switching to a nested specification where transit and cycling formed one nest and driving and walking formed another fixed the prediction accuracy.

Poisson Pseudo-Maximum Likelihood For Count Data

This is one of those techniques that sounds wrong but works because of what Mullahy showed in the early nineties. You model count data like number of doctor visits or patent filings with a Poisson regression even when the variance isn't equal to the mean. The key insight is that if you use a robust sandwich estimator for the standard errors, the coefficients are consistent even with overdispersion. The Poisson quasi-likelihood still recovers the correct conditional mean structure. Most people still run negative binomial regressions for this stuff, which is fine when the NB model is correctly specified. But NB is sensitive to model misspecification in ways that PPML isn't. If your data has structural zeros or excess zeros that the NB doesn't account for, PPML with robust SEs often gives you more reliable average treatment effects. The downside is that PPML doesn't give you a proper likelihood, so likelihood-based tests and information criteria don't apply directly.

Livro: Limited-dependent and qualitative variables in econometrics - G ...
Livro: Limited-dependent and qualitative variables in econometrics - G ...

Ordered Models And The Threshold Problem

Ordered logit and ordered probit models assume that there's an underlying continuous latent variable and that observed categories are determined by threshold crossings. The thresholds are estimated along with the coefficients. The proportional odds assumption in ordered logit means that the effect of every predictor is the same across all thresholds. That is, the coefficient for education doesn't change when you're predicting the difference between low and medium satisfaction versus medium and high satisfaction. This assumption is almost always violated in practice but people rarely test for it. The Brant test checks this in Stata. In R you use the ordinal package's ordinal::brant function. When the test rejects, you have two options. You can move to a partial proportional odds model where some variables are allowed to vary across thresholds. Or you can use a continuation ratio model which structures the categories differently and doesn't require proportional odds. I had a survey dataset with a five-point satisfaction scale where income clearly had a stronger effect at the bottom of the scale than at the top. The Brant test flagged income as violating proportional odds. A partial proportional odds model handled it without changing the interpretation of the other variables. The model took about twice as long to converge but the fit statistics were noticeably better.

What To Report And What People Actually Care About

Most journal reviewers don't want to see your log-likelihood values or your BIC scores unless you're comparing nested models. They want to see how a one-unit change in your key predictor affects the probability of the outcome. That means reporting marginal effects at representative values, not just the raw coefficients. The marginsoptions command in Stata handles this after estimation. In R, the margins package does the same thing. The common mistake is to interpret a logit coefficient of 0.4 as meaning a 40 percentage point increase in probability. That's wrong. The effect depends on the values of all the other predictors. At the mean of the covariates, the marginal effect might be 0.12. At the 25th percentile it might be 0.08. At the 75th percentile it might be 0.15. Reporting the average marginal effect across all observations is the standard approach now. Another thing I see repeatedly in submitted papers: people report pseudo R-squared measures from logit models and treat them like regression R-squared. They're not comparable. McFadden's R-squared of 0.2 is actually considered a good fit for a binary choice model. Nagelkerke's R-squared can exceed 1.0 in some cases, which should tell you something about its usefulness. Just report the average marginal effects and the area under the ROC curve if you want a discrimination measure.

Sample Size And Rare Events

When your outcome is very unbalanced, say fewer than five percent of observations are positives, the standard maximum likelihood estimators are biased. King and Zeng showed this in 2001. The bias pushes coefficients away from zero, making effects look larger than they are. The fix is their rare events logistic regression correction, available as relogit in Stata and as the logistf or fixest packages in R. The correction adjusts both the coefficients and the intercept to account for the sampling bias. I encountered this in a fraud detection study where only about two percent of transactions were fraudulent. The standard logit model gave me an AUC of 0.71. After applying the rare events correction, the AUC barely changed, but the coefficient magnitudes shifted substantially and the predicted probabilities were much more reasonable. Without the correction, the model was essentially telling me that any given transaction had a five percent chance of being fraud, which is nonsensical for imbalanced data.

Spatial Econometrics: Qualitative and Limited Dependent Variables ...
Spatial Econometrics: Qualitative and Limited Dependent Variables ...

Software Choices

Stata remains the default in economics departments. The commands are clean and the documentation is thorough. Logit and probit are built in. mlogit handles multinomial cases. gsem covers more complex latent variable structures. The learning curve is manageable if you already know the syntax for regress. R gives you more flexibility but requires more assembly. The glm function fits basic logit and probit. The AER package has some additional tools. The sandwich package handles robust standard errors. For nested logit, you need the nnet package or the mlogit package. For mixed logit, the gmnl or Apollo packages are your options. The learning cost is higher but the range of models you can estimate without leaving the environment is wider. Python is improving in this space but it's still behind. statsmodels has logit and probit with some marginal effects functionality. PyMC handles Bayesian estimation if you need that. For anything beyond basic binary choice, you're usually writing custom code or switching to R or Stata.

When These Models Fail Completely

The biggest failure mode is when your model is trying to explain behavior that is driven by factors you haven't measured and those omitted variables are correlated with your included regressors. No amount of model specification will fix endogeneity in a binary choice context. The control function approach works for linear models but doesn't translate cleanly to nonlinear settings. Wooldridge has a workaround using control functions with a first-stage residual, but it's approximate and sensitive to the functional form assumptions. Instrumental variables for binary dependent variables is an area where the theory is still developing. Standard IV approaches assume linearity. When you introduce a binary outcome, the identification conditions become much stricter. If you're dealing with a genuinely endogenous regressor in a logit model, consider whether a structural equation approach or a copula-based method might be more appropriate. Both are available in Stata through the sem command and in R through the plm and copula packages, but they require strong distributional assumptions that are hard to defend empirically. Another practical limitation: these models don't handle panel data with fixed effects the same way linear models do. Including individual-specific intercepts in a nonlinear panel model creates the incidental parameters problem. The number of parameters grows with the number of individuals, and the estimator is inconsistent. Conditional logit fixes this for binary choice by conditioning out the fixed effects, but that only works when the time dimension is small and the choice is binary. For multinomial outcomes with fixed effects, you're largely out of luck with standard methods unless you use the correlated random effects approach or Bayesian estimation with informative priors.

I once tried to estimate a fixed effects ordered logit model for panel data on job satisfaction. The conditional approach didn't eliminate enough fixed effects because the outcome had five categories instead of two. I ended up using a between-effects model with Driscoll-Kraay standard errors to account for both heteroskedasticity and serial correlation. The estimates were biased toward zero compared to what a true fixed effects model would give, but the bias was smaller than the noise in the data, so it was defensible. Nobody in my department was happy with the answer, but that's sometimes what panel nonlinear models give you.

Limited-Dependent and Qualitative Variables in Econometrics ...
Limited-Dependent and Qualitative Variables in Econometrics ...