Why Your Effect Definitions Keep Getting Reviewed Back

You define the effect, run the regression, and the reviewer says your identification strategy doesn't actually recover what you claimed. This happens constantly because economists skip the definition step or define it too loosely. I spent three years chasing this problem in labor economics before it clicked. You bring me a dataset on a job training program, you want the employment effect, and everyone assumes that's straightforward. It isn't. Not even close.

Effect Definition Economics

The core principle is simple but easily ignored. Before any estimation touches your data, you need to write down exactly which causal quantity you are trying to recover. Not a paragraph of motivation. A mathematical definition with the population, the treatment, the outcome, and the counterfactual all explicitly stated. When you do this properly, you immediately see whether your data can support it. When you don't, you waste months on estimators that estimate something close but not what you said you wanted. The standard framework breaks down into three components. The treatment indicator D, which can be binary or continuous. The potential outcomes Y(1) and Y(0), representing what happens under each treatment state. And the parameter of interest, which might be the average treatment effect, the treatment effect on the treated, or something more specific like the local average treatment effect from Imbens and Angrist 1994. Here is where most people slip up. They define the effect as E[Y(1) - Y(0)] and call it a day. That is the ATE, yes, but it is an unconditional expectation over the entire population. In practice, your sample might only cover working-age adults in one metropolitan area, your treatment might be a policy that only reached certain sectors, and your data might only observe one of the two potential outcomes for each person. The ATE you estimated is not the ATE you claimed. It is something else entirely.

I ran into this explicitly in 2019 working on a minimum wage study. I wanted the effect of a wage increase on restaurant employment. I defined it as the ATE. My data came from county-level census microdata, which meant I was actually estimating something closer to a weighted average of effects across counties with different baseline wages and labor market structures. The reviewer caught this immediately. The fix was not to re-estimate. The fix was to redefine the target parameter to match what the data actually supported, which in that case was the effect among counties where the policy variation had meaningful pre-post differences. That narrowed the scope considerably but made the whole analysis defensible. The workaround I use now is to write the effect definition on a separate page before touching Stata or R. I specify the population S, the treatment rule d(x), the outcome function y(d), and then I write out the estimator I plan to use and prove in a few lines that it converges to the parameter I claimed. If the proof does not work, I change the definition, not the estimator. Most people do the reverse and wonder why their results look weird. Another common trap involves heterogeneous effects. People love to report a single number and call it the effect. But if the treatment effect varies across subpopulations, the single number is meaningless without knowing which subpopulation it applies to. This is especially damaging in policy work where the treated group and the untreated group have systematically different response patterns. Rosenbaum and Rubin made this point clearly back in 1983, and it still gets ignored regularly.

Get the Full Details

What is the Income Effect in Economics? - FlyingMachineArena
What is the Income Effect in Economics? - FlyingMachineArena

The LATE framework from instrumentally variable approaches helps here, but it introduces its own definitional problem. The local average treatment effect only applies to compliers, and compliers are not randomly selected. They are the subset of the population whose treatment choice actually responds to the instrument. In my experience, about 30 to 40 percent of published IV papers in applied micro either misstate what LATE means or treat it as if it were the ATE without justification. It is not. Continuous treatments add another layer. The standard binary framework does not apply cleanly. You need a functional form assumption or a generalized method of moments approach, and the effect you define becomes path-dependent. If you define the effect as the derivative of the outcome with respect to treatment at the mean, you get one number. If you define it as the average change from zero to the observed level, you get another. Both are valid definitions. They just answer different questions. There are also edge cases where the effect definition collapses entirely. If your treatment is sticky and there is no variation in a reasonable timeframe, you cannot identify any causal effect no matter how sophisticated your estimator. I saw this in a study of educational reform where the policy had been in place for fifteen years and no comparable control existed. The proposed difference-in-differences design was fundamentally unidentifiable. The only honest move was to frame it as a descriptive exercise or find a completely different identification strategy, like regression discontinuity if there was a cutoff in implementation timing.

So the practical checklist is this. State your target parameter with explicit notation. Verify your data actually contains the variation needed to identify it. Check whether your sample matches your population. Make sure your estimator recovers the parameter you defined, not a close approximation. If any of those fail, go back to step one and redefine. This approach usually cuts the iteration time down from weeks to a couple of days because you stop re-estimating the same model with different controls and start fixing the actual definition. Most of the problems in empirical economics are not estimation problems. They are definition problems wearing estimation costumes.