Econometric policy evaluation: A critique
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
Can economists reliably predict what a new government policy will do by plugging it into a statistical model estimated from past data? This famous 1976 essay says no, and explains why in a way that changed how economics is done. Any model built from real people's behavior bakes in their expectations about the old policy; once the policy changes, people's behavior changes too, so the old model's numbers stop applying. Using consumption, investment, and inflation-unemployment examples, the paper argues this is not a technical glitch but a basic flaw in comparing policies with fixed-parameter models -- valid predictions require modeling the policy rule itself, not just its past effects.
What this paper finds — and why it matters
This 1976 Carnegie-Rochester Conference Series paper by Robert Lucas argues that the standard “theory of economic policy” – simulating a fixed, estimated econometric model under alternative hypothetical policies – is invalid for policy evaluation, because the model’s parameters are themselves the optimal decision rules of economic agents, and those decision rules change systematically whenever the policy regime they were estimated under changes. Lucas first documents that actual forecasting practice already departs from the textbook theory of economic policy – econometricians largely ignore pre-1947 data, frequently re-fit relationships (the wage-price sector being a running example), and revise intercepts based on recent runs of residuals – and shows that Cooley and Prescott’s “adaptive regression,” in which the parameter vector follows a random walk, matches this behavior and reconciles good short-term forecast accuracy with essentially unbounded variance for long-run policy simulations built on the same models. He then works through three canonical building blocks of large macroeconometric models – the Friedman-Muth permanent-income consumption function, a Jorgensonian tax-and-investment model in the style of Hall and Jorgenson, and an expectational Phillips curve – to show concretely that a policy change which is understood in advance by agents (e.g., an income tax surcharge announced as temporary, or an investment tax credit believed to be transitory rather than permanent) produces behavioral responses that a fixed-parameter extrapolation of the estimated relationship gets systematically wrong, sometimes by a factor of several times in the investment-credit case. The paper’s central syllogism, stated in the concluding section, is that because econometric relationships encode optimal decision rules that vary with the stochastic structure of the variables agents must forecast, “any change in policy will systematically alter the structure of econometric models,” so that comparisons of alternative policy rules using models estimated under a different, unstated policy regime are invalid regardless of how well those models fit historical data or forecast in the short run. Lucas’s proposed remedy is not to abandon policy evaluation but to model policy itself as a parameterized rule generating the forcing variables, so that the behavioral parameters become an estimable function of the policy parameters – feasible, he argues, only for policy changes that are openly discussed, understood, and expected to be enforced as stable rules, not for ad hoc or deliberately concealed interventions.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What is “the theory of economic policy” that the paper targets, and what does the paper claim is wrong with it?
The theory of economic policy, following Tinbergen, describes the economy by y_{t+1} = F(y_t, x_t, θ, e_t), where F is a fixed (if initially unknown) function, θ a fixed parameter vector to be estimated, x_t a vector of “arbitrary” policy/forcing variables, and e_t i.i.d. shocks; once F and θ are known, policy evaluation consists of specifying alternative sequences of x_t and computing or simulating the resulting moments of y_t, such as a “long-run Phillips curve” (Sec. 2, pp. 19-21). Lucas’s overarching claim, stated in the introduction, is that “simulations using these models can, in principle, provide no useful information as to the actual consequences of alternative economic policies” (Sec. 1, p. 20) – not because the models are badly estimated, but because the very assumption that (F, θ) is stable “under arbitrary changes in the behavior of the forcing sequence {x_t}” is theoretically unjustified once agents’ decision rules (which is what F and θ encode) depend on their own expectations about the behavior of x_t.
Q2. What evidence does Lucas offer that actual econometric forecasting practice already departs from the theory of economic policy?
Three signs (Sec. 3, pp. 21-24): (1) forecasters largely ignore pre-1947 annual data even though the theory says more observations, especially “extreme” ones, should sharpen estimates; (2) econometric relationships – the wage-price sector “now in progress” being a good example – are frequently and often substantially refitted, rather than showing the continuously improving precision within a fixed structure that the theory predicts; and (3) forecasters routinely revise intercept estimates based on recent runs of residuals, which “accounts… for the superiority of the actual Wharton forecasts as compared to forecasts based on the published version of the model.” Lucas stresses this is not a criticism of forecasters – adapting to new information is sensible – but evidence that “current forecasting practice is not conducted within the framework of the theory of economic policy,” so forecasting success provides no evidence for the theory’s reliability for policy simulation.
Q3. What is Cooley and Prescott’s “adaptive regression,” and how does Lucas use it?
Cooley and Prescott model θ as following a random walk, θ_{t+1} = θ_t + η_{t+1}, rather than being fixed; maximum-likelihood forecasting under this structure resembles exponential smoothing, closely matching actual econometric practice (Sec. 3, pp. 22-24). Lucas uses the model positively, “as an idealized model of the behavior of large-model forecasters,” arguing that if it is roughly accurate it reconciles two facts that otherwise seem to conflict: that current models forecast well and are improving, and that their long-run policy simulations are meaningless. “Under the adaptive structure, a small standard error of short-term forecasts is consistent with infinite variance of the long-term operating characteristics of the system” (Sec. 3, p. 24) – i.e., good tracking says nothing about the reliability of counterfactual policy simulations built on the same fitted structure.
Q4. Why does the consumption-function example show that “understood” policy changes break fixed-parameter forecasting?
Using the Friedman-Muth permanent-income model, in which permanent income is the discounted sum of expected future income (eq. 4) and the empirical consumption function (eq. 6) is derived under a specific stochastic process for income, Lucas shows that if a policy adds a supplement x_t to income from date T onward and the consumer knows this in advance, the standard method – insert forecasted income plus x_t into the fitted equation (6) – gives the wrong answer (Sec. 5.1, pp. 25-29). For a permanent increase x_t = ξ, the true consumption response from (1) and (4) is kξ, but the forecast from (6) reaches this only asymptotically; for an exponentially growing or stochastic supplement, the ratio of the (6)-based forecast to the true effect need not even tend to one as t → ∞, and “for any policy change which is understood in advance, extrapolation or simulation based on (6) yields an incorrect forecast, and what is more, a correctibly incorrect forecast” (p. 29).
Q5. What does the investment tax credit example add, quantitatively, about the size of the bias?
Building an explicit accelerator-with-taxes model of investment demand (eqs. 7-9), Lucas shows the effect of a tax credit switch depends critically on whether firms believe the new credit regime is permanent or transitory (Sec. 5.2, pp. 29-34). Hall and Jorgenson’s method implicitly assumes the credit is permanent; but if instead firms correctly expect a Markov-switching credit regime, comparing the case the credit is almost never offered but permanent once granted (q near 0, p near 1) against the case it is frequently imposed but always transitory (q near 1, p near 0), the ratio of investment-stimulus effects is about 7 using r=.14, δ=.15 (p. 33), and for a more realistic calibration – a credit “off” for an average of five years and “on” for one year – the true stimulus effect is roughly 4.5 times larger than the Hall-Jorgenson estimate would suggest (pp. 33-34). Lucas concludes that only policies generated by a fixed, publicly understood rule (his aside references Article I, Section 7 of the U.S. Constitution, i.e., the requirement that tax legislation be passed by Congress) permit this kind of correction; the effects of genuinely novel, ad hoc tax interventions cannot be accurately forecast even in principle by this method (p. 34).
Q6. How does the Phillips-curve example connect to the Phelps-Friedman natural-rate hypothesis?
Lucas constructs an island/incomplete-information model of aggregate supply (following his own 1972-73 work) in which cyclical supply responds to the gap between the actual local price and traders’ rational estimate of the aggregate price level, yielding y_t = θβ(p_t - p̄_t) + y^P_t (eq. 11), where θ is a signal-extraction parameter reflecting the ratio of aggregate to relative price variance (Sec. 5.3, pp. 34-38). If actual prices follow a random walk with drift π, the model implies a stable-looking empirical Phillips curve over any sample period with roughly constant π and variance, yet “it is evident… that a sustained increase in the inflation rate… will not affect real output” (p. 38) – because a change in the mean or variance of inflation changes θ itself, so the apparently stable trade-off is an artifact of the sample, not a stable structural relationship. He also shows a distributed-lag variant (autoregressive inflation) can produce a nonzero apparent “long-run” trade-off in-sample despite the same underlying invariance failure, warning that both configurations are observationally possible and neither licenses treating the fitted relationship as policy-invariant (pp. 38-39).
Q7. What is the paper’s central syllogism, and how general is its scope?
Stated explicitly in the concluding section: “given that the structure of an econometric model consists of optimal decision rules of economic agents, and that optimal decision rules vary systematically with changes in the structure of series relevant to the decision maker, it follows that any change in policy will systematically alter the structure of econometric models” (Sec. 7, p. 41). Lucas is careful to scope the claim: for short-term forecasting the conclusion is “of only occasional significance,” since a fixed-parameter model can still track well if drift is slow; but “for issues involving policy evaluation… it is fundamental,” since it implies that comparisons of alternative policy rules using current macroeconometric models are invalid “regardless of the performance of these models over the sample period or in ex ante short-term forecasting” (p. 41).
Q8. What alternative structure does Lucas propose for valid conditional forecasting, and what are its limits?
Lucas proposes writing policy and other disturbances as x_t = G(y_t, λ, η_t) for known G and fixed λ, so that the system becomes y_{t+1} = F(y_t, x_t, θ(λ), e_t) – i.e., policy is a parameterized rule, and the behavioral parameters θ are an estimable function θ(λ) of the policy parameters, rather than fixed under arbitrary policy (Sec. 6, pp. 39-40). The catch is epistemic, not just technical: “if the policy change occurs by a sequence of decisions following no discussed or pre-announced pattern, it will become known to agents only gradually… and the movement to a new θ(λ)… will be unsystematic, and econometrically unpredictable,” whereas “if… policy changes occur as fully discussed and understood changes in rules, there is some hope that the resulting structural changes can be forecast” (p. 40). Lucas is explicit that this preference for rule-based over discretionary policymaking is not based on any general optimality claim about rules (“There seems to be no theoretical argument ruling out the possibility that… delegating economic decision-making authority to some individual or group might not lead to superior… economic performance”); rather, only comparisons among alternative rules are, in principle, scientifically assessable (Sec. 6, p. 40).
Q9. Does the paper claim originality for these arguments, and to whom does it credit the underlying ideas?
No – Lucas’s second disclaimer states plainly: “There is little in this essay which is not implicit… in Friedman [1957], Muth [1961] and, still earlier, in Knight [1921]” (Sec. 1, p. 20), and the concluding section again credits the point of view “due originally to Knight and, in modern form, to Muth” (Sec. 6, p. 40). The consumption example explicitly builds on Friedman’s (1957) permanent-income hypothesis and Muth’s (1960, 1961) rational-expectations formalization of optimal forecasting of permanent income (Sec. 5.1, pp. 26-27); the paper is thus presented as tracing an existing, under-appreciated logic to its foundation and re-deriving its quantitative bite for specific, widely used applied models, rather than as introducing a wholly new idea.
Q10. What is the paper’s concluding position on rules versus discretion and democratic policymaking?
The essay’s final paragraph reframes the argument as, in part, an argument for transparency: “it appears that policy makers, if they wish to forecast the response of citizens, must take the latter into their confidence,” since only policies whose rule is publicly known and believed can, even in principle, be evaluated scientifically in advance (Sec. 7, p. 42). Lucas notes this conclusion, “if ill-suited to current econometric practice, seems to accord well with a preference for democratic decision making” – tying the paper’s technical argument about parameter invariance to a normative preference for openly rule-based rather than concealed, ad hoc policymaking, while explicitly declining to claim this preference follows from any proof that rules dominate discretion in general.
Key terms in this paper
Definitions below follow the paper's own usage.
- The theory of economic policy
- the standard framework the paper attacks (following Tinbergen): the economy is described by y_{t+1} = F(y_t, x_t, θ, e_t), with F known (or estimable), θ a fixed parameter vector, x_t a vector of policy/forcing variables treated as "arbitrary," and e_t i.i.d. shocks; policy evaluation consists of simulating this fixed structure under alternative hypothetical sequences of x_t and comparing the resulting moments of y_t.
- The central syllogism (policy-invariance failure)
- the paper's closing syllogism (Sec. 7): "given that the structure of an econometric model consists of optimal decision rules of economic agents, and that optimal decision rules vary systematically with changes in the structure of series relevant to the decision maker, it follows that any change in policy will systematically alter the structure of econometric models" -- i.e., the parameter vector θ is not policy-invariant, so simulating a fixed-θ model under a new policy rule yields no valid inference about that policy's actual consequences.
- Adaptive regression (parameter drift)
- Cooley and Prescott's proposed alternative to the fixed-parameter view, in which the parameter vector θ follows a random walk (θ_{t+1} = θ_t + η_{t+1}); Lucas uses it positively, as a descriptive model of how large-model forecasters actually behave (revising intercepts based on recent runs of residuals, frequent re-fitting), reconciling good short-term forecast accuracy with the paper's claim that long-run policy simulations from the same models are meaningless.
- Policy as parameters of a forcing rule
- the paper's proposed alternative structure for valid conditional forecasting (Sec. 6): policies and other disturbances are written as x_t = G(y_t, λ, η_t) for known G and fixed parameter λ, so that the remaining behavioral parameters become a function of the policy parameters, θ(λ), in y_{t+1} = F(y_t, x_t, θ(λ), e_t) -- treating a policy change as a change in λ (or in the rule G itself) whose effect on θ(λ) can, in principle, be estimated, rather than treating x_t as an arbitrary sequence overlaid on a fixed θ.