Macro Paper Warehouse
Published Classic [American Economic Journal: Macroeconomics] doi:10.1257/mac.1.1.242 Vol. 1, No. 1, pp. 242-266

New Keynesian Models: Not Yet Useful for Policy Analysis

V. V. Chari

Patrick J. Kehoe

Ellen R. McGrattan

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Can the sticky-price models central banks rely on be trusted for quantitative policy advice? This 2009 critique says not yet. Its core point is that one pattern in the data can come from two rival stories with opposite implications — one where the disturbance is a distortion the government should offset, another where it is an efficient shift the government should leave alone. Four of the seven disturbances in a leading estimated model fail that interpretability test, yet they account for most of the movement in inflation and hours, and one implies swings in bargaining power thousands of percent wide. That matters because policy counterfactuals inherit whichever story was assumed.

What this paper finds — and why it matters

This 2009 American Economic Journal: Macroeconomics paper by V. V. Chari, Patrick Kehoe, and Ellen McGrattan is a theoretical and methodological critique rather than an empirical study: it asks whether state-of-the-art New Keynesian DSGE models, as typified by Smets and Wouters (2007), are yet reliable enough for quantitative policy analysis, and argues they are not. Building on the “business cycle accounting” framework of Chari, Kehoe, and McGrattan (2007) – which decomposes aggregate fluctuations into an efficiency wedge, a labor wedge, an investment wedge, and a government-consumption wedge that together capture essentially all the movement in US output – the authors first show that the same labor wedge can be generated by two observationally equivalent structural models with opposite policy implications: one in which the wedge reflects fluctuating union monopoly power over wages (a “bad” shock the government should offset, so that “relentless union busting is optimal”), and one in which it reflects fluctuating utility of leisure (a “good,” efficient shock, so that laissez-faire is optimal). They argue a model is useful for policy only if its shocks are both invariant to the policy interventions being evaluated and interpretable as “good” or “bad” in this sense, and that four of the seven shocks in the Smets-Wouters model – the wage markup, price markup, exogenous spending, and risk premium shocks – fail this test and are therefore “dubiously structural.” These four shocks are far from a side issue: per the paper’s own forecast-error variance decomposition (their Table 1), they account for 39.6/53.9/86.9 percent of the variance of output/hours/inflation at a 4-quarter horizon, rising to 60.3/86.0/88.0 percent at a 1,000-quarter horizon. Taken at face value, the wage markup shock implies a standard deviation of the wage markup of 2,587 percent – “several orders of magnitude outside of a reasonable range” if interpreted literally as variation in workers’ elasticity of substitution; the exogenous spending shock, defined residually from the national income identity and including net exports, has 3.5 times the variance of measured US government spending; and the risk premium shock has more than six times the variance of short-term nominal rates, which the authors argue is best read as a flight-to-quality shock – and therefore plainly not invariant to monetary policy. The paper separately criticizes the model’s backward price-indexation assumption, used to generate inflation persistence, as inconsistent with microeconomic price-duration evidence (Bils and Klenow 2004 report about 4 months between price changes; Nakamura and Steinsson 2008 report about 11 months; backward indexation instead implies every price changes every period), and questions the standard Taylor-rule specification as hard to reconcile with the smooth, trending behavior of long-term interest rates implied by the expectations hypothesis. The authors are explicit that the critique targets the Smets-Wouters (2007) implementation specifically and does not claim New Keynesian models are wrong – only that they are “not yet useful” for the kind of quarter-to-quarter policy counterfactual exercises for which they are increasingly used – and they note that, despite the critique, New Keynesian and neoclassical economists have in practice converged on similar broad policy recommendations (commitment to rules, low average inflation).

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What, in the authors’ own framework, makes a model’s shocks “structural” enough to support policy analysis?

A shock is usable for policy analysis only if it satisfies two properties: it must be invariant to the class of policy interventions being evaluated, and it must be interpretable as either a “good” shock (one the government should accommodate) or a “bad” shock (one it should try to offset). (Section I.A, pp. 246-247) If a shock fails either test – if it would itself change under the policy being contemplated, or if economists cannot agree whether it reflects an efficient disturbance or a distortion – then simulating the model’s response to a policy change is not actually informative about what would happen under that policy, because the “shock” is not a fixed, policy-invariant object.

Q2. What is the “business cycle accounting” framework this paper builds on, and what does it establish about US fluctuations?

Chari, Kehoe, and McGrattan’s (2007) business cycle accounting framework represents any economy’s aggregate data as if generated by a stand-in neoclassical growth model hit by four time-varying “wedges” – an efficiency wedge (acting like a TFP shock), a labor wedge (acting like a tax on labor income), an investment wedge (acting like a tax on capital accumulation), and a government-consumption wedge – and CKM (2007) show these four wedges together capture essentially all the movement in US output, with the labor wedge in particular central to accounting for US labor movements in both the Great Depression and postwar business cycles. (Section I, pp. 246-248; Figures 1A-1B) Critically, a “wedge” is a purely reduced-form, measured object: by construction it is consistent with many different underlying structural stories, and the accounting exercise itself is silent on which story is correct.

Q3. How does the paper’s “two models, one wedge” result illustrate the danger of treating a wedge as if it were structural?

The paper presents two full structural models that generate the identical labor wedge process yet imply opposite policy prescriptions: in a model with fluctuating government limits on union monopoly power (Proposition 1), the wedge is an inefficient distortion and the optimal policy is “relentless union busting”; in a model with an exogenously fluctuating utility of leisure (Proposition 2), the identical wedge is an efficient response to preference shifts and the optimal policy is laissez-faire. (Section I.B, pp. 248-253) Because the two models are observationally equivalent at the aggregate level, no amount of aggregate data can tell a researcher which one is true – illustrating, before the paper even turns to the New Keynesian model, that a reduced-form wedge (or an equivalently-specified DSGE shock) is not by itself informative about the right policy response.

Q4. Which shocks in the Smets-Wouters (2007) model do the authors single out as “dubiously structural,” and how much of the model’s action do they drive?

Of the seven shocks in Smets-Wouters (2007), the authors treat three as arguably structural – total factor productivity, investment-specific technology, and monetary policy – and four as dubiously structural: the wage markup, price markup, exogenous spending, and risk premium shocks. (Section II.A, pp. 253-256) These four are quantitatively dominant: the paper’s own forecast-error variance decomposition (Table 1, p. 256) attributes to shocks 4-7 a 39.6 percent share of output variance, 53.9 percent of hours, and 86.9 percent of inflation at a 4-quarter horizon; 44.0/69.2/86.5 percent at 10 quarters; and 60.3/86.0/88.0 percent at a 1,000-quarter horizon – meaning the model’s long-run predictions for inflation and hours, and much of its predictions for output, are driven mostly by shocks the authors argue cannot be signed as good or bad.

Q5. What specifically is wrong with the wage markup shock?

The wage markup shock is, mechanically, identical to inserting an exogenous labor wedge into the model, and literally interpreted as a fluctuating elasticity of substitution between worker types it implies an estimated standard deviation of 2,587 percent – “several orders of magnitude outside of a reasonable range.” (Section II.A, p. 257) Beyond its implausible literal magnitude, the shock is open to multiple, mutually exclusive interpretations that the model cannot distinguish: it could reflect fluctuating union bargaining power (a bad shock implying union-busting is optimal) or fluctuating value of leisure (a good shock implying laissez-faire is optimal), echoing exactly the indeterminacy from Q3; the authors note that under the leisure interpretation, the model’s implied “potential output” falls sharply in 1979-1984, implying the era’s recessions resulted from workers experiencing “a contagious attack of laziness” – a reading they clearly find implausible. (pp. 258-260; Figure 3)

Q6. What do the authors find wrong with the exogenous spending shock and the risk premium shock?

The exogenous spending shock has 3.5 times the variance of measured US government spending because it is defined residually from the national income identity and absorbs items such as net exports, which are plainly not invariant to monetary policy; the risk premium shock has more than six times the variance of short-term nominal rates and is best interpreted as capturing “flight to quality” episodes in financial markets – again “hardly likely to be invariant to monetary policy.” (Section II.A, pp. 260-261; Figure 4) In both cases the problem is the same as with the wage markup shock: the shock’s magnitude and behavior do not match what its name suggests it should be measuring, and its most economically sensible interpretation fails the policy-invariance requirement the authors set out in Q1.

Q7. Why do the authors object to the model’s backward price-indexation assumption, and why does it matter for policy?

Smets-Wouters (like Christiano-Eichenbaum-Evans 2005) assume that non-adjusting firms mechanically index their prices to lagged inflation, which implies every price changes every period – directly at odds with micro evidence that individual prices change only every 4 months on average (Bils and Klenow 2004) or every 11 months (Nakamura and Steinsson 2008). (Section II.B, pp. 261-263; Figure 5) This is not a purely academic mismatch: the authors argue that backward indexation makes the modeled costs of disinflation look very high, whereas if inflation persistence instead comes from a random-walk component in the Fed’s own policy rule, disinflation costs are low – and they cite Cogley and Sbordone (2005) and Ireland (2007) as showing the New Keynesian Phillips curve fits without backward indexation once such a policy component is included. (p. 264)

Q8. What is the authors’ separate objection to the model’s Taylor-rule specification?

Standard New Keynesian models assume the short-term nominal interest rate is stationary and ergodic, as in a Taylor rule, but the authors note that short and long rates share common secular trends and that the expectations hypothesis implies changes in the short rate shift long-run rate expectations – both features that instead point to a large random-walk component in the Fed’s actual policy behavior, which is “hard to reconcile with the smooth long-run rates implied by the use of the Taylor rule.” (Section II.B, p. 264) This is presented as a second, independent instance of the paper’s broader complaint: a convenient modeling assumption embedded in the state-of-the-art model that does not survive contact with the time-series behavior of the data it is meant to describe.

Q9. Given this critique, do the authors think New Keynesian and neoclassical economists actually disagree about what policy should do?

No – the authors emphasize that despite their critique of the Smets-Wouters model’s structural credentials, New Keynesian and neoclassical policy recommendations have converged on broad principles such as commitment to rules and low average inflation, citing Correia, Nicolini, and Teles (2008) showing that in a sticky-price model with a rich enough set of fiscal instruments, optimal policy coincides with the flexible-price neoclassical optimum. (Section III, pp. 264-265) They frame this convergence as New Keynesian modelers moving toward positions long held by neoclassical economists such as Lucas and Stokey, and are careful to state that their critique is about the reliability of these models for fine-grained, quarter-to-quarter quantitative policy counterfactuals – not a claim that the models’ broad qualitative policy lessons are wrong. (pp. 245-246, 264-265)

Key terms in this paper

Definitions below follow the paper's own usage.

Structural shock (in this paper's sense)
a shock usable for policy analysis must be invariant to the policy intervention under consideration and interpretable as unambiguously "good" (to be accommodated) or "bad" (to be offset); a shock that would itself change under the policy being evaluated, or whose sign of desirability cannot be determined, is "dubiously structural" even if it appears as a primitive disturbance in an estimated DSGE model.
Business cycle accounting wedges
the four reduced-form objects from Chari-Kehoe-McGrattan (2007) -- the efficiency wedge (like a TFP shock), the labor wedge (like a labor income tax), the investment wedge (like a capital-accumulation tax), and the government-consumption wedge -- that, fit to the stand-in neoclassical growth model, together reproduce essentially all the movement in US aggregate output; the paper stresses that a wedge is purely a measured summary statistic, consistent with many different underlying structural models.
Labor wedge
the object 1 minus the effective labor tax rate in the prototype economy's intratemporal first-order condition; the paper's Propositions 1 and 2 show the identical labor wedge process can be generated either by fluctuating union monopoly power (an inefficient, "bad" wedge calling for union-busting) or by fluctuating utility of leisure (an efficient, "good" wedge calling for laissez-faire), so the wedge alone cannot tell a researcher which policy is optimal.
Dubiously structural shock
the paper's label for the four Smets-Wouters shocks (wage markup, price markup, exogenous spending, risk premium) that, when interpreted literally, imply implausible magnitudes (e.g., a 2,587 percent standard deviation for the wage markup) or absorb residual, policy-sensitive components (e.g., net exports in the spending shock), and that admit multiple economically distinct interpretations the aggregate data cannot adjudicate between.
Flight to quality
the authors' preferred economic interpretation of the Smets-Wouters risk premium shock -- an abrupt increase in consumers' preference for holding government debt during financial market stress -- offered as the "only sensible" reading of a shock whose variance is more than six times that of short-term nominal rates, but one the authors argue is still not invariant to monetary policy and therefore not truly structural.
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.