Macro Paper Warehouse
Published Classic [Journal of Economic Literature] doi:10.1257/jel.49.3.673 Vol. 49, No. 3, pp. 673-685

Can Government Purchases Stimulate the Economy?

Valerie A. Ramey — University of California, San Diego, and National Bureau of Economic Research

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

How much extra output does a dollar of government purchases buy? Writing after the 2008 crisis, Ramey surveys theory and three empirical literatures: aggregate time series, and newer cross-state and cross-locality natural experiments. Neoclassical models can give multipliers from strongly negative to modestly above one depending on how spending is financed; New Keynesian models give small multipliers in normal times but potentially much larger ones at the zero lower bound. She concludes the multiplier relevant to a temporary, deficit-financed stimulus is probably between 0.8 and 1.5, likely higher in severe recessions -- while stressing that standard errors are always large and the transmission mechanism remains disputed.

What this paper finds — and why it matters

This essay surveys the state of knowledge on the government-spending multiplier as of 2011, drawing together theory, U.S. aggregate time-series evidence, and the newer cross-state/cross-locality “natural experiment” literature, to answer the question most relevant to the post-2008 stimulus debate: how much output does a temporary, deficit-financed increase in government purchases buy? Ramey first shows that “the multiplier” is not one number: in neoclassical models the multiplier operates through a negative wealth effect on labor supply and can range from as low as -2.5, when spending is temporary and financed by concurrent distortionary taxes, to about 1.2 for a permanent increase financed by lump-sum taxes, because the size and even sign of the multiplier depends critically on the persistence of the spending and how it is financed; in New Keynesian models the multiplier is typically smaller than in traditional Keynesian models (often at or below one) unless the model assumes widespread rule-of-thumb consumers and demand-determined employment, or unless the nominal interest rate is stuck at the zero lower bound, in which case rising expected inflation lowers the real interest rate and multipliers well above one become possible (Christiano, Eichenbaum, and Rebelo’s estimate peaks at 2.3). Turning to the data, she reviews a range of aggregate U.S. time-series studies – from the large-scale econometric models of the 1960s through modern structural VARs using military-spending instruments, narrative “news” shocks, and Blanchard-Perotti-style identification – and finds that, despite substantial differences in sample, experiment, and identification strategy, most estimates cluster between roughly 0.6 and 1.8, with standard errors that are always large; she argues the lower end of this range mostly reflects episodes (such as the Korean War) in which spending was partly tax-financed, and that after truncating for that and for concerns about persistence and measurement in the highest estimates, the range most relevant to a temporary, deficit-financed increase is 0.8 to 1.5, with multipliers likely toward the top of that range during severe recessions. She also surveys the newer literature using cross-state and cross-region panel data (military contracts, ARRA allocations, New Deal spending), which mostly finds income multipliers of about 0.5 to 2.0 and implied costs of roughly $25,000-$35,000 per job created, while cautioning that these local multipliers answer a conceptually different question from the aggregate multiplier and require an explicit model (of the kind developed by Nakamura and Steinsson 2011 and Shoag 2010) to translate into an aggregate estimate. The essay closes by noting that despite the proliferation of estimates, there remains no consensus on the underlying transmission mechanism – some studies find consumption falls in response to spending increases, consistent with the neoclassical wealth effect, while others find it rises, consistent with rule-of-thumb consumers – and that none of the multiplier estimates reviewed speaks to the welfare consequences of stimulus spending, which would require additional assumptions about mechanisms and about whether government purchases enter the utility function.

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What is the paper’s bottom-line estimate of the government-spending multiplier?

“I conclude that the multiplier for this type of spending [a temporary, deficit-financed increase in government purchases] is probably between 0.8 and 1.5” (Abstract, p. 673), a conclusion the author restates in the body text: “the range of plausible estimates for the multiplier in the case of a temporary increase in government spending that is deficit financed is probably 0.8 to 1.5” (Section 3, p. 680-681). She is explicit about the uncertainty surrounding this range: “reasonable people could argue that the multiplier is 0.5 or 2.0 without being contradicted by the data” (p. 681), and that if the spending increase occurs during a severe recession, estimates are likely to sit at the upper end of the 0.8-1.5 range (p. 681).

Q2. Why does Ramey insist that “the multiplier” is not a single, well-defined number?

Because it “depends very much on the type of government spending, its persistence, and how it is financed” (Introduction, p. 673). She restricts her headline estimate specifically to spending that is temporary, deficit-financed (rather than tax-financed), and that “enter[s] separately in the utility function and have[s] no direct effect on private sector production functions” – i.e., not productive public investment or transfers, both of which she treats separately later in the essay (Introduction, p. 673).

Q3. How does the neoclassical model generate multipliers as low as -2.5 or as high as 1.2, and what is the common mechanism?

In neoclassical models, government spending affects output only through its effect on equilibrium hours worked, which operates through a negative wealth effect: a rise in g reduces the value function’s implied wealth, and because leisure is a normal good, “the function h is strictly increasing in g” (Section 2.1, p. 674, following Aiyagari, Christiano, and Eichenbaum 1992). Absent distortionary taxes, Ricardian equivalence means financing is irrelevant; but Baxter and King (1993) show that once distortionary financing is introduced, the multiplier can be “as low as negative 2.5” for a temporary spending increase financed by concurrent distortionary taxes, “substantially below unity” for temporary spending financed by future lump-sum taxes, and around 1.2 in the long run for a permanent increase financed by lump-sum taxes, because “the greater negative wealth effect raises labor supply more and the steady-state capital stock rises” (p. 675). Burnside, Eichenbaum, and Fisher (2004) further show that when distortionary taxes rise only gradually (a hump-shaped path, as observed historically), intertemporal substitution toward present labor supply pushes short-run multipliers higher than under the lump-sum case (p. 675).

Q4. What does the New Keynesian literature say about the multiplier, and when can it exceed one?

Standard New Keynesian models typically generate multipliers smaller than the traditional Keynesian cross would predict, “since the New Keynesian model builds a sticky-price edifice on a neoclassical foundation,” and Cogan et al. (2010) using the Smets-Wouters model find multipliers at or below unity (Section 2.2, p. 675-676). Galí, López-Salido, and Vallés (2007) obtain multipliers as high as 2.0, but only by assuming at least 50 percent rule-of-thumb consumers and fully demand-determined employment – assumptions that “essentially convert the New Keynesian model back into a traditional Keynesian model” (p. 676). The one channel that can raise multipliers without resorting to such assumptions is the zero lower bound: “a deficit-financed increase in government spending leads expectations of inflation to increase. When nominal interest rates are held constant, this increase in expected inflation drives the real interest rate down, spurring the economy,” and Christiano, Eichenbaum, and Rebelo (2011) show that if the nominal rate is held constant for twelve quarters while spending rises, the multiplier peaks at 2.3 (p. 676).

Q5. What other features do most multiplier models abstract from, and how much do they matter?

Ramey flags three omissions – productive government spending, transfers, and underutilized resources – and finds mixed evidence that any of them substantially raises short-run multipliers (Section 2.3, pp. 676-677). Baxter and King (1993) find that public capital raising private marginal products can generate long-run multipliers of 4.0 to 13.0, “but much lower in the short run,” so it “does not raise the predicted short-run stimulus effects.” Because most of the 2009-2010 U.S. stimulus was allocated to transfers rather than purchases, and standard models treat temporary transfers like a negative lump-sum tax with no effect under the permanent-income hypothesis and Ricardian equivalence, Oh and Reis (2011) explore relaxations of these assumptions but are “not able to generate much bigger effects” (p. 677). Whether underutilized capacity raises multipliers is, in Ramey’s assessment, “a promising area for more theoretical research” rather than a settled question (p. 677).

Q6. What does the aggregate U.S. time-series evidence show, and how consistent is it across studies?

“Despite significant differences in samples, experiments, and identification methods, most aggregate studies estimate a range of multipliers from around 0.6 to 1.8,” with within-study ranges “almost as wide as the range across studies” and consistently large standard errors (Section 3, p. 679, Table 1). The evidence spans Evans’s (1969) large-scale Keynesian econometric models (multipliers “slightly above 2.0”), military-spending-instrument studies by Barro, Hall, and Barro-Redlick (0.6 to 1.25), Ramey-Shapiro-style narrative/anticipation-based VARs (0.6 to 1.57), Blanchard-Perotti SVARs (0.9 to 1.29), and Auerbach-Gorodnichenko’s regime-switching estimates, which imply multipliers around 2.2 in recessions versus -0.3 in expansions under one specification, and a more plausible 1.0-1.5 in recessions versus 0-0.5 in expansions when regimes are allowed to switch endogenously (pp. 679-680).

Q7. Why does Ramey exclude the lowest and highest aggregate estimates from her headline range?

She “truncated the lower estimates because they were usually accompanied by increases in distortionary taxes” – for instance, her own lower estimates come from samples dominated by the Korean War, which was substantially tax-financed, and Barro and Redlick (2011), controlling for the average marginal tax rate, find a multiplier of only 0.6 – “and truncated the very highest estimates because of the various concerns” raised earlier, including Fisher and Peters’s (2010) finding of an unusually persistent spending shock (which a neoclassical model would predict raises the multiplier above the value relevant to a genuinely temporary stimulus) and possible anticipation effects contaminating Gordon and Krenn’s (2010) 1940-41 estimates (Section 3, pp. 679-681).

Q8. Does the historical zero-lower-bound period (1939-1947) show larger multipliers, as the New Keynesian theory would predict?

No – restricting the sample to 1939-1949, a period during which Treasury bill rates never rose above 0.38 percent despite six percent average inflation, Ramey (2011) finds a multiplier of 0.7, “though with even larger than normal standard errors,” so “I find no evidence of larger multipliers during the extended period in which interest rates were held virtually constant at the zero lower bound” (Section 3, p. 680). This is presented as evidence against, or at least not clearly in favor of, the theoretical prediction that multipliers should be systematically larger when the zero bound binds, though the author notes the small effective sample size limits how much weight this particular comparison can bear.

Q9. What does the cross-state/cross-locality literature find, and why does Ramey caution against equating it with the aggregate multiplier?

Cross-state income-multiplier estimates mostly fall between 0.5 and 2.0, with an implied cost of roughly $25,000-$35,000 per job created across several studies (Section 4, pp. 681-683, Table 2), and several papers (Shoag 2010; Suárez Serrato and Wingender 2011; Nakamura and Steinsson 2011) find multipliers are significantly larger during periods of greater economic slack. But Ramey illustrates with a simple example why these are not directly the aggregate multiplier: if the federal government transfers $1 to one state, financed by lump-sum taxes spread across all states, “the true aggregate multiplier is 0, since the taxes and transfers cancel in the aggregate,” yet a panel regression with time fixed effects – which nets out the economywide tax increase – recovers a spurious multiplier of mpc/(1 - mpc), e.g., 1.5 if the marginal propensity to consume is 0.6, “even though the aggregate multiplier for this experiment is 0” (Section 4, p. 682). Translating local multipliers into an aggregate figure therefore requires an explicit theoretical model of the kind developed by Nakamura and Steinsson (2011) and Shoag (2010), which treat individual states as small open economies within a currency union.

Q10. What unresolved questions does Ramey flag for future research?

Despite the proliferation of multiplier estimates, “there is still no consensus on the mechanism by which government spending raises GDP”: some studies find consumption falls in response to a spending increase, consistent with the neoclassical negative wealth effect, while others find it rises, consistent with rule-of-thumb consumers, and Nekarda and Ramey (2011) present evidence that industry markups do not fall in response to government spending as the New Keynesian model requires (Conclusions, p. 683). She also stresses that “none of these estimates sheds light on the welfare consequences of temporary increases in government spending to stimulate the economy,” which would require both a resolved mechanism and explicit assumptions about whether government purchases enter household utility (p. 683).

Key terms in this paper

Definitions below follow the paper's own usage.

The temporary, deficit-financed multiplier
the paper's preferred experiment for assessing stimulus-relevant multipliers -- an increase in government purchases that is (1) temporary, (2) financed by deficits rather than concurrent taxes, and (3) enters the utility function separately with no direct effect on private production; Ramey stresses that "the multiplier" is not a single number but depends fundamentally on this choice of experiment, and that other commonly estimated experiments (permanent increases, tax-financed increases, productive public capital) imply different, sometimes very different, multiplier values.
Neoclassical wealth-effect decomposition of the multiplier
the paper's account of the shared neoclassical mechanism (Aiyagari, Christiano, and Eichenbaum 1992): government spending acts as a negative wealth shock, inducing the household to substitute out of leisure into labor ("static effect") and, if spending is persistent, to raise desired future capital as well ("dynamic effect"); because leisure is normal and government spending raises hours, output can rise even without demand-side frictions, but the same logic implies multipliers as low as -2.5 when spending is financed by concurrent distortionary taxes and as high as 1.2 for permanent, lump-sum-financed spending.
The zero-lower-bound channel for larger multipliers
the paper's review of Eggertsson, Eggertsson-Woodford, Christiano-Eichenbaum-Rebelo (2011), and Woodford (2011): in a New Keynesian model in which the nominal interest rate is held at zero, a deficit-financed spending increase raises expected inflation, which drives down the real interest rate and stimulates private spending -- the one channel by which the New Keynesian model can generate multipliers well above one (Christiano-Eichenbaum-Rebelo's estimate peaks at 2.3) without resorting to widespread rule-of-thumb, nonoptimizing behavior.
Cross-state multipliers versus the aggregate multiplier
Ramey's methodological point that estimates from panel or cross-section regressions of state or local outcomes on state-level spending or transfers answer a different question than the aggregate multiplier: in a simple Keynesian illustration where $1 is transferred to one state and financed by taxes spread across all states, the true aggregate multiplier is zero, yet a panel regression with time fixed effects (which nets out the economywide tax increase) recovers a spurious multiplier of mpc/(1 − mpc); translating cross-state estimates into aggregate multipliers requires an explicit open-economy/currency-union model, as in Nakamura and Steinsson (2011) and Shoag (2010).
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.