Macro Paper Warehouse
Published Classic [Handbook of Macroeconomics] doi:10.1016/bs.hesmac.2016.03.003

Macroeconomic shocks and their propagation

Valerie A. Ramey

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Are economists doomed never to know what causes recessions? This 2016 survey takes stock, comparing the main strategies for isolating surprise changes in monetary policy, government spending, taxes and technology, often re-estimating them on updated data. Estimates of a one-percentage-point rate rise range from no effect at all to output falls of several percent, with no consensus magnitude, and a popular market-based measure of policy surprises turns out to be partly predictable from the Federal Reserve's own internal forecasts. Government-spending multipliers cluster between 0.6 and 1.5. Why it matters: it shows how much disagreement is method rather than economics.

What this paper finds — and why it matters

This 2016 Handbook of Macroeconomics chapter by Valerie Ramey is a survey and methodological stock-taking exercise on how macroeconomists identify structural shocks to monetary policy, fiscal policy (government spending and taxes), and technology, and how those shocks propagate through the economy — organized around the question Cochrane (1994) posed of whether researchers are “destined to remain forever ignorant of the fundamental causes of economic fluctuations,” to which Ramey’s answer is “no.” Rather than introducing a new identification scheme, the chapter sets out a common reduced-form/structural framework (reduced-form residuals e_t = Bepsilon_t, impulse responses C(L) = A(L)^-1B) and classifies the literature into roughly nine identification strategies — recursive/Cholesky orderings (with the policy variable ordered either last, as in Christiano-Eichenbaum-Evans, or first), non-recursive structural restrictions, narrative methods, high-frequency identification, external-instrument/proxy SVARs, long-run restrictions, sign restrictions, and factor-augmented VARs — and then re-estimates or directly compares many of them on overlapping or updated data. For monetary policy, a comparison of eleven leading published methods on their original samples finds trough effects of a 100-basis-point funds-rate increase on output/industrial production ranging from a positive, statistically insignificant response (Uhlig’s 2005 sign-restriction estimate) to declines of several percent (Romer-Romer’s 2004 narrative measure implies -4.3% at 24 months; Barakchian-Crowe’s 2013 high-frequency estimate implies -5% at 23 months), with no consensus on the “true” magnitude. Ramey’s own re-estimation of the Christiano-Eichenbaum-Evans-style recursive VAR shows the effect shrinking from a -0.48% trough (accounting for 6.6% of the output forecast-error variance at 24 months) over 1965-1996 to a -0.20% trough (0.5% of variance) over 1959-2007, and turning expansionary when the sample is restricted to 1983-2007. She then subjects the Gertler-Karadi (2015) high-frequency proxy instrument to validity checks and finds it has a nonzero mean (-0.013), is serially correlated (AR(1) coefficient of 0.31, robust SE 0.11), and is predictable from Federal Reserve Greenbook forecasts (R-squared of 0.21, p=0.027) — evidence that the instrument is contaminated by the Fed’s private information rather than being a clean exogenous surprise — leading her to conclude that what is now identified as a monetary policy shock is “really mostly the effects of superior information on the part of the Fed, foresight by agents, and noise.” For fiscal shocks, government-spending multiplier estimates cluster mostly between 0.6 and 1.5 once computed correctly as the integral of the output response over the integral of the spending response (rather than a peak-to-impact ratio), while tax-multiplier estimates vary more across method, from roughly -0.5 to -0.8 (SVAR-based estimates using outside output elasticities) up to about -3 (Romer-Romer 2010 narrative and Mertens-Ravn 2014 proxy-SVAR estimates). For technology shocks, the chapter documents low correlations between shocks extracted from different identification approaches and finds that investment-specific technology (IST) and marginal-efficiency-of-investment (MEI) shocks — including news about future IST change (Ben Zeev-Khan 2015, estimated to explain 73% of output variance at 8 quarters in one specification) — tend to matter more for business-cycle fluctuations than neutral total-factor-productivity shocks, whose estimated importance ranges from negligible (Galí 1999) to more than 40% of output variance depending on method. A combined VAR that Ramey estimates herself, incorporating leading fiscal, technology, and monetary shock measures together over 1954-2005, finds that the Ben Zeev-Khan IST-news shock is the single most important driver of both output and hours (accounting for up to 40% of the hours forecast-error variance at 8 quarters), that the Justiniano-Primiceri-Tambalotti MEI shock explains 42% of output variance on impact falling to 24% within a year, that the federal funds rate accounts for up to 8% of output and 18% of hours variance, and that all the included shocks together explain 63-79% of the variance of output and hours at horizons of 4-20 quarters — evidence Ramey reads as showing genuine, if incomplete, progress against the identification problem Cochrane posed two decades earlier, even as she stresses that none of the individual identification strategies commands full consensus.

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. How does the chapter organize the identification literature, and what common framework does it use to compare otherwise very different methods?

Ramey classifies the empirical monetary, fiscal, and technology-shock literatures into roughly nine identification strategies — Cholesky/recursive VARs with the policy variable ordered last (the Christiano-Eichenbaum-Evans convention) or first, non-recursive structural VARs, narrative methods, high-frequency identification (HFI), external-instrument/proxy SVARs, long-run restrictions, sign restrictions, and factor-augmented VARs (FAVARs) — and evaluates all of them against a single reduced-form/structural framework. The framework starts from a reduced-form VAR, A(L)Y_t = e_t, with reduced-form residuals related to structural shocks by e_t = Bepsilon_t (epsilon_t having mean zero and identity covariance), so that the structural impulse response is C(L) = A(L)^-1B. Each identification strategy amounts to a different way of pinning down enough of B (or of avoiding needing all of it) to isolate a single structural shock. The chapter’s stated purpose is not to argue for one method over the others but to lay out their assumptions transparently and then, in later sections, actually re-estimate or otherwise directly compare leading methods on the same or updated data — an unusually empirical form of survey.

Q2. What is a proxy SVAR / external-instrument approach, and how does it differ from the traditional Cholesky ordering it is often used to avoid?

An external-instrument (proxy SVAR) approach identifies a structural shock using an outside variable Z_t that is correlated with the true structural shock of interest (relevance: E[Z_tepsilon_1t] is not 0) but uncorrelated with all the other structural shocks in the system (exogeneity: E[Z_tepsilon_it] = 0 for i not equal to 1), and does not require imposing any particular causal ordering among the VAR’s own variables. This matters because a Cholesky/recursive identification implicitly assumes that the policy variable does not respond contemporaneously to some variables (if ordered early) or that other variables do not respond contemporaneously to policy (if ordered late) — a strong assumption whose validity is rarely tested. The chapter describes a three-step implementation: estimate the reduced-form VAR, use the instrument to identify the column of B corresponding to the shock of interest, and then compute impulse responses from that identified column. Examples applied later in the chapter include Gertler-Karadi’s (2015) use of federal-funds-futures surprises as a monetary policy instrument and Mertens-Ravn’s (2013, 2014) use of Romer-Romer narrative tax shocks as an instrument in a fiscal VAR.

Q3. Why does the chapter treat “foresight” — the idea that private agents or the Fed anticipate policy moves before they occur — as a deep, unresolved identification problem rather than a technical nuisance?

When policymakers follow a forward-looking rule that responds to their own anticipated future values of the economy, the residual that a standard VAR labels a “shock” can be shown algebraically to conflate the true policy surprise with (a) the gap between the true and the econometrician’s estimated policy-rule coefficients and (b) the Fed’s anticipated future information, so that what looks like an identified shock is partly just the Fed reacting to information the econometrician’s VAR does not contain. Formally, if the true rule is x_1t = sum of alpha_k * E_t[x_2,t+k] plus a genuine shock epsilon_ft, then the residual an econometrician recovers by regressing on only the current value of x_2 equals (alpha_0 - alpha_0-hat)*x_2t + the sum of the forward-looking terms + epsilon_ft — a mixture, not a clean shock. Ramey identifies this “foresight” or information problem as the central obstacle behind the disappointing performance of many recursive monetary-shock measures (see Q6), and notes that the literature’s main responses are to bring in survey expectations, narrative evidence, high-frequency surprises around policy announcements, or richer lag structures that better proxy for the private information the policymaker is reacting to.

Q4. When Ramey compares roughly a dozen leading, separately published monetary-shock identification methods, how much do they agree on the output effect of a monetary policy tightening?

They disagree substantially: across the compared studies, the estimated trough effect of a 100-basis-point rise in the federal funds rate on output/industrial production ranges from a positive (expansionary), statistically insignificant response under Uhlig’s (2005) sign-restriction VAR to declines of roughly 4-5% under narrative and high-frequency methods, with several recursive and FAVAR-based estimates clustering in the -0.6% to -2% range. Christiano et al.’s (1999) original recursive SVAR (1965q3-1995q3 sample) implies about a -0.7% trough at 8 quarters; Bernanke, Boivin, and Eliasz’s (2005) FAVAR implies about -0.6% at 18 months; Coibion’s (2012) “robust” Romer-Romer respecification implies about -2% at 18 months; Barakchian and Crowe’s (2013) high-frequency-instrument VAR implies about -5% at 23 months; and Gertler and Karadi’s (2015) high-frequency proxy SVAR implies about -2.2% at 18 months. The chapter treats this spread as itself an important empirical fact: methods that look individually reasonable and are estimated by careful researchers nonetheless disagree by close to an order of magnitude on the central quantity of interest.

Q5. How fragile is the classic Christiano-Eichenbaum-Evans-style recursive monetary VAR to the sample period used, based on Ramey’s own re-estimation?

Re-estimating essentially the same recursive specification but updating the sample, Ramey finds the estimated output effect of a monetary contraction shrinks sharply and eventually flips sign: a -0.48% trough (accounting for 6.6% of the output forecast-error variance at 24 months) over 1965m1-1996m6 falls to a -0.20% trough (0.5% of variance) over 1959m1-2007m12, and over the shorter 1983m1-2007m12 window the estimated response of unemployment to a funds-rate increase is actually expansionary. She interprets the shrinking and eventual sign reversal as consistent with two, not mutually exclusive, explanations: monetary policy has genuinely become more systematic and less prone to large discretionary surprises since the Volcker disinflation, and/or later-sample VARs increasingly pick up the Fed’s forward-looking, information-based responses rather than true exogenous shocks (the foresight problem of Q3). A parallel re-estimation using the Romer-Romer (2004) narrative shock in a VAR shows a similar pattern of a materially larger effect in the original 1969m3-1996m12 sample (-1.38% trough, 8.8% of variance) than in an extended 1969m3-2007m12 sample (-0.83% trough, 2.7% of variance).

Q6. What specific validity problems does Ramey document with the Gertler-Karadi (2015) high-frequency proxy instrument, and what conclusion does she draw about monetary shock identification generally?

Ramey finds three problems with the Gertler-Karadi instrument that each point toward contamination by the Fed’s private information rather than a clean exogenous surprise: its mean is -0.013 rather than zero; regressing it on its own lag gives an AR(1) coefficient of 0.31 (robust SE 0.11), i.e., it is serially correlated when a true shock should not be; and regressing it on the Greenbook forecast variables the Romers used to construct their own shock gives an R-squared of 0.21, rejecting the joint null that those coefficients are zero at p=0.027 — meaning the instrument is statistically predictable from the Fed’s own internal forecasts. She concludes from this evidence, together with the sample fragility documented in Q5 and the wide cross-method disagreement documented in Q4, that “true monetary shocks are now rare” and that what current methods identify as a monetary policy shock is “really mostly the effects of superior information on the part of the Fed, foresight by agents, and noise” — a conclusion she frames as bad news for the econometrics of monetary identification but good news for the conduct of monetary policy itself, since it implies policy has become more systematic and less prone to discretionary surprises.

Q7. What does the chapter’s survey of fiscal-shock estimates conclude about the size of government-spending and tax multipliers, and what methodological correction does Ramey emphasize?

Ramey argues that many published multiplier estimates are not properly comparable because they are calculated as a peak-to-impact ratio rather than as the integral (present value) of the output response divided by the integral of the spending response — the peak-to-impact convention tends to overstate the multiplier relative to the integral measure — and, once this correction and a related log-scaling issue are taken into account, most government-spending-multiplier estimates in the aggregate US literature cluster between about 0.6 and 1.5, with military-news-based approaches (Ramey 2011; Ben Zeev-Pappa) and Blanchard-Perotti-style SVARs (roughly 0.9 to 1.3 on a peak basis) both represented within that range, while some state- or region-level panel estimates and recession-conditional estimates (Auerbach-Gorodnichenko) come in noticeably higher. Tax-multiplier estimates vary more by identification strategy: SVAR approaches using externally imposed output elasticities of tax revenue (Blanchard-Perotti, Caldara-Kamps) imply multipliers of roughly -0.65 to -1.3, while narrative and proxy-SVAR approaches built on Romer-Romer’s (2010) legislated tax-change series imply substantially larger multipliers of about -3. Ramey highlights work by Mertens and Ravn showing that the assumed output elasticity of tax revenue is itself a key driver of these differences, since setting it too low leaves a positive reverse-causality contamination (output raising measured tax revenue) inside what is meant to be an exogenous tax shock.

Q8. What does the chapter’s comparison of technology-shock identification methods find about agreement across approaches, and which type of technology shock emerges as most important for business cycles?

Cross-method correlations among identified technology shocks are low, and the chapter interprets this as evidence that different SVAR and DSGE approaches to identifying “the” technology shock are not recovering the same underlying object; estimates of the share of output forecast-error variance explained by neutral TFP shocks range from “very little” (Galí 1999, using long-run restrictions) to over 40% (Christiano et al.’s long-run-restriction SVAR reports 31-45% for horizons up to 20 quarters; Francis et al.’s medium-horizon-restriction approach reports 15-40% for horizons up to 32 quarters), while investment-specific technology (IST) and marginal-efficiency-of-investment (MEI) shocks are repeatedly found to matter more. In particular, Ben Zeev and Khan’s (2015) IST news shock is estimated to account for 73% of output variance at 8 quarters in their preferred medium-horizon-restriction specification, and DSGE-based estimates of MEI shocks reach as high as 60% of output variance at business-cycle frequencies (Justiniano et al. 2011) even as estimates of the same broad shock category vary enormously across studies (from near zero to over 60%, depending on model and horizon). Ramey reads this pattern — low cross-correlation of “the same” shock across methods, but a recurring finding that investment-related and news-based technology shocks dominate neutral TFP shocks — as one of the chapter’s more robust qualitative conclusions, even though the exact magnitude is not pinned down.

Q9. In Ramey’s own combined VAR spanning fiscal, technology, and monetary shocks simultaneously, which shocks matter most for output and hours, and how much of their variance is jointly explained?

Estimating a single VAR over 1954q3-2005q4 that includes leading measures of government-spending news, anticipated and unanticipated tax shocks, TFP shocks, IST news, and MEI shocks alongside output, hours, prices, and the federal funds rate (four lags, ordered with the funds rate last), Ramey finds that the Ben Zeev-Khan IST-news shock is the single most important shock for both output and hours — for example, accounting for 40% of the hours forecast-error variance at 8 quarters (90% CI: 25-54) — and that the combined contribution of the three fiscal shocks, the three technology shocks, and the federal-funds-rate shock together explains 63-79% of the variance of output and hours at horizons of 4-20 quarters. The Justiniano-Primiceri-Tambalotti MEI shock contributes 42% of output variance on impact (90% CI: 34-50), falling to 24% within a year, while the federal funds rate — interpreted as the monetary policy shock — accounts for up to 8% of output variance and up to 18% of hours variance at longer horizons. Because data availability for some shocks limits the sample to start after the Korean War, government-spending shocks contribute relatively little in this particular exercise, which Ramey explicitly flags as a sample-driven limitation rather than evidence that spending shocks are unimportant generally.

Q10. What other, more briefly surveyed shock categories does the chapter discuss, and what overall conclusion does Ramey draw about progress on Cochrane’s original question?

The chapter briefly surveys oil supply shocks (noting an unresolved debate over asymmetric effects and whether Cholesky orderings with oil ordered first actually isolate supply rather than demand-driven price movements), credit shocks (citing Gilchrist and Zakrajšek’s 2012 finding that innovations to the excess bond premium, orthogonal to current conditions, have significant macroeconomic effects), uncertainty shocks (flagged as an area needing more work to distinguish an independent shock from an endogenous propagation mechanism), and labor-supply/wage-markup shocks (citing Shapiro-Watson 1988 and DSGE estimates, including Khan-Tsoukalas 2012 and Schmitt-Grohe-Uribe 2012, finding wage-markup news shocks can account for up to 60% of the variance share of hours) — each treated as an active but less mature area than the monetary, fiscal, and technology literatures that anchor the chapter. Returning to the question posed in the introduction — whether researchers are “destined to remain forever ignorant of the fundamental causes of economic fluctuations” — Ramey answers “no,” pointing to the combined VAR exercise (Q9) as evidence that identified shocks now jointly account for a substantial majority of business-cycle variance in output and hours. She is careful to qualify this as real but incomplete progress: no single identification strategy commands consensus, several of the most-used monetary instruments fail validity checks (Q6), and she explicitly calls for more work testing the plausibility of identifying assumptions and the robustness of shock estimates before treating the combined-VAR shares as settled.

Key terms in this paper

Definitions below follow the paper's own usage.

Recursive/Cholesky identification
in this chapter, a VAR identification scheme that imposes a causal ordering on the contemporaneous impact matrix so shocks are recovered via a Cholesky factorization of the reduced-form covariance matrix; Ramey distinguishes "Type A" orderings (policy variable last, as in Christiano-Eichenbaum-Evans, so real variables and prices are contemporaneously unaffected by policy) from "Type B" orderings (policy variable first, as in some government-spending identifications, so policy is predetermined with respect to the rest of the system).
External instrument / proxy SVAR
Ramey's term for an identification approach that uses an outside variable Z_t satisfying relevance (correlated with the shock of interest) and exogeneity (uncorrelated with all other structural shocks) to identify a column of the structural impact matrix B without imposing a Cholesky ordering; implemented in the chapter via Gertler-Karadi's (2015) high-frequency funds-futures instrument for monetary shocks and Mertens-Ravn's (2013, 2014) narrative-based instrument for tax shocks.
Foresight / information problem
Ramey's diagnosis, formalized via a forward-looking policy rule, that a VAR residual meant to capture an exogenous policy shock instead conflates the true shock with the policymaker's response to privately anticipated future conditions and with any misspecification of the assumed policy rule; she treats this as the central unresolved obstacle to reliable monetary-shock identification in her own sample-fragility and instrument-validity findings.
Integral (present-value) multiplier
the way Ramey argues a fiscal multiplier should be computed -- the integral of the output response to a fiscal shock divided by the integral of the spending (or tax) response over the same horizon -- as opposed to the more common but, in her view, less meaningful peak-to-impact ratio, which can materially overstate the multiplier relative to the integral measure.
Local projections (LP) as a comparison tool
in this chapter, Jorda's (2005) horizon-by-horizon regression method used alongside conventional VAR impulse responses, particularly in the technology-shock and monetary-narrative sections, to check whether qualitative disagreements across identification approaches are an artifact of VAR misspecification or persist under a less parametric estimator; Ramey reports that the qualitative disagreements across methods generally persist under LP as well.
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.