Macro Paper Warehouse
Published Classic [American Economic Journal: Macroeconomics] doi:10.1257/mac.4.2.1 Vol. 4, No. 2, pp. 1-32

Are the Effects of Monetary Policy Shocks Big or Small?

Olivier Coibion

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

How much does a surprise interest-rate change actually move the economy? Two standard approaches disagreed sharply, one implying a 0.7 percent fall in industrial production after a one-percentage-point tightening, the other 4.3 percent. Using monthly United States data from 1970 to 1996, this 2012 paper traces the gap to three fixable causes: the methods imply different paths for the policy rate, the 1979-82 reserve-targeting episode supplies shocks that were partly predictable, and one method is unusually sensitive to how many lags are included. Corrected, both converge near 1.8 to 2.8 percent. It matters because that middle range is what quantitative models must reproduce.

What this paper finds — and why it matters

This 2012 American Economic Journal: Macroeconomics paper by Olivier Coibion asks why standard recursive VARs and the Romer-Romer (2004) narrative approach produce such different estimates of the real effects of monetary policy shocks, and argues the discrepancy is driven by three identifiable and largely correctable factors rather than genuine disagreement about the underlying policy shock. Using monthly US data (1970:1-1996:12, following Romer and Romer’s sample and variable choices), Coibion compares a standard recursive (Cholesky) VAR with the funds rate ordered last and 12 lags, the Romer-Romer single-equation approach (a macro variable regressed on 24 autoregressive lags and 36 lags of the R&R narrative shock series), and a “hybrid VAR” that replaces the funds rate in the VAR with the cumulative R&R shock series. At baseline lag specifications the three diverge sharply (Table 2, Panel A): a 100-basis-point shock produces a peak industrial production decline of about 0.7% and a 0.16 percentage-point rise in unemployment in the standard VAR, versus a 4.3% IP decline and 0.93 percentage-point unemployment rise in the R&R single-equation approach, with the hybrid VAR intermediate (1.6% IP decline, 0.40 points). Three mechanisms account for most of this gap: (1) an R&R shock corresponds to a larger, more persistent counterfactual path for the funds rate than a VAR Cholesky shock, and rescaling impulse responses to a common FFR path largely eliminates the difference; (2) R&R shocks during the 1979:10-1982:9 non-borrowed-reserves (NBR) targeting episode are statistically predictable from lagged macro variables (violating the shocks’ assumed exogeneity) and disproportionately drive the large R&R estimates; and (3) the R&R two-step procedure’s results rise nearly monotonically with the number of shock lags included, while AIC- or model-averaging-based lag selection picks shorter lags and shrinks the estimated effects (reducing the peak IP effect from -4.3% to -3.4%), whereas the VAR is comparatively insensitive to lag length. Once the NBR-targeting period is excluded and lag length is chosen by AIC or model averaging, the three approaches converge (Table 2, Panel B) on a “medium-range” peak effect of roughly a 1.8-2.8% industrial production decline and a 0.19-0.70 percentage-point unemployment increase; three additional Taylor-rule-residual shock measures constructed in Section III (GARCH, time-varying-coefficient, and Smets-Wouters DSGE-filtered shocks) corroborate this medium range and are markedly less sensitive to lag length or episode selection than the original R&R series, though the paper is explicit that its evidence covers only the pre-1997 sample and does not speak to the post-1996 or zero-lower-bound period.

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What puzzle motivates the paper, and what is its headline answer?

The paper asks why monetary VARs and the Romer-Romer (2004, “R&R”) narrative approach imply such different real effects of monetary policy shocks, and answers that the gap is largely explained by three identifiable, correctable factors rather than a genuine disagreement about the true policy shock (pp. 1-3). Once these three factors — a different implied path for the federal funds rate (FFR), contamination from the non-borrowed-reserves (NBR) targeting episode, and lag-length sensitivity — are controlled for, the paper argues the competing approaches converge on a “medium range” estimate of the real effects of monetary policy.

Q2. What three baseline approaches does the paper compare, and how different are their estimates at face value?

At baseline lag specifications, a 100-basis-point policy shock produces wildly different peak effects across methods (Table 2, Panel A, p. 21): a standard recursive VAR (Cholesky-identified, FFR ordered last, 12 lags, following R&R’s variable choices) implies industrial production (IP) falls about 0.7% and unemployment rises 0.16 percentage points; the R&R single-equation approach (a macro variable regressed on 24 autoregressive lags plus 36 lags of the R&R narrative shock, following R&R’s own baseline procedure) implies a much larger 4.3% IP decline and 0.93-point unemployment rise; and a “hybrid VAR” (the baseline VAR with the cumulative R&R shock series substituted for the FFR as the policy variable) lands in between, at a 1.6% IP decline and 0.40-point unemployment rise. Both the sample (monthly US data, 1970:1-1996:12, with R&R shocks derived over 1969:3-1996:12) and prices move similarly: -0.1% (VAR), -4.2% (R&R), -1.8% (hybrid).

Q3. Why does the “different contractionary impetus” mechanism matter, and how much of the gap does it close?

A VAR Cholesky shock and an R&R narrative shock are not the same object: an R&R innovation generates a larger and more persistent rise in the FFR than an equivalent VAR shock, and once impulse responses are rescaled onto a common counterfactual FFR path, the difference in estimated macroeconomic effects “largely disappears” (Section II.B, pp. 10-12). Coibion shows this by applying the single-equation approach to all three shock series and normalizing by each series’ implied FFR response (Figure 4); doing so brings the VAR and hybrid-VAR results into line with the R&R results, indicating that part of the apparent disagreement is really a disagreement about how big a “shock” is, not about the transmission mechanism.

Q4. What role does the non-borrowed-reserves targeting period (1979:10-1982:9) play in the discrepancy?

R&R shocks during the Federal Reserve’s 1979:10-1982:9 non-borrowed-reserves (NBR) targeting experiment are statistically predictable from lagged macroeconomic variables — F-statistics of 2.28 and 2.45 (both significant at the 5% level) for industrial production and inflation respectively over the full sample (Table 1, pp. 12-16) — which violates the identifying assumption that the shocks are exogenous, and this episode disproportionately drives the large estimated R&R effects. Bernanke-Mihov and Boschen-Mills alternative measures of policy stance show much smaller swings over this period than R&R do, and dropping the NBR period from the sample brings all three approaches’ estimates together (Figure 7); Coibion concludes the original discrepancy is “largely driven by the early Volcker era” (p. 20).

Q5. How sensitive are the results to the number of lags used, and what happens under AIC or model-averaged lag selection?

The R&R single-equation results rise nearly monotonically as the number of shock lags increases (Figure 8), while the VAR estimates are comparatively insensitive to lag length; selecting lags by AIC instead of the R&R baseline choice (24 autoregressive, 36 shock lags) reduces the R&R peak IP effect from -4.3% to -3.4%, with model averaging (Kass-Raftery bootstrap weights) giving essentially the same result (Section II.D, pp. 17-20; Table 2, Panel A, p. 21). This is presented as a purely methodological source of disagreement: the R&R two-step procedure’s sensitivity to an atheoretical tuning choice, rather than the VAR’s, generates part of the apparent gap in shock magnitude.

Q6. Once all three factors are controlled for, what is the paper’s “consensus” estimate of the real effects of a monetary policy shock?

With the NBR-targeting period excluded (1979:10-1983:12 omitted) and lag length chosen by AIC or model averaging, peak industrial production estimates converge across all three approaches to roughly a 1.8% to 2.8% decline, with unemployment rising by about 0.19 to 0.70 percentage points, depending on approach and lag method (Table 2, Panel B, p. 21). The paper states its key conclusion in these terms (p. 20, the paper’s own emphasis): “the real effects of monetary policy shocks appear to be in the medium range, with industrial production dropping by approximately 2-3 percent and unemployment rising by around half a percentage point.”

Q7. What do the three alternative Taylor-rule-residual shock measures (Section III) add, and how do they compare?

Three additional shock measures constructed from Taylor-rule residuals — a GARCH(1,1) version of the R&R rule, a time-varying-coefficient (TVC) version re-estimated at each FOMC meeting, and Smets-Wouters (2007) DSGE-filtered shocks — imply 100-basis-point effects of roughly a 2% IP decline and 0.60-point unemployment rise (GARCH); a 4% IP decline, 0.75-point unemployment rise, and 4% price decline (TVC, at baseline lags); and about a 1.5% IP decline, 0.40-point unemployment rise, and 1.5% price decline (Smets-Wouters) (Section III.A-C, pp. 22-27). All three are markedly less sensitive than the original R&R series to lag length, to dropping individual historical episodes, or (for GARCH and Smets-Wouters especially) to excluding the NBR period; correlations with the original R&R shock series are 0.93 (GARCH), 0.82 (TVC), and only 0.40 (Smets-Wouters), and the Smets-Wouters shocks are 69% more volatile than R&R’s.

Q8. What does the paper conclude about the historical contribution of monetary policy shocks to US fluctuations?

All four measures (original R&R, GARCH, TVC, and Smets-Wouters) point to a “nontrivial” contribution of monetary policy shocks to historical US fluctuations that is “considerably smaller than implied by the original R&R results but also considerably larger than those implied by the baseline VAR” (Section III.D, p. 28). Exogenous monetary policy is estimated to have contributed approximately 2-3 percentage points of the roughly 5-percentage-point rise in unemployment between 1979 and 1982, and much of the disinflation of the early 1980s is attributed to exogenous policy — a feature the paper notes standard monetary VARs do not share. The contribution of monetary policy shocks to real fluctuations is described as having “drastically declined” since the early 1980s, which the paper interprets as consistent with improved monetary policy conduct (the Great Moderation).

Q9. What caveats and limitations does the paper flag about its own approach?

The sample ends in 1996:12, limited by the availability of the Greenbook forecast data the R&R approach requires, so the paper explicitly cannot speak to monetary policy effects after 1996 or in a zero-lower-bound environment (p. 31: measuring policy stance “is rife with difficulties even in calm periods,” and doing so “during extraordinary circumstances such as the Volcker experiment or the current period presents a particularly challenging exercise”). The TVC specification also imposes a narrower information set than the original R&R rule (contemporaneous rather than lagged forecasts, p. 25 fn. 27), and the Smets-Wouters filtered shock uses realized data rather than Greenbook forecasts, so it combines true monetary innovations with rational-expectations forecast errors (p. 26, fn. 28). Online Appendix Monte Carlo simulations show model averaging has a small downward bias relative to the truth, though smaller than the bias from under- or over-fitting lag length.

Key terms in this paper

Definitions below follow the paper's own usage.

Recursive (Cholesky) VAR shock
in this paper, the "standard" small-effects benchmark — a monetary policy shock identified by ordering the federal funds rate last in a Cholesky decomposition of a 12-lag monthly VAR (following Romer and Romer's variable choices), so the shock is orthogonal to contemporaneous values of the other VAR variables by construction.
Romer-Romer (R&R) narrative shock
the shock series from Romer and Romer (2004), constructed as the residual from a regression of intended FFR changes around FOMC meetings on the Federal Reserve's own Greenbook forecasts; used in this paper's "single-equation approach," where a macro variable is regressed on 24 autoregressive lags and 36 lags of this shock series (Eq. 1).
Hybrid VAR
the paper's own intermediate construction — the same baseline recursive VAR specification, but with the cumulative R&R narrative shock series substituted for the funds rate as the policy variable, isolating how much of the VAR-versus-R&R gap comes from the identification scheme itself versus the funds-rate variable.
Non-borrowed reserves (NBR) targeting period
the Federal Reserve's October 1979-September 1982 operating procedure (part of the "Volcker experiment") that generated unusually large funds-rate swings; in this paper it is the specific episode during which R&R shocks are shown to be statistically predictable from lagged macro data (Table 1), violating the exogeneity assumption needed for identification and disproportionately driving the largest R&R effect estimates.
Time-varying-coefficient (TVC) Taylor rule shock
a Taylor-rule-residual shock measure (following Boivin 2006 and Coibion-Gorodnichenko 2011) in which the policy-rule coefficients are re-estimated at each FOMC meeting rather than held fixed, with two volatility breaks (1979, 1982) imposed so that regime changes in the rule's coefficients are classified as rule changes rather than as policy shocks.
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.