Monetary policy matters: Evidence from new shocks data
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
Why do the standard ways of measuring monetary policy surprises stop working after 1988? This 2013 paper shows all four leading methods then imply output rises after a tightening, because the Federal Reserve turned more forward-looking and the models omit what it reacts to. The authors build a new measure from movements in interest-rate futures around 157 policy meetings between 1988 and 2008; industrial production then falls significantly, peaking near two years, and the measure accounts for some 40 to 50 percent of output variation at long horizons, a share the authors note is inflated by the period's unusual calm. Why it matters: an apparent breakdown was measurement.
What this paper finds — and why it matters
This 2013 Journal of Monetary Economics paper by S. Mahdi Barakchian and Christopher Crowe argues that the standard toolkit for identifying monetary policy shocks – recursive VARs (Christiano-Eichenbaum-Evans), over-identified VARs (Bernanke-Mihov), non-recursive VARs (Sims-Zha), and the Romer-Romer narrative measure – stops delivering plausible results once the sample extends past 1988, because the Federal Reserve’s increasingly forward-looking behavior violates the identifying assumptions these methods rely on: when a VAR omits the forward-looking variables the Fed actually reacts to, its endogenous response to anticipated conditions gets misread as an exogenous shock, biasing the estimated output effect upward. Using each method’s own original specification, the authors show all four produce the “wrong” sign in the post-1988 period – a contractionary shock followed by rising, not falling, output (for example, the CEE recursive VAR shows output declining in 1960Q1-1992Q4 but increasing significantly in 1988Q4-2007Q3, and the Bernanke-Mihov, Sims-Zha, and Romer-Romer methods show the same reversal over their own comparison samples). To sidestep the problem, they construct a new high-frequency-identification (HFI) shock measure from the innovations in six Fed Funds futures contracts (current month through five months ahead) around each of 157 FOMC meetings between 1988:12 and 2008:06, extracting a two-factor structure by maximum likelihood in which the first factor – explaining 92% of the variance, interpreted as a “level” shift in the expected medium-term policy path – is used as the shock, while a second “slope” factor (9% of variance, associated with forward-guidance content) is set aside. Feeding the cumulated new shock into a three-variable monthly VAR (log industrial production, log CPI, and the cumulated shock, with 36 lags and the shock ordered last) over 1988:12-2008:06, industrial production shows a statistically significant, sustained decline after a contractionary shock, with the maximum impact around a two-year horizon – recovering the sign that the conventional methods lose in this period – while prices exhibit a milder price puzzle (turning significantly negative only after about four years) that the authors cannot fully resolve even after adding a commodity price index or survey-based inflation expectations. Forecast error variance decompositions show the new shock accounts for roughly 40-50% of industrial production’s variance at a three-year-plus horizon, around twice the share attributed to existing shock measures over the same period, though the authors note this partly reflects the historically low volatility of the “Great Moderation” sample. A regression of the new shock on the Fed’s exclusive information (Greenbook-Blue Chip forecast gaps for 17 variables, following Romer-Romer, over 113 meetings in 1988-2002) finds no significant joint explanatory power (R-squared = 0.185, F(17) = 1.50, p = 0.132), suggesting the measure is relatively free of the simultaneity bias that could contaminate a directly observed policy-rate innovation – though two individual coefficients (current-quarter output growth and current-quarter GDP deflator) are significant, pointing to some remaining contamination that the authors argue would bias the estimated effects toward zero rather than away from it.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What puzzle motivates the paper, and what do the authors argue is its cause?
All four leading approaches to identifying monetary policy shocks – the Christiano-Eichenbaum-Evans (CEE) recursive VAR, the Bernanke-Mihov over-identified VAR, the Sims-Zha non-recursive VAR, and the Romer-Romer narrative measure – produce economically implausible results once applied to post-1988 US data: a contractionary shock is followed by a rise, not a fall, in output. The authors attribute this to the Fed becoming “more forward-looking” after 1988, which “invalidat[es] the identifying assumptions in conventional methods of measuring monetary policy’s effects, leading to spurious and unlikely results for this period” (Abstract, p. 950). Two structural breaks in the estimated reaction function, dated to 1979:10 and 1982:10, and the growing importance of forward-looking elements in the Fed’s information set are cited as the mechanism (Section 2.2).
Q2. How exactly does each conventional method fail, and over what samples?
Each of the four methods is run on both an earlier comparison sample and a post-1988 sample using its own original specification, and each flips sign in the later sample. The CEE recursive VAR shows output declining over 1960Q1-1992Q4 but rising significantly over 1988Q4-2007Q3 (Fig. 1); the Bernanke-Mihov over-identified VAR shows industrial production (IP) declining over 1965:01-1996:12 but rising over 1988:12-2007:11 (Fig. 2); the Sims-Zha non-recursive VAR shows GNP declining over 1964Q1-1994Q4 but rising over 1988Q4-2007Q4 (Fig. 3); and the Romer-Romer narrative measure shows IP declining over 1969:01-1996:12 but rising significantly over 1988:12-2008:06 (Fig. 4). The authors summarize: “the existing identification schemes lead to estimated impulse responses in the post-1988 sample that are both different from those found in the earlier samples, and counter to what most central bankers would find plausible” (p. 955).
Q3. Why does increased forward-looking behavior by the Fed break these identification schemes specifically?
The mechanism is an omitted-variable problem: when the Fed’s reaction function depends on anticipated future output and inflation (elements of its information set, Omega_t) that a VAR does not include, the VAR’s residual conflates the genuine policy shock with the Fed’s endogenous response to those anticipated conditions, biasing the estimated output effect upward. As the authors put it, “the key change is to the estimated effect rather than the actual effect, because identification problems become more pronounced when the Fed’s policy becomes more forward looking” (p. 962). This same logic implies a fourth identifying assumption for any shock measure: the measured innovation should not itself reveal information from the Fed’s private forecasts or reaction function beyond what is already priced in by the market.
Q4. How is the new high-frequency-identification (HFI) shock measure constructed?
Following Kuttner (2001) and Gurkaynak-Sack-Swanson (2005), the shock is the revision in the private sector’s expectation of the policy path implied by Fed Funds futures prices around each of 157 FOMC meetings from 1988:12 to 2008:06 (the futures market itself only began in October 1988), formalized as the pre-announcement expectation of the policy path minus its realized value (Eq. 6, p. 958), under the assumptions that private-sector beliefs prior to the announcement are correct on average and that a background “noise” term is unchanged around the announcement (Eqs. 4-5, p. 958). Rather than using a single futures contract, the authors combine daily price innovations from six Fed Funds futures contracts spanning the current month through five months ahead.
Q5. What does the two-factor structure of the futures innovations look like, and why use only the first factor as the shock?
A maximum-likelihood two-factor model fit to the six futures-contract innovations yields a first factor explaining 92% of total variance, on which all six contracts load positively, interpreted as a “level” shift in the expected medium-term policy-rate trajectory; a second factor explains the remaining 9%, with loadings switching from positive to negative as contract maturity increases, interpreted as a “slope” or yield-curve effect (Eq. 11 and footnote 14, p. 959). Only the first (level) factor is used as the policy shock measure; the authors note the second factor plausibly captures forward-guidance content in FOMC communications but do not incorporate it into the shock series.
Q6. When fed into a VAR, how does the new shock’s implied effect on output and prices compare to the conventional methods’ post-1988 results?
In a three-variable monthly VAR (log industrial production, log CPI, and the cumulated new shock, with 36 lags, a constant and linear trend, and the shock ordered last) over 1988:12-2008:06, industrial production falls significantly and persistently after a contractionary shock, “with a maximum impact at a horizon of around two years” – similar in shape to the pre-1988 CEE and Romer-Romer results, but the opposite of what those same methods find in the post-1988 period (Section 4.1, p. 962). Prices are more problematic: “the effect becomes significantly negative only after four years; the positive response over the medium term, although small, contrasts with the negative effect that has generally been found in the literature” (p. 962). This price puzzle is described as milder than in the conventional post-1988 results but is not resolved by adding a commodity price index or one-quarter-ahead Blue Chip inflation expectations (Section 4.2).
Q7. How much of output’s variance does the new measure explain relative to existing shock measures, and what caveat applies to that comparison?
Forecast error variance decomposition shows the new shock accounts for roughly 40-50% of industrial production’s forecast error variance at a horizon of three years or more, “around twice the proportion using existing shock measures” over the same period, with the Romer-Romer and change-in-the-funds-rate measures each accounting for roughly 10-20% at the same horizon (Fig. 7, p. 964; Abstract, p. 950). The authors caveat that this comparison is made “in an era of low overall output volatility” (the Great Moderation), so part of the new measure’s relatively large share may reflect the low volatility of the denominator rather than a larger absolute effect (Section 4.1).
Q8. Is the new shock measure contaminated by the Fed’s private information, and how might that bias the results?
Regressing the new shock on the gap between Greenbook and Blue Chip forecasts for 17 variables in the Romer-Romer reaction function (113 FOMC meetings, 1988-2002) yields R-squared = 0.185 and F(17) = 1.50 (p = 0.132), so the authors cannot reject the hypothesis that the Fed’s exclusive private information has no joint explanatory power for the new shock at conventional levels (Table 1, p. 961). Two individual coefficients are nonetheless significant – current-quarter output growth (-2.37, significant at 1%) and current-quarter GDP deflator (2.34, significant at 5%) – suggesting some contamination by the Fed’s reaction to contemporaneous conditions. The authors argue any resulting simultaneity bias works toward attenuation (biasing estimated effects toward zero), so the paper’s reported effects “likely understate the true effect” (Section 4.3, pp. 963-965).
Q9. How robust are the results, and what illustrative episode do the authors offer as face validity for the new measure?
Across eight robustness checks – alternative VAR ordering (shock first), an alternative price index (PPI), alternative lag lengths (6, 12, 24 lags), subsample stability, adding a commodity price index, adding inflation expectations, excluding FOMC dates that coincide with the Employment Report, and a single-equation Romer-Romer-style specification – the qualitative results are described as essentially unchanged, with the price puzzle persisting throughout (Section 4.2). As an illustrative episode, the authors point to the June 27, 2001 FOMC meeting, where the Fed cut rates by 25 basis points, less than the 50 basis points the market had priced in: conventional VAR and narrative methods would record this as an expansionary shock, but the new measure records it as contractionary (roughly two standard deviations above zero), consistent with the subsequent rise in bond yields and the dollar’s climb to a ten-week high (Section 3.5, pp. 961-962). The authors also show that using only the current-month (spot) futures innovation, as in Kuttner (2001), produces an imprecise, positive output response, underscoring the value of combining information across the six-contract term structure rather than relying on the spot-month shock alone.
Key terms in this paper
Definitions below follow the paper's own usage.
- High-frequency-identification (HFI) shock (this paper's construction)
- the revision in the private sector's expectation of the medium-term policy-rate path around an FOMC announcement, measured as the first of two maximum-likelihood factors extracted from the daily innovations in six Fed Funds futures contracts (current month through five months ahead); it is distinguished from a single-contract (spot-month) surprise because it aggregates information across the futures term structure.
- Level factor vs. slope factor
- in the paper's two-factor decomposition of futures innovations, the level factor (92% of variance) has uniformly positive loadings across all six contract horizons and represents a parallel shift in the expected policy path; the slope factor (9% of variance) has loadings that switch sign with maturity and is interpreted as capturing forward-guidance content distinct from the immediate policy move -- only the level factor is used as the shock measure in this paper.
- Forward-looking reaction function / Omega_t
- the set of variables -- including anticipated future output and inflation -- that the Fed is argued to respond to when setting policy after 1988; because conventional VARs omit these forward-looking elements, the Fed's endogenous response to them is misattributed to an exogenous policy shock, which the authors identify as the source of the post-1988 identification failure.
- Price puzzle (as it appears in this paper)
- the finding that prices respond positively, rather than negatively, to a contractionary monetary shock over the medium term before eventually turning significantly negative after about four years; the paper's new HFI measure produces a milder version of this puzzle than conventional post-1988 estimates, and the authors are unable to resolve it by adding commodity prices or survey inflation expectations.
- Simultaneity (attenuation) bias
- the component of the new shock measure that reflects the Fed's systematic response to its own internal (Greenbook) forecasts rather than a pure exogenous innovation; the paper argues this contamination, to the extent it is present, biases the estimated output and price effects toward zero, so the paper's reported magnitudes are likely conservative rather than overstated.