Measuring Monetary Policy
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
How can you tell whether the Federal Reserve surprised the economy or simply reacted to it? This 1998 paper argues the standard statistical recipes imposed arbitrary assumptions, and instead writes down how the Fed actually operates in the market for bank reserves, fitted to monthly United States data from 1965 to 1996. Rival recipes give similar-looking responses but magnitudes differing by factors of two to four, because each mixes the true policy surprise with shocks to banks' demand for reserves. That contamination can explain why policy seemed to stop moving interest rates after the mid-1980s. It matters because the measured size of policy's effects depends on the measuring instrument.
What this paper finds — and why it matters
Bernanke and Mihov (1998, QJE) address a measurement problem in the monetary-VAR literature: the era’s dominant methods for extracting a “monetary policy shock” – Cholesky orderings built around the federal funds rate (FFR), nonborrowed reserves (NBR), or the ratio of nonborrowed to total reserves (NBR/TR, following Strongin 1995) – impose ad hoc restrictions rather than an explicit model of how the Fed actually operates in the market for bank reserves. The paper builds a “semi-structural” six-variable VAR (total reserves, nonborrowed reserves, and the federal funds rate in a policy block; real GDP, the GDP deflator, and a commodity price index in a block-exogenous macro block) over monthly U.S. data, 1965:1-1996:12, with lag length chosen by a general-to-specific procedure (13 lags retained for the full sample). Within the policy block they specify structural equations for reserves demand, borrowed-reserves supply, and the Fed’s nonborrowed-reserves supply/policy rule, estimated by two-step efficient GMM using the reduced-form VAR residuals; this yields five nested identifications – FFR, NBR, NBR/TR, borrowed-reserves (BR), and a just-identified (JI) benchmark that restricts only the reserves-demand elasticity to zero – each corresponding to different assumptions about how much the Fed accommodates reserves-demand shocks (phi^d) and borrowing shocks (phi^b). Overidentifying-restriction tests reject the NBR model in every subsample except the 1979:10-1982:10 Volcker NBR-targeting episode, favor NBR/TR as the best overall fit across subsamples, and favor FFR pre-1979 and post-1988; a Hamilton (1989) regime-switching model applied directly to (phi^d, phi^b) independently locates two structural breaks matching the conventional 1979/1982 Volcker-experiment dates, with phi^d falling from roughly 0.86 outside that window to roughly 0.18 within it. Although the five identifications generate qualitatively similar impulse responses (output rising with a 12-18 month peak, prices moving more slowly and persistently), they diverge sharply in magnitude – the cumulative output response under NBR/TR is more than twice as large as under NBR, and at a four-year horizon the price response under NBR is roughly four times that under FFR – because, the authors show algebraically, any overidentified indicator is a contaminated mixture of the true policy shock with reserves-demand and borrowing shocks whenever the Fed accommodates those shocks. This contamination framework offers an explanation for the post-1984 “vanishing liquidity effect” as a measurement artifact of using the NBR indicator rather than a change in the economy: a derived bias formula, evaluated at the JI model’s full-sample GMM estimates, is negative in the full sample (bias term -0.777) and grows more negative in 1988-1996 (-0.904) and 1984-1996 (-1.035, where it flips the estimated liquidity effect’s sign). The paper closes by constructing an overall monetary-policy-stance index from the full vector of structural innovations, normalized by a 36-month moving average, which correlates 0.71 (monthly) with the Boschen-Mills (1991) index; the authors caution that the empirical illustrations are meant only to be illustrative, since the Figure II impulse responses pool full-sample JI parameter estimates across regimes that the paper’s own regime-switching results show are unstable.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What problem motivates this paper, and what does it offer instead of the standard recursive VAR indicators?
Bernanke and Mihov argue that the standard recursive VAR indicators of monetary policy – orderings built around the federal funds rate, nonborrowed reserves, or the nonborrowed/total reserves ratio (Strongin 1995) – are Cholesky-style restrictions that do not correspond to an explicit model of how the Federal Reserve actually operates in the market for bank reserves, so empirical results can be sensitive to which indicator is chosen. Instead of picking one ordering on priors, they build a “semi-structural” model of the reserves market with parameters governing how much the Fed accommodates non-policy disturbances, and show that the standard indicators are special cases of this model corresponding to particular parameter restrictions – which lets the indicators be tested against each other rather than one being assumed correct.
Q2. How is the underlying VAR specified?
The system is a six-variable monthly VAR for the United States, 1965:1-1996:12, split into a “policy block” of total reserves (TR), nonborrowed reserves (NBR), and the federal funds rate (FFR), and a “macro block” of real GDP, the GDP deflator, and a Dow-Jones spot commodity price index, with the macro block assumed block-exogenous within the month (policy-block shocks cannot affect the macro variables contemporaneously) – a timing assumption the authors say “seems consistent with… monthly data” (Section II, p. 874). Lag length is chosen by a general-to-specific procedure starting from 15 lags, with 13 retained for the full monthly sample; the model is also re-estimated at biweekly frequency and across four subsamples (1965:1-1979:9, 1979:10-1996:12, 1984:2-1996:12, 1988:9-1996:12) as robustness checks.
Q3. How does the reserves-market model identify the policy shock, and what are the five candidate indicators?
Within the policy block, BM specify three structural equations – reserves demand (elastic in the funds rate, slope alpha), borrowed-reserves supply (elastic in the funds-rate/discount-rate spread, slope beta), and a nonborrowed-reserves/policy equation in which the Fed’s accommodation of demand and borrowing shocks is governed by parameters phi^d and phi^b – and invert these to recover the exogenous policy shock v^s as a linear combination of the three VAR residuals. Setting (phi^d, phi^b) to particular values nests five identifications: FFR (phi^d=1, phi^b=-1: the Fed fully accommodates all shocks to hold the funds rate fixed), NBR (phi^d=0, phi^b=0: pure nonborrowed-reserves targeting), NBR/TR a la Strongin (alpha=0, phi^b=0), BR (phi^d=1, phi^b=alpha/beta: borrowed-reserves targeting), and a just-identified (JI) model that imposes only alpha=0 and estimates phi^d and phi^b freely.
Q4. How are the structural parameters estimated?
Estimation proceeds in two steps: the reduced-form VAR is estimated by OLS, and its residuals are then used as data to estimate the structural parameters (alpha, beta, phi^d, phi^b) by efficient GMM, with the overidentifying restrictions of each of the four restricted models (FFR, NBR, NBR/TR, BR) tested against the data. Because the JI model imposes only one restriction (alpha=0), it has no overidentifying restrictions of its own to test and instead serves as the benchmark the other four are checked against.
Q5. Which identification fits best, and does the answer change across the Volcker period?
The NBR model’s overidentifying restrictions are rejected at conventional significance levels in every subsample except 1979:10-1982:10 – the conventional dates of the Volcker nonborrowed-reserves-targeting experiment – while the NBR/TR (Strongin) specification has the best overall fit across subsamples and the FFR model performs well before 1979 and after 1988. Consistent with this, a Hamilton (1989) two-state regime-switching model fit directly to the structural accommodation parameters (phi^d, phi^b), without imposing subsample break dates by assumption, independently locates two breaks – one in late 1979, one in 1982 – bracketing the same Volcker episode: phi^d and phi^b are estimated at roughly 0.863 and -0.778 outside that window (an interest-rate-focused regime that largely accommodates reserves-demand shocks) versus roughly 0.183 and -0.216 within it (an NBR-targeting regime with little demand accommodation) (Figure I, p. 891).
Q6. How much do the five identifications’ implied effects of policy actually differ?
All five identifications generate impulse responses that are qualitatively “reasonable” – an expansionary shock raises output relatively quickly, peaking around 12-18 months, and raises prices more slowly and persistently – but the magnitudes diverge sharply: the cumulative output response under NBR/TR is more than twice as large as under NBR, and at a four-year horizon the price response under the NBR identification is about four times the response under FFR. The authors conclude “these differences are certainly large enough to have important effects on the inferences one draws about the effects of monetary policy” (p. 894); the exercise itself is explicitly “meant only to be illustrative” (p. 896), since it applies full-sample JI parameter estimates to a VAR that pools across the regime changes documented in Q5.
Q7. Why do the indicators diverge, and what does this imply for the “vanishing liquidity effect”?
Algebraically, each overidentified indicator (FFR, NBR, NBR/TR, BR) can be written as a linear combination of the true policy shock v^s plus reserves-demand shocks v^d and borrowing shocks v^b, weighted by phi^d and phi^b (Eqs. 13-16); when the Fed accommodates demand shocks heavily (phi^d close to 1, as estimated outside the Volcker window), the NBR indicator in particular loads heavily on v^d rather than v^s, biasing its implied impulse responses. Formalizing this as a bias term in the NBR-based liquidity-effect estimate (Eq. 17), the authors evaluate it at the JI model’s GMM estimates and find it negative and mostly growing in magnitude – -0.777 in the full sample, -0.904 in 1988-1996, and -1.035 in 1984-1996 (a sign reversal, i.e. the NBR indicator implies the wrong sign for the liquidity effect in that subperiod) – leading them to conclude that “the ‘vanishing liquidity effect’ may well be the result of using a biased indicator of monetary policy, rather than of a change in the economy” (p. 896).
Q8. What overall policy-stance measure does the paper propose, and how does it compare to existing indicators?
Beyond the five discrete indicators, BM construct a continuous overall policy-stance index from the full vector of structural innovations (A^{-1}(I-G)P), in which one element’s VAR innovations correspond exactly to the JI model’s policy shocks; because the funds rate and nonborrowed-reserves growth are not measured in comparable units across regimes, the raw series is normalized by subtracting a trailing 36-month moving average, with zero defined as “normal” policy. This index correlates 0.71 with the Boschen-Mills (1991) qualitative index at monthly frequency, and the Romer and Romer (1989) narrative dates line up with peaks of maximum policy tightness in BM’s index “rather than points at which policy changed from expansionary to contractionary” (p. 898).
Q9. What limitations do the authors themselves flag?
The authors note that the block-exogeneity timing assumption is shared with the standard Cholesky approach and is not independently tested beyond its plausibility for monthly data; that the discount rate is set to zero in estimation because discount-rate changes are infrequent, a simplification that could matter in periods of active discount-rate adjustment; and that the model’s institutional structure is calibrated to the U.S. reserves market, so applying it elsewhere requires adapting the reserves-market equations to local institutions (p. 900). They also flag, as noted in Q6, that the illustrative impulse responses pool across regimes their own regime-switching results show are structurally unstable – though a robustness check using regime-switching rather than full-sample parameter estimates gives similar results (fn. 24, p. 892).
Key terms in this paper
Definitions below follow the paper's own usage.
- semi-structural VAR
- BM's term for their reserves-market model -- a VAR in which the macro block is treated as block-exogenous (Cholesky-style, not separately identified) but the policy block (total reserves, nonborrowed reserves, federal funds rate) is identified using an explicit economic model of reserves demand, borrowed-reserves supply, and the Fed's reaction function, rather than an arbitrary causal ordering.
- accommodation parameters (phi^d, phi^b)
- the two structural parameters in BM's nonborrowed-reserves/policy equation that measure how much the Fed lets nonborrowed reserves adjust to absorb reserves-demand shocks (phi^d) and borrowing shocks (phi^b) rather than holding NBR (or the funds rate) fixed; setting these parameters to specific values nests the FFR, NBR, NBR/TR, and BR indicators as special cases, and their estimated values -- and the regime shift in them BM document around 1979-1982 -- are what identify how the Fed's operating procedure has changed over time.
- just-identified (JI) model
- BM's benchmark specification, which imposes only the single restriction that reserves demand has zero interest-rate elasticity in the structural sense (alpha=0, so that total-reserves innovations equal the demand shock v^d) and leaves phi^d and phi^b to be estimated freely by GMM from the data, rather than fixed a priori as in the four overidentified indicators.
- contamination
- BM's term for the way any of the four restricted policy indicators (FFR, NBR, NBR/TR, BR), if the JI model is the correct one, mixes the true monetary policy shock v^s with non-policy reserves-market shocks (demand shock v^d, borrowing shock v^b) in proportions governed by phi^d and phi^b -- so an indicator's estimated impulse responses partly reflect responses to these other shocks rather than to actual policy.
- overall policy-stance index
- BM's continuous summary measure of monetary policy, constructed from the full vector of structural VAR innovations A^{-1}(I-G)P and normalized by subtracting a trailing 36-month moving average (so that zero corresponds to "normal" policy), designed to be comparable across the different Fed operating-procedure regimes the paper documents, unlike the funds rate or NBR growth taken alone.