What Does Monetary Policy Do?
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
What share of the economy's ups and downs actually comes from the Federal Reserve's own decisions, rather than the Fed simply reacting to what's already happening? Using U.S. data from 1960 to 1996 and comparing many statistical models built the same way for a fair test, this paper finds that only a modest portion — sometimes almost none — of the swings in output and prices traces back to monetary policy surprises. Most of the movement in interest rates and the money supply is the Fed responding normally to the economy, not driving it. That undercuts simple stories that credit or blame the Fed for most economic swings.
What this paper finds — and why it matters
This 1996 Brookings Papers on Economic Activity paper by Eric Leeper, Christopher Sims, and Tao Zha asks what monetary policy actually does, and argues that the answer the identified-VAR literature had been giving was fragile because different papers used different samples, variable sets, and identification schemes; the authors instead estimate an escalating sequence of Bayesian structural VARs – first small 4- and 5-6-variable systems that reinterpret existing recursive and non-recursive identifications (Strongin’s NBR/TR system, Christiano-Eichenbaum-Evans’s NBR system), then integrated 13- and 18-variable systems – all on one common monthly U.S. data set (January 1960-March 1996, six lags, quarterly series interpolated via Chow-Lin) so that conflicting conclusions can be checked for robustness on equal footing. Identification proceeds through a mix of exact zero restrictions, probabilistic (“soft-zero”) priors, and informal plausibility screening on impulse responses, organized around a sectoral block structure that splits variables into slow-moving private (“P”), fast-moving auction-market information (“I”), Federal Reserve policy (“F”), and banking (“B”) blocks, with the Fed’s reaction function restricted to respond within the month only to fast-moving financial information, not to CPI or GDP (which are measured with a lag); estimation uses a Bayesian “reference prior” (Sims-Zha) that downweights explosive, poorly-identified dynamics, and published error bands are 68% (about one standard error) probability bands from posterior simulation, not classical confidence intervals. Across the model sequence, a policy-tightening shock in the integrated 13- and 18-variable systems produces a broadly coherent contractionary picture – short and long rates rise, reserves and M1 fall smoothly, output and its components fall, unemployment rises, commodity prices fall, and the dollar appreciates – but the CPI response is very small (briefly slightly positive, only weakly negative after several years), and policy shocks account for only a modest share of output and (in the 13-variable model) a “negligible” share of CPI variance, while a separate private-sector shock – one the model allocates to the private sector with nothing in its estimated form suggesting misallocation, and which the authors judge it “unlikely” is “mistakenly incorporat[ing] much of an expansionary monetary policy disturbance” – is nonetheless the dominant source of M1 and reserves variation, showing that most of the observed variation in these aggregates is genuinely nonpolicy in origin and so is unsatisfactory as a one-dimensional policy indicator. Specifications that imply large real effects of policy (notably Strongin’s NBR/TR system, where policy shocks explain about half of output variance at three-plus-year horizons) tend to come bundled with an implausible sustained price-level rise after a contraction (a severe price puzzle), which the authors read as a sign of policy misspecification rather than a genuine large policy effect; the CEE NBR system, by contrast, produces believable but small real effects. The authors’ robust cross-specification conclusions are that only a modest portion (sometimes essentially none) of U.S. output and price variance since 1960 traces to monetary policy shocks, that most of the observed variation in policy instruments and monetary aggregates is the systematic, endogenous response of policy to the state of the economy rather than exogenous disturbance, that treating reserves or a monetary aggregate as moving mainly in response to policy is therefore unreliable, and that specifications implying large real effects tend to be exactly the ones with implausible price responses. The authors caution that their identified policy shocks may still be absorbing some adverse-supply-shock variation (which would understate price effects and overstate output effects), that the published hard-zero restrictions are a computationally tractable substitute for soft-zero priors the authors initially found more plausible, and that with small changes in specification a shock formally labeled “monetary policy” can rotate into an information-sector shock and vice versa – a fragility warning central to the paper’s message.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What problem is this paper trying to solve, and how does its approach differ from the existing identified-VAR literature on monetary policy?
The paper’s premise is that the existing literature’s disagreements about what monetary policy does are partly an artifact of different papers using different samples, variable sets, and identification schemes – so the authors re-estimate a whole sequence of competing identifications on one common monthly U.S. data set (1960-1996) to see which conclusions actually survive being run on equal footing. They explicitly reject the idea that policy-relevant models and well-fitting models are distinct, aiming “to show that it is possible to construct economically interpretable models with superior fit to the data” (Abstract; Introduction, pp. 1-2). Rather than proposing one final model, the paper is organized as an escalating sequence – 4-variable recursive and non-recursive systems, then 5-6-variable reserves-market reinterpretations of Strongin and Christiano-Eichenbaum-Evans (CEE), then integrated 13- and 18-variable systems – explicitly designed to test whether small-model conclusions hold up as the information set is enriched.
Q2. What data and identification machinery underlies all the models in the sequence?
All models are estimated on monthly U.S. data from January 1960 to March 1996 (using data from July 1959 onward) with six lags, with quarterly series such as GDP components interpolated to monthly frequency via the Chow-Lin (1971) procedure; all variables are in natural logs except interest rates and unemployment (Appendix A, pp. 59-63). Identification of the structural form A(L)y(t) = ε(t) (with Var(ε) = I) is achieved through a combination of (i) exact zero restrictions on the contemporaneous matrix A₀, (ii) probabilistic “soft-zero” priors expressing that some coefficients are more likely near zero than others, and (iii) informal plausibility screening that rejects models implying strongly positive price or output responses to a monetary contraction (Method, pp. 3-16). A recurring sectoral block structure groups variables into private “sluggish” (P), fast-moving auction-market “information” (I), Federal Reserve policy (F), and banking (B) blocks, with the key restriction that the Fed’s reaction function does not respond within the month to CPI or real GDP (measured with a lag) but does respond contemporaneously to fast-moving financial signals such as commodity prices, the 10-year rate, stock prices, and the dollar (pp. 25-26, 39-47). Estimation is Bayesian, using a Sims-Zha “reference prior” that downweights models with large coefficients on distant lags or explosive dynamics, which the authors need because the likelihood is ill-behaved with many parameters and near-nonstationary data; published error bands are 68% (roughly one-standard-error) posterior probability bands from simulation, not classical confidence intervals (Inference, pp. 15-16; fn. 12, p. 16).
Q3. In the smallest (4-variable) models, what do a recursive and a non-recursive identification each imply, and do they agree?
They do not fully agree. In a recursive 4-variable system (CPI, output, the federal funds rate, and M1), a positive M1 innovation produces a smooth, slow rise in prices and a quicker rise in output, but M1 shocks account for only a small share of output variance, and the system exhibits both the liquidity puzzle (the funds rate barely falls, and only temporarily, when M1 rises) and the price puzzle (prices rise following a contractionary interest-rate innovation) (pp. 21-24). Re-identifying the same variables non-recursively with the P/P/I/F sectoral scheme instead isolates an “F-column” shock that behaves like a genuine monetary tightening: the funds rate rises and decays back to baseline over about a year, M1 declines (with most of M1’s variance attributed to policy in this specification), and output declines persistently but only by about a tenth of a percent, while CPI moves negligibly (very slightly down). In the other (private/supply-shock) columns, prices and interest rates move together, implying that most observed interest-rate variation is an endogenous response to the economy rather than an exogenous policy shock (pp. 25-28). Replacing M1 with total reserves in this small system, the already-modest output response to a “money” shock “almost completely disappeared” and the price response weakened further – an early warning against reading any single aggregate as a one-dimensional indicator of policy (pp. 28-29).
Q4. When the paper re-estimates Strongin’s and Christiano-Eichenbaum-Evans’s reserves-market identifications on this common sample, do their original large-effect and small-effect conclusions survive?
Strongin’s identification (a 5-variable NBR/TR system) implies policy shocks explain about half of output variability at horizons of three-plus years, but comes bundled with a severe price puzzle – a substantial, sustained rise in prices following a contraction – which the authors interpret as evidence that the specification “involves unreasonable characterizations of policy behavior” that confound inflationary supply shocks with policy shocks (pp. 31-37). The CEE-style identification (a 6-variable NBR system), by contrast, produces effects the authors describe as “believable”: output rises persistently after an expansionary shock, but policy shocks account for only a small fraction of output variability; the price level does not fall after an expansion (essentially no effect on finished-goods prices); commodity prices move as expected; and the endogenous response of policy to information about future inflation dominates funds-rate variability. Even so, the authors judge that this specification rests on “assumptions about policy behavior and the reaction of the economy that do not seem plausible” (pp. 37-39). The pattern that emerges – large implied real effects paired with an implausible price response, versus small implied real effects paired with a more plausible price response – becomes one of the paper’s central diagnostic tools in the larger models.
Q5. What does moving to the integrated 13-variable model add, and does the earlier small-model picture hold up?
In the 13-variable system (short and long rates, reserves, M1, CPI, output, unemployment, investment components, consumption, stock prices, commodity prices, and the dollar), the identified contractionary policy shock produces an internally coherent contraction – short and long rates rise, reserves and M1 fall smoothly, output and all its GDP components fall, unemployment rises, commodity prices fall, and the dollar appreciates – but the CPI response stays very small (briefly slightly positive, only weakly and slowly turning negative, within about 2 standard errors of zero), and the combined policy-plus-banking block explains only a modest share of overall variance and a “negligible contribution” to CPI variance specifically (pp. 44-52). A separate private-sector (P) shock, not the policy shock, is the single most important source of variation in M1 and total reserves. The authors find nothing in its estimated form to suggest the model misallocates it to the P sector in error, and judge it “unlikely that it mistakenly incorporates much of an expansionary monetary policy disturbance” – that is, this is genuinely a private disturbance (one that expands the money supply and raises interest rates without affecting output), correctly identified as nonpolicy rather than a policy shock mislabeled as private. It is precisely because the model is led by the data to allocate so much variation in M1 and reserves to this legitimate nonpolicy disturbance that using either aggregate alone as a one-dimensional policy indicator is unsatisfactory – reinforcing the small-model finding (Q3), now in a richer, more defensible system.
Q6. Does adding still more variables (the 18-variable model) overturn any of these conclusions?
No – expanding to 18 variables (adding the discount rate, an alternative funds-rate specification, wages, the M2 deposit rate, bank securities and loans, and the prime rate) yields two distinct policy shock columns, a short-lived and a longer-lived tightening, but the results are “similar” to the 13-variable findings and described by the authors as “basically easy to defend” (pp. 52-57). The escalation from 4 to 13 to 18 variables therefore does not reverse the paper’s qualitative conclusions; if anything it stabilizes them, since the larger, richer information sets are precisely what let the authors distinguish a genuine policy shock from private-sector or information shocks that a smaller system would conflate.
Q7. What are the paper’s headline conclusions once robustness is checked across the whole model sequence?
The authors draw five conclusions they present as robust across most of the specifications (Conclusion, pp. 57-59): (1) “most of the specifications imply that only a modest portion (in some cases, essentially none) of the variance in output or prices in the United States since 1960 can be attributed to shifts in monetary policy”; (2) “a large fraction of the variation in monetary policy instruments can be attributed to the systematic reaction of policy authorities to the state of the economy,” meaning most interest-rate and money-aggregate movement is endogenous rather than exogenous shock – which the authors note is “what one would expect of good monetary policy” but is also precisely what makes policy’s effects hard to read off historical time series; (3) treating reserves or a monetary aggregate as moving mainly in response to policy is unreliable, since most of their movement reflects policy accommodating private-sector money demand; (4) specifications implying large real effects (like Strongin’s) tend to also imply implausible price puzzles, and correcting the price puzzle tends to shrink the implied real effects; and (5) a theoretical aside shows that a policy rule reacting too sensitively to a market long rate risks indeterminacy, by abandoning the rule’s role in anchoring the term structure (pp. 41-43).
Q8. What limitations and sources of fragility do the authors themselves flag?
Several. The identified policy disturbances “may still be attributing to policy disturbances some variation that actually originates in adverse supply shocks,” which would understate the price-reducing and overstate the output-reducing effects of a contraction (p. 58). The published results use hard-zero restrictions rather than the soft-zero priors the authors initially regarded as more plausible, because the soft-zero likelihood was highly non-Gaussian (multiple peaks) and defeated simulation of error bands; substituting hard zeros “greatly diminished” this problem “without substantial changes” to the impulse responses, but the authors note that resolving the numerical difficulties properly “may [reveal] much more statistical uncertainty in our results than this presentation would suggest” (pp. 46, 57). Most strikingly, the authors warn of “impulse response rotation”: with small changes in specification or sample, “what looked like a monetary policy shock might emerge as an I sector shock, while the shock formally identified as a monetary policy shock might cease to make sense” – a fragility that is especially acute when both the 3-month T-bill and the funds rate are included together, since term-structure arbitrage between them can weaken identification without improving accuracy (pp. 57-58). Error bands throughout are 68% Bayesian probability bands reflecting only the reference prior, not classical confidence intervals or the full range of prior information (fn. 12, p. 16). Finally, the authors note that other countries’ data (Kim 1996; Cushman-Zha for the open economy) show very small real effects and larger price effects than found here, suggesting the U.S. findings do not automatically generalize.
Key terms in this paper
Definitions below follow the paper's own usage.
- Sectoral block identification (P/I/F/B)
- this paper's non-recursive identification scheme, which groups variables into private "sluggish" variables that cannot respond within the month to financial signals (owing to planning and data lags), fast-moving auction-market "information" variables that respond contemporaneously to everything, the Federal Reserve's policy/reaction-function block, and a banking block -- used to identify a monetary policy shock without imposing a single arbitrary Cholesky ordering on all variables.
- Price puzzle (as a diagnostic here)
- the empirical pattern in which prices rise, rather than fall, following an identified monetary contraction; the paper treats the presence of a strong, sustained price puzzle (as in Strongin's reserves-market identification) not merely as a technical embarrassment but as evidence that a given specification mischaracterizes actual policy behavior and is likely confounding inflationary supply shocks with policy shocks.
- Liquidity puzzle
- in the paper's small recursive 4-variable system, the finding that the federal funds rate barely falls, and only temporarily, in response to a positive M1 innovation -- the empirical failure of the textbook liquidity effect in this specification.
- Impulse response rotation
- the paper's term for a specific fragility of structural VAR identification in which a shock that looks like a monetary policy shock under one specification or sample "rotates" into looking like an information-sector shock under a small change in specification, while the shock formally labeled "monetary policy" ceases to make economic sense -- used by the authors to explain why seemingly minor modeling choices (e.g., including both the 3-month bill and the funds rate) can produce sharply different conclusions about what monetary policy does.
- Reference prior (Sims-Zha Bayesian estimation)
- the Bayesian prior used to estimate all the VARs in this paper, which downweights parameter configurations implying large coefficients on distant lags or explosive/near-nonstationary dynamics; needed because the likelihood surface in these large, highly parameterized systems is otherwise ill-behaved, and because it produces the posterior-simulation-based 68% error bands reported throughout the paper.