Macro Paper Warehouse
Published Classic [Journal of Economic Dynamics and Control] doi:10.1016/j.jedc.2024.104999

Modeling inflation expectations in forward-looking interest rate and money growth rules

Zhengyang Chen

Victor J. Valcarcel

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Which indicator best captures what monetary policy is doing — a short-term interest rate or a measure of money? This 2025 paper embeds forward-looking expectations in a small three-equation system and runs it across 241,865 parameter combinations on monthly United States data from 1988 to 2020, counting how often results come out with the theoretically wrong sign. With a shadow interest rate, output responses are wrong in about 99 percent of combinations; with a broad weighted money measure, about 4 percent. Why it matters: conclusions about policy may hinge on the indicator chosen, though validity rests on the assumed three-equation structure.

What this paper finds — and why it matters

This 2025 Journal of Economic Dynamics and Control paper by Zhengyang Chen and Victor J. Valcarcel proposes a “rational expectations structural VAR” (RE-SVAR) — a way to embed rational expectations directly into a low-dimensional structural VAR without mapping the system to a fully specified DSGE model — and uses it to compare two candidate monetary-policy indicators, the Wu and Xia (2016) shadow federal funds rate and Divisia M4 money growth, on which produces a larger share of theoretically sensible (“non-puzzling”) impulse responses. The model starts from a three-equation consensus New Keynesian system (a forward-looking monetary policy rule written in either an interest-rate or a money-growth form, an IS equation, and a Phillips-curve AS equation, all estimated simultaneously) and identifies the monetary policy shock using a rational-expectations forecast-revision restriction that recovers it as a linear combination of reduced-form VAR residuals — not a Cholesky/recursive scheme, and without imposing any delayed-reaction exclusion restriction on the policy indicator. Rather than estimating the policy rule’s forward-looking parameters, the paper runs a “pseudo-calibration” grid search over the inflation- and output-response coefficients (φ_π, φ_y, each cycling over 61 values from 0 to 4) and the horizons over which expectations are formed (h_π = 0,…,12 months for inflation, h_y = 0,…,4 months for output), yielding 241,865 structural VAR specifications (8 lags, monthly U.S. data) and a resulting “cloud” of impulse responses; a response counts as a “puzzle” if output or inflation turns negative at any point within the first year following an expansionary shock to the policy indicator. In the main October 1988-February 2020 sample (377 monthly observations), the shadow federal funds rate produces output puzzles in 98.68% of the 241,865 specifications and inflation puzzles in 99.13%, with only 0.87% (2,109 specifications) surviving a no-joint-puzzle criterion; replacing it with Divisia M4 growth as the policy indicator cuts output puzzles to 4.02% and inflation puzzles to 4.13%, with 95.85% (231,825 specifications) surviving. This Divisia advantage holds up under a coarser 25,137-specification grid, a post-Global-Financial-Crisis effective-lower-bound sample (December 2008-February 2020, where even the shadow rate’s best-case output and inflation puzzle rates remain high at 72% and 93%), a long historical sample back to January 1967, a narrower Divisia M2 aggregate, and a PCE-based inflation measure — with one partial exception: in the long historical sample under PCE, the output-puzzle gap between indicators narrows substantially (53.3% for the shadow rate vs. 56.0% for DM4 vs. 47.9% for DM2). Longer inflation-expectation horizons help Divisia far more than the shadow rate (at a 12-month horizon, 18,430 of 241,865 DM4 specifications are non-puzzling versus only 5 for the shadow rate), and extending the model to four variables by adding the Gilchrist-Zakrajsek excess bond premium (July 1979-February 2020 sample) still yields an 81.45% no-joint-puzzle survival rate with DM4 as the indicator, alongside IS- and AS-shock responses that are broadly consistent with textbook signs. The authors read the results as evidence that the federal funds rate — especially its shadow-rate extension through the effective lower bound — is an inadequate standalone policy indicator in low-dimensional VARs, and that Divisia M4’s richer information content lets it capture forward-looking central bank behavior without added information variables, while cautioning that the RE-SVAR’s validity is conditional on the sensibility of its underlying three-equation theoretical structure and that, unlike a standard recursive VAR, it is not modular: adding any variable (as they do, only loosely, for the excess bond premium) requires specifying a full structural equation for it.

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What problem motivates the paper, and what exactly is it testing?

The paper asks whether rational expectations can be built directly into a small structural VAR — without first mapping the system to a fully specified DSGE model — and then uses that framework as a testbed to compare two candidate monetary policy indicators: the Wu and Xia (2016) shadow federal funds rate versus Divisia M4 nominal money growth. The motivation is a documented empirical problem with the federal funds rate (and its shadow-rate extension through the effective lower bound) in low-dimensional recursive VARs: it tends to produce “puzzling” impulse responses (e.g., inflation rising after a contractionary shock), a pattern the authors say is “amply documented in Ramey (2016) and references therein.” Standard fixes add more information variables to the VAR, but commodity prices, futures data, and similar variables have no accepted structural equation in the consensus New Keynesian model the authors want to preserve.

Q2. How does the RE-SVAR’s identification scheme differ from a standard recursive (Cholesky) SVAR?

Identification does not rely on a Cholesky ordering or on any delayed-reaction exclusion restriction on the policy indicator — even though the indicator is still ordered first in the system. Instead, the authors use a rational-expectations forecast-revision relationship, E_t v_{t+j} − E_{t-1} v_{t+j} = S_v Ψ^j D e_t, to express the monetary policy shock as a specific linear combination of the reduced-form VAR’s residuals (e.g., ε_t^MP = e_t^i − (φ_π S_π Ψ^{h_π} D e_t)′ − (φ_y S_y Ψ^{h_y} D e_t)′ for the interest-rate rule, and an analogous expression for the money-growth rule). This recovers the structural shock without ever estimating the structural parameters themselves, and the approach is explicitly frequentist rather than Bayesian.

Q3. What is the underlying theoretical structure of the RE-SVAR?

The model is a three-equation “consensus” New Keynesian system, estimated simultaneously: a forward-looking monetary policy rule (written either as a Taylor-type interest-rate rule or as a money-growth rule), a forward-looking IS equation, and a Phillips-curve AS equation. The policy rule responds to expected future inflation and output (E_t π_{t+h_π}, E_t y_{t+h_y}) rather than to current or lagged values, which is what lets the RE-SVAR be “forward-looking” without adding extra observable variables to the system. The paper treats this three-equation structure as a deliberate simplification: it is not modular in the way a standard reduced-form VAR is, since each additional variable would need its own fully specified structural equation.

Q4. How does the paper handle the policy rule’s forward-looking parameters, and what counts as a “puzzle”?

Rather than estimating the policy rule’s response coefficients and expectation horizons, the paper conducts a “pseudo-calibration” grid search across 241,865 combinations of them — 61 values each of φ_π and φ_y (0 to 4, in 1/15 increments), 13 inflation-expectation horizons (h_π = 0,…,12 months), and 5 output-expectation horizons (h_y = 0,…,4 months) — producing a “cloud” of structural impulse responses rather than a single point estimate. A response is counted as a “puzzle” if, following an expansionary shock to the policy indicator, the output or inflation response goes negative at any point within the first year afterward; a “joint” puzzle is recorded if either variable (or both) misbehaves this way. The lag length is fixed at 8 (AIC-selected for the shadow-rate specification) across all specifications.

Q5. What are the headline results comparing the shadow federal funds rate to Divisia M4?

In the main October 1988-February 2020 sample, the shadow federal funds rate produces output puzzles in 98.68% of the 241,865 specifications and inflation puzzles in 99.13%, so only 0.87% (2,109 specifications) survive a no-joint-puzzle criterion; switching the policy indicator to Divisia M4 growth cuts output puzzles to 4.02% and inflation puzzles to 4.13%, so 95.85% (231,825 specifications) survive. The gap between the two indicators is thus very large across virtually the entire grid of plausible forward-looking policy-rule parameterizations, not just at a single calibrated point.

Q6. How robust is the Divisia-over-shadow-rate finding to alternative samples, price indices, and monetary aggregates?

The finding holds up across a coarser 25,137-specification grid, a post-Global-Financial-Crisis effective-lower-bound sample (December 2008-February 2020), a long historical sample back to January 1967, a narrower Divisia M2 aggregate, and a PCE-based inflation measure — with one partial exception. Even in the shadow rate’s best-performing subsample (the post-GFC ELB period), its output and inflation puzzle rates remain high, at 72% and 93% respectively, versus low single digits for the Divisia aggregates in the corresponding comparisons. The exception is the long historical sample (1967-2020) under PCE inflation, where the output-puzzle advantage of Divisia over the shadow rate narrows considerably: 53.3% for the shadow rate versus 56.0% for DM4 and 47.9% for DM2. The authors also report that extending the sample through April 2022 yields qualitatively similar conclusions, but flag that COVID-period estimates may be severely distorted and so exclude that period from the main analysis.

Q7. Does the assumed inflation-expectation horizon matter for these results?

Yes — longer inflation-expectation horizons substantially help Divisia M4 but barely move the shadow rate’s puzzle incidence. At a 6-month inflation-expectation horizon, 17,973 of 241,865 DM4 specifications are non-puzzling compared with only 195 for the shadow rate; at a 12-month horizon the gap widens further, to 18,430 non-puzzling DM4 specifications versus just 5 for the shadow rate. This pattern is consistent with the paper’s broader claim that Divisia money growth is better able to embed forward-looking policy behavior than the federal funds rate.

Q8. What happens when the model is extended to four variables with a credit-spread measure?

Adding the Gilchrist-Zakrajsek (2012) excess bond premium (EBP) as a fourth variable, using Divisia M4 as the policy indicator over a July 1979-February 2020 sample (dictated by EBP data availability), still yields a high no-joint-puzzle survival rate of 81.45% (196,988 of 241,865 specifications). In this extended system, median responses to IS shocks show industrial production and CPI both rising, while median responses to AS shocks show industrial production falling and CPI rising — patterns the authors describe as broadly consistent with textbook New Keynesian predictions. The EBP’s own structural equation, however, is specified in a deliberately unrestricted form, since the authors note there is “a lack of theoretical foundation for a law of motion” for the excess bond premium.

Q9. What limitations and caveats do the authors themselves flag?

The authors flag four main limitations: the RE-SVAR is non-modular, its validity is conditional on the underlying three-equation theory being sensible, its frequentist grid-search approach does not report cloud-median responses (following Inoue and Kilian’s 2022 critique of doing so from a Bayesian perspective), and the excess bond premium’s structural equation in the four-variable extension is underspecified for lack of theory. On non-modularity, they write that adding variables to the RE-SVAR requires a fully specified structural equation for each one, unlike a standard recursive VAR where variables can simply be appended — an explicit disadvantage relative to modular n-variable recursive SVARs. On the theoretical dependency, they state plainly that “if the theoretical model we construct our RE-SVAR from is not sensible, it renders the whole enterprise a non-starter.” The sample also excludes the COVID period, on the grounds that pandemic-era estimates may be severely distorted even though results extending through April 2022 are reported as qualitatively similar.

Key terms in this paper

Definitions below follow the paper's own usage.

RE-SVAR (rational expectations structural VAR)
the paper's proposed identification framework — a structural VAR that embeds rational-expectations forecast-revision restrictions directly into a low-dimensional VAR system to recover a monetary policy shock as a linear combination of reduced-form residuals, without mapping the system to a full DSGE model and without relying on a Cholesky ordering or delayed-reaction exclusion restriction.
Pseudo-calibration grid search ("cloud" of responses)
instead of estimating the policy rule's forward-looking response coefficients (φ_π, φ_y) and expectation horizons (h_π, h_y), the paper fixes a grid of values for each and computes a structural impulse response for every combination — 241,865 in the main specification — yielding a distribution ("cloud") of responses whose puzzle incidence is then tabulated across the whole grid rather than at one calibrated point.
Puzzle criterion
the paper's specific empirical test for a "sensible" impulse response — following an expansionary shock to the policy indicator, any output or inflation response that turns negative at any point within the first year afterward counts as a puzzle; a "joint" puzzle is recorded whenever either variable (or both) puzzles.
Divisia M4 (DM4)
a monetary aggregate, constructed by the Center for Financial Stability, that weights component monetary assets by their user cost rather than summing them dollar-for-dollar; used here as the money-growth policy indicator, μ_t = m_t − m_{t-1} + π_t, where m_t is the log real Divisia M4 balance.
Non-modularity
the paper's own acknowledged structural limitation — because every variable in the RE-SVAR must correspond to a fully specified structural equation, variables cannot simply be appended to the system the way they can in a standard reduced-form or recursive VAR, making extensions (such as the four-variable version with the excess bond premium) require new theoretical justification each time.
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.