The long and variable lags of monetary policy: Evidence from disaggregated price indices
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
Why does the overall price level take years to respond to a central bank tightening? Splitting United States consumer prices into as many as 136 categories over 1982 to 2008, this paper finds no uniform sluggishness: roughly half the individual price series rise significantly at some point after a tightening, and many never turn down. The overall index shows no significant decline for about 43 months, then falls to a trough near 4 percent around 54 months. Why it matters: the long delay reflects offsetting movements across sectors, which standard accounts based on infrequent price changes do not explain.
What this paper finds — and why it matters
This 2024 Journal of Monetary Economics paper by S. Borağan Aruoba and Thomas Drechsel asks why the aggregate U.S. price level responds to monetary policy only after a long delay, and locates the answer in the cross-section of price categories rather than in uniform price stickiness. Using monthly Personal Consumption Expenditures Price Index (PCEPI) data disaggregated into as many as 136 subcomponents at five levels of aggregation (1982-2008, a sample chosen to exclude the zero lower bound and COVID episodes), the authors estimate a Jordà-style local projection for each individual price series against an external, non-market-based monetary policy shock constructed in a companion paper (Aruoba and Drechsel 2023) from natural-language processing of Federal Reserve Greenbook/Tealbook documents, following the Romer-Romer (2004) narrative approach. Because the individual-series impulse responses are estimated separately, the authors re-aggregate them into an implied aggregate PCEPI response using a weighted sum of the disaggregated coefficients, and correct the standard error of that re-aggregated response for cross-sectional correlation across price series with a seemingly-unrelated-regressions/feasible-generalized-least-squares (SUR-FGLS) procedure — showing that because price movements are mostly positively correlated across categories, naive standard errors that ignore this covariance are too small and would spuriously reject a zero response earlier than the correct SUR-corrected bands do. The headline empirical result is that, following a 100-basis-point monetary contraction, the aggregate (and core) PCEPI shows no statistically significant decline for roughly 43 months, only turning significantly negative after that point and reaching a peak reduction of roughly 4% (headline) and roughly 2.5% (core) around 54 months — consistent with Christiano, Eichenbaum, and Evans’s (1999) finding of a multi-year delay before the aggregate price level falls significantly. Critically, the paper shows this aggregate flatness does not reflect uniform stickiness: at the finest (136-series) level of disaggregation, roughly half of all price series show a significantly positive response to a monetary tightening at some horizon, and a substantial share never turn significantly negative at all, so the eventual aggregate decline emerges only once enough individual series’ responses turn negative to outweigh the persistently positive ones — a pattern the authors argue is inconsistent with standard Calvo or menu-cost stickiness models (which predict faster responses for more frequently adjusting prices, a relationship the data does not show) and better rationalized by a heterogeneous cost-channel/demand-substitution mechanism in which a monetary tightening acts as a demand shock for some sectors but raises costs (and hence prices) for financially constrained sectors in the short run. The authors also show that re-aggregating with 1959 versus 2023 expenditure shares yields similar aggregate timing and magnitude, indicating that decades of structural change in consumption patterns have not accelerated the aggregate lag — leading them to connect the finding explicitly to Friedman’s (1960, 1961) claim that monetary policy operates with “long and variable lags,” here documented in the cross-section of prices as well as in the aggregate time series.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What question does this paper ask, and how does it reframe Milton Friedman’s claim that monetary policy works with “long and variable lags”?
The paper asks why the aggregate U.S. price level takes so long to respond to a monetary policy shock, and argues that Friedman’s (1960, 1961) “long and variable lags” doctrine applies across the cross-section of individual prices, not only to the aggregate time series. Using highly disaggregated Personal Consumption Expenditures Price Index (PCEPI) data, the authors document “stark differences in the timing and magnitude of the responses across price categories, including some prices that show an initially positive response” to a monetary tightening, and use this heterogeneity to explain why the aggregate response looks slow and muted even though individual prices are often moving substantially.
Q2. What data and monetary policy shock does the paper use, and why avoid a market-based surprise measure?
The analysis uses monthly PCEPI subcomponents for the United States from 1982 to 2008 — a sample chosen to exclude the zero lower bound, non-standard monetary policy, and the COVID period — disaggregated into as many as five levels of granularity, down to 136 individual price categories at the finest level (level 5), which form a balanced panel. The identified shock is the external instrument constructed in the authors’ companion paper, Aruoba and Drechsel (2023): a natural-language-processing/machine-learning measure applied to the Federal Reserve’s Greenbook (Tealbook) forecasting documents, following the narrative logic of Romer and Romer (2004), designed to back out exogenous variation in the federal funds rate target rather than relying on market interest-rate surprises around FOMC announcements — the latter avoided because such surprises can reflect “information effects” or responses to non-standard policy rather than a clean policy shock.
Q3. How is the impulse response of an individual disaggregated price series estimated?
Each price component i is analyzed with a separate local projection (following Jordà 2005) at every horizon h from 0 up to 60 months, regressing the log price h periods ahead on the contemporaneous monetary policy shock, a control vector, and the lagged log price level; the coefficient on the shock at each horizon is the impulse response, scaled to a 100-basis-point policy tightening. The control vector for each series is chosen by a combinatorial search over macroeconomic variables and their first 10 lags — selecting up to 7 controls that maximize fit at the 24-month horizon, a procedure the authors note runs on the order of 2.4 billion candidate regressions per price series — and standard errors are computed with Newey-West HAC correction using a bandwidth equal to h+1.
Q4. How are the individual-series impulse responses re-aggregated into an implied aggregate PCEPI response, and why does the standard error correction matter?
Because the aggregate PCEPI is approximately an expenditure-share-weighted sum of the subcomponent price levels, the re-aggregated impulse response at each horizon is simply the same weighted sum of the individual-series coefficients — but because price responses are correlated across categories, the authors estimate a Seemingly Unrelated Regressions (SUR) system for each horizon and use a feasible generalized least squares (FGLS, via GMM with a Cholesky factorization) estimator to obtain the correct joint covariance matrix. The correctly computed SUR standard error of the aggregate response includes cross-series covariance terms that a naive standard error (built only from the individual OLS variances) ignores; since price movements are mostly positively correlated, the naive approach understates the true standard error and “would reject a zero response early on” at the 90% confidence level where the SUR-corrected bands do not — meaning the correction changes the substantive conclusion about when the aggregate response becomes significant, not just its precision.
Q5. What does the aggregate PCEPI response to a monetary tightening actually look like?
Following a 100-basis-point contractionary shock, the point estimate for headline PCEPI is essentially flat (even slightly positive, though not statistically distinguishable from zero at the 90% confidence interval) for roughly the first three years, turning significantly negative only after about 43 months and reaching a peak reduction of roughly 4% around 54 months; core PCEPI (excluding food and energy) follows a similar timing but with a smaller peak reduction of roughly 2.5%, estimated more precisely. This pattern reproduces Christiano, Eichenbaum, and Evans’s (1999) finding that “the aggregate price level does not fall to a statistically significant extent for 4 years after an FFR increase”; the authors themselves flag that these specific magnitudes should ultimately be checked against their Figure 1 rather than treated as exact to the decimal.
Q6. How heterogeneous are the responses once prices are disaggregated, and what is the paper’s central interpretation of that heterogeneity?
Disaggregating further reveals sharply different response profiles: at the 2-category level, durable goods prices decline fastest (significant after about 2.5 years, peaking around a 5% reduction near 4 years) while services move more slowly and NPISH (nonprofit institution) prices show no significant response at all; at the finest 136-series level, roughly 13% of series are never significant in either direction, and fully half of the series show a significantly positive response to the tightening at some horizon. The authors’ central interpretive claim is that “the reason why the headline and core response is flat and fall after 4 years is not because all prices are unchanged for 4 years” — instead, many individual prices move substantially in both directions, and the aggregate only turns clearly negative once enough series’ declines outweigh the persistently positive ones; used motor vehicle prices, for example, rise initially because higher rates suppress new-vehicle purchases and push substitution demand toward used vehicles.
Q7. Which price categories actually drive the eventual aggregate decline?
A contribution-level decomposition — which weights each subcomponent’s price change by its expenditure share — shows that a small number of categories account for most of the eventual aggregate movement, with motor vehicle fuels/lubricants/fluids as the single largest contributor, alongside food and nonalcoholic beverages purchased off-premises, hospital and nursing-home services, financial services, housing, and household utilities. Consistent with the aggregate pattern, almost all of these contribution-level impulse responses are flat for the first few years and only turn strongly negative after roughly 40 months.
Q8. Is this cross-sectional pattern consistent with standard sticky-price models, and what mechanism do the authors favor instead?
The data show no meaningful negative relationship between how quickly a price category’s response to the shock peaks and how frequently that category’s prices adjust empirically — if anything the relationship is weakly positive, the wrong sign for a Calvo-style sticky-price model, which predicts more frequently adjusting prices should respond faster. Menu-cost models fare no better, the authors argue, since in that framework a firm that does adjust must move its price in the direction implied by the money-supply change, which cannot generate the positive responses observed in about half the series; instead, the authors favor a heterogeneous cost-channel mechanism (building on Barth and Ramey 2001 and Ravenna and Walsh 2006) combined with demand substitution across goods, in which a tightening acts as a demand shock for some sectors but raises marginal costs — and hence prices — for financially constrained sectors in the short run, with Baumeister, Mumtaz, and Peersman’s (2013) multi-sector sticky-price-with-cost-channel model described as coming “closest to what we envision as a suitable framework.”
Q9. Does the paper’s finding hold up to changes in expenditure weights, and what are the main caveats?
Re-aggregating the disaggregated impulse responses using 1959 expenditure shares instead of 2023 shares yields a similar aggregate response in both magnitude and timing, even though the underlying consumption basket shifted substantially over that period (the food-and-beverages share fell by about 12 percentage points while the health-care share rose by about 11 points) — indicating that decades of structural change in what households buy have not accelerated the economy’s aggregate inflation lag. The authors are explicit about scope limits: the 1982-2008 sample excludes the zero lower bound, COVID, and the 2022-2023 tightening cycle, so the results may not generalize to those episodes; the underlying shock series itself comes from a companion paper’s NLP/ML methodology that treats Fed decisions as exogenous, even though the authors note elsewhere that the 2022-2023 decisions were largely systematic rather than exogenous; and a stylized calculation of the price reduction implied by applying their estimates to the 2022-2023 hiking cycle is explicitly described as “intentionally provocative” and “not a serious policy counterfactual.”
Key terms in this paper
Definitions below follow the paper's own usage.
- Local projection (per price series)
- in this paper, a separate OLS regression estimated for each individual PCEPI subcomponent and each horizon h (0 to 60 months), regressing the h-months-ahead log price on the contemporaneous monetary policy shock, a series-specific set of controls chosen by combinatorial search, and the lagged log price — used here as the estimation engine that generates hundreds of horizon-by-horizon impulse responses that are subsequently re-aggregated, rather than as a single aggregate-variable exercise.
- SUR-FGLS re-aggregation
- the paper's procedure for combining the individually estimated, correlated price-level impulse responses into a single aggregate PCEPI impulse response with a correctly sized standard error — a Seemingly Unrelated Regressions system is specified at each horizon and estimated by feasible generalized least squares (via GMM with a Cholesky factorization) so that the reported standard error of the aggregate response incorporates the cross-series covariances that a naive weighted-sum-of-variances calculation would omit.
- Aruoba-Drechsel monetary policy shock
- the external instrument used throughout, constructed in the authors' companion paper (2023) by applying natural-language-processing/machine-learning methods to Federal Reserve Greenbook (Tealbook) documents in the spirit of Romer and Romer's (2004) narrative approach; distinguished in this paper from market-based high-frequency surprise measures because it is meant to isolate exogenous variation in the funds-rate target without embedding information effects.
- Cross-sectional "long and variable lags"
- the paper's reframing of Friedman's (1960, 1961) doctrine — instead of (or in addition to) the aggregate price level responding to policy only after a long, variable delay, individual price categories respond at markedly different speeds and even in different directions, so the long aggregate lag is a composition effect of heterogeneous, sometimes offsetting, sectoral responses rather than evidence that all prices are simply slow to move.
- Heterogeneous cost-channel mechanism
- the authors' preferred explanation for why some prices rise (rather than fall) after a monetary tightening — in sectors exposed to financing constraints, a tightening raises firms' marginal costs (a "cost channel," per Barth and Ramey 2001 and Ravenna and Walsh 2006) and can push prices up in the short run even as it acts as a demand-dampening (price-reducing) shock elsewhere, with demand substitution across goods (e.g., toward used vehicles) reinforcing some of the positive responses.