Local Projections
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
How should economists trace an economy's response to a shock over the following months and years? This 2025 survey makes the case for estimating each horizon with its own ordinary regression, rather than extrapolating a single fitted time-series system forward. The authors argue it gives the same answer when that compact system is correct, but handles nonlinearities, differences across regions or firms, and policies adopted at staggered dates far more easily — and still works in some cases where the compact approach cannot identify the shock. They then walk through the practical choices: smoothing, honest error bands, and very persistent series. That matters because those choices change published conclusions.
What this paper finds — and why it matters
This 2025 Journal of Economic Literature survey by Òscar Jordà (who introduced the method in a 2005 AER paper) and Alan M. Taylor takes stock of local projections (LP) as a general framework for estimating impulse responses, arguing that LPs “help bridge the divide between current best practices in applied microeconomics and standard time series methods in macroeconomics.” The baseline LP estimates the impulse response at horizon h directly as the coefficient beta_h in a single OLS regression of y_{t+h} on the intervention s_t and lagged controls x_t, requiring one separate regression per horizon rather than iterating forward a fitted VAR; the authors identify four practical advantages over VARs – no cross-equation system constraints, straightforward accommodation of nonlinearities and heterogeneity, direct estimation of cumulative multipliers, and a natural extension to panel and difference-in-differences settings – while showing LP remains asymptotically equivalent to VAR-implied impulse responses (and to VAR-based Cholesky and long-run identification) whenever the underlying process truly is a VAR of the assumed order (citing Plagborg-Møller and Wolf 2021), and retains identification via LP-IV even in “non-invertible” settings where VAR identification from the reduced-form covariance matrix breaks down. The survey then works through the full LP toolkit: levels versus long-difference specifications (long-differencing “considerably reduces” the O(T^-1) small-sample bias that is severe for persistent, near-unit-root outcomes in short samples); horizon-by-horizon smoothing via B-splines or a Gaussian basis function; inference via Newey-West HAC standard errors, lag-augmented LP (which restores ordinary White standard errors with, per Montiel Olea and Plagborg-Møller 2021, correct uniform coverage across stationary, near-unit-root, and unit-root processes), system GMM for joint hypothesis tests across horizons, sup-t simultaneous confidence bands, and Lagrange-multiplier significance bands constructed under the null; identification via selection on observables, inverse-propensity weighting, long-run restrictions, and LP-IV; projection minimum distance (PMD) for recovering structural parameters from LP-estimated auxiliary parameters; counterfactual policy-path analysis subject to a Mahalanobis-distance “modesty” test; state-dependent and nonlinear LPs; and panel LP / LP-DiD for staggered-adoption settings. Worked illustrations throughout convey magnitudes without claiming new empirical results of their own: a state-dependent local-projection estimate of the OECD fiscal-consolidation multiplier (1978-2019) finds the average four-year output response is more than twice as large in slumps (-1.78, chi-squared(5)=58.3, p=0) as in booms (-0.80, chi-squared(5)=17.2, p=0.004); a projection-minimum-distance estimate of a UK Phillips curve (1975m1-2007m12, using the Cloyne-Hürtgen 2016 monetary shock) yields coefficients of 0.838 (SE 0.593) on inflation and -1.990 (SE 0.180) on the output term; and an LP-IV estimate of the unemployment response to a Romer-Romer monetary shock (1985:1-1999:12) peaks at roughly 1.25 percentage points around 24 months before returning to zero by about 48 months.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What is a local projection, and how does it differ mechanically from estimating an impulse response through a VAR?
A local projection (LP) estimates the impulse response at each horizon h directly as the coefficient beta_h in a single OLS regression of the future outcome y_{t+h} on the current intervention s_t and controls x_t – one separate regression per horizon, rather than fitting one VAR and iterating it forward to trace out the whole impulse-response path. Formally the impulse response is defined as the difference in expected future outcomes under an intervention of size delta versus none, R_{s->y}(h,delta) = E[y_{t+h}|s_t=s_0+delta;x_t] - E[y_{t+h}|s_t=s_0;x_t] for h=0,1,…,H (Eq. 1); the baseline LP regression y_{t+h} = alpha_h + beta_h s_t + gamma’h x_t + v{t+h} (Eq. 2) delivers beta_h as a direct estimate of that object, with v_{t+h} a moving-average error of order h by construction and H+1 separate regressions needed to trace the full horizon profile. The authors identify four practical advantages LPs have over VARs: it is a single-equation method with no cross-equation system constraints to impose; nonlinearities and heterogeneity are straightforward to accommodate; cumulative responses and multipliers are directly obtained; and the method extends naturally to panel data and difference-in-differences settings.
Q2. In what formal sense are LPs and VARs equivalent, and what can LP-IV identify that VAR-based identification cannot?
LPs and VARs are asymptotically equivalent whenever the underlying data-generating process truly is a VAR of the assumed order – citing Plagborg-Møller and Wolf (2021) – and this equivalence extends to identification: Cholesky and long-run identification schemes implemented within an LP framework are asymptotically equivalent to their VAR counterparts. For Cholesky identification, adding contemporaneous values of variables ordered earlier in the Wold causal chain as additional controls in x_t makes LP-Cholesky asymptotically equivalent to VAR-Cholesky. Where LP has a genuine edge is in “non-invertible” settings – where the number of underlying shocks exceeds the number of observables (for example, news shocks about future technology) – because “VAR identification methods based on the covariance matrix of reduced-form residuals will not work directly” there, whereas LP-IV, using an external instrument, “achieves identification even in non-invertible settings.” LP-IV was first introduced in Jordà, Schularick, and Taylor (2015).
Q3. Levels or long differences – why does this specification choice matter in finite samples?
The baseline LP can be run in levels (Eq. 11) or in long differences of the outcome (Eq. 12), and the levels specification carries a small-sample bias of order O(T^-1) that becomes considerable for persistent outcomes (autoregressive root rho near 1) combined with a small sample size T. Monte Carlo experiments (T=100, rho in {0.95, 1.0}, Figure 1) illustrate that “small sample biases appear to be considerably reduced when using long differencing” instead of levels – a practical recommendation rather than a claim that long-differencing dominates in every setting.
Q4. How are cumulative multipliers estimated from a set of local projections?
The cumulative multiplier at horizon h is the ratio of the cumulative outcome response to the cumulative response of the intervention itself, m(h) = R^c_{sy}(h)/R^c_{ss}(h) (Eq. 16), and when estimation uses LP-IV with a single instrument z_t, this cumulative multiplier can be recovered from a single regression (Eq. 19) that instruments the cumulative treatment path with z_t, rather than requiring separate numerator and denominator regressions at each horizon.
Q5. Why and how are local projections smoothed across horizons?
Horizon-by-horizon LP estimates can be noisy, so the survey describes two ways to impose a parametric functional form on beta_h as a function of h: B-splines (piecewise polynomials with continuity constraints at the knots, following Eilers and Marx 1996) and a three-parameter Gaussian basis function (GBF), phi(h;Theta) = a*exp{-[(h-b)^2/c^2]}, where a is the amplitude, b the location of the peak response, and c governs its width. GBF parameters are estimated jointly within the same GMM system used for inference; a worked example fits a = 1.388 (SE 0.278), b = 26.189 (SE 0.922), c = 12.997 (SE 0.891) to an unemployment response to a Romer-Romer shock (1985:1-1999:12), implying a peak response around month 26.
Q6. Why do standard OLS standard errors fail for LP coefficients, and what are the two leading pointwise fixes?
Because the LP error v_{t+h} is a moving average of order h by construction, it is serially correlated in a way that standard OLS standard errors ignore, so the survey recommends either Newey-West HAC standard errors with J = h+p lags (Bartlett kernel, p the LP’s own lag length) or “lag-augmented” LP, which adds extra lags of the outcome variable to the regression (Eq. 28) so that ordinary White heteroskedasticity-robust standard errors apply. Following Montiel Olea and Plagborg-Møller (2021), lag-augmented LP is reported to give “correct uniform probability coverage under a wide range of scenarios (stationarity, near unit roots, nonstationarity) and even for long-distance horizons (as long as the sample size is large enough relative to the horizon).”
Q7. Beyond single-horizon standard errors, how does the survey handle joint and simultaneous inference across the whole impulse-response path?
For testing hypotheses across multiple horizons at once, the survey estimates the full set of LP coefficients as a system by GMM (Eqs. 30-33), which nests standard 2SLS/IV as the single-horizon special case; for simultaneous confidence bands that keep joint coverage at the nominal level (rather than the pointwise 1.96-SE band, which “understates actual joint coverage”), it uses the sup-t bands of Montiel Olea and Plagborg-Møller (2019), computed by simulating from the joint normal distribution of all H+1 coefficients – bands the authors describe as “relatively conservative, since they accommodate unspecified nulls.” A third tool, significance bands (Inoue, Jordà, and Kuersteiner, forthcoming), is constructed under the null hypothesis beta_h=0 via the Lagrange-multiplier principle – analogous to the standard +/-1.96/sqrt(n) bands used for autocorrelation functions – and, unlike error bands, straddles zero rather than the point estimate; the authors suggest “current practice could be extended to display significance bands alongside error bands.” An empirical illustration (price level and inflation responses to an extended Romer-Romer shock, 1969:I-2007:IV) reports joint-significance test p-values of 1.72e-23 for the price level and 1.12e-17 for inflation.
Q8. What identification strategies does the survey lay out for LPs beyond plain OLS?
The survey organizes LP identification into four approaches: LP-OLS (“selection on observables”), which assumes variation in the intervention s_t is as good as random conditional on controls x_t; LP-IPW, which reweights by the estimated propensity score for a binary intervention (with a doubly robust LP-IPWRA variant adding regression adjustment); long-run restrictions, where Plagborg-Møller and Wolf (2021) show Blanchard-Quah-style long-run identification can be implemented as a two-step LP procedure; and LP-IV, which uses an external instrument and requires relevance and lead-lag exogeneity (Assumption 1). As noted in Q2, LP-IV’s distinguishing advantage is that it remains valid in non-invertible settings where covariance-matrix-based VAR identification does not.
Q9. How does the survey extend LPs to structural estimation and to evaluating whether observed policy is optimal?
Projection minimum distance (PMD, Jordà and Kozicki 2011) estimates structural parameters theta = g(pi) by minimum distance from LP-estimated auxiliary parameters pi-hat, minimizing (theta - g(pi-hat))‘W_T(theta - g(pi-hat)) to get an asymptotically normal theta-hat – a method the authors note “does not require the user to use LPs to estimate the impulse responses” in the first place, and which extends to rational-expectations and DSGE models. A worked illustration estimates a UK Phillips curve (using the Cloyne-Hürtgen 2016 monetary policy shock, residualized against UK oil-price and exchange-rate controls, sample 1975m1-2007m12) with PMD point estimates theta-hat_pi = 0.838 (SE 0.593) and theta-hat_x = -1.990 (SE 0.180). Building on the same PMD logic, Barnichon and Mesters (2023) use LP-estimated impulse responses to test whether a policymaker minimizing a quadratic loss over inflation and unemployment deviations is at its optimum (H0: delta=0), with the notable feature that “nowhere in the discussion did we have to explicitly write down the policy rule.”
Q10. How do LPs handle state-dependence, nonlinearities, counterfactual policy paths, and panel/difference-in-differences settings?
For a predetermined binary state D_{t-r}, the survey recommends estimating separate state-specific long-difference LPs (Eq. 64) rather than pooling, since “the correct approach is to condition on the state as well, and to estimate a state-dependent LP”; a worked fiscal-consolidation illustration (OECD, 1978-2019, treatment = change in the cyclically adjusted primary balance, instrumented with the Guajardo et al./Adler et al. narrative measure) finds the average output response over years 0-4 is -1.78 (chi-squared(5)=58.3, p=0) in economic slumps versus -0.80 (chi-squared(5)=17.2, p=0.004) in booms, i.e., “much larger when fiscal consolidations are implemented during slumps.” The survey cautions this state-dependent approach breaks down “if the state includes current information,” since then the intervention and the state are mutually influenced and instruments may be needed for both. A related but distinct nonlinear-LP specification (Eq. 67, following Cloyne, Jordà, and Taylor 2023) decomposes the total response into a direct effect and an indirect “state-amplification” effect that varies continuously with x_t rather than switching between two discrete regimes. Separately, counterfactual policy-path analysis (following Leeper and Zha’s “modest policy interventions,” 2003) uses the joint normality of LP coefficient estimates to compute a counterfactual response to an alternative policy path (Eq. 61), valid only when a Mahalanobis-distance “modesty test” (Eq. 63) is not too large; a worked example (funds-rate GBF parameters for 1985:1-2000:12) shifting the peak-timing parameter by one standard deviation yields a borderline modesty p-value of 0.08, and the authors caution that results in such cases “should be interpreted with some caution.” Finally, panel LP (Eq. 73) adds individual and time fixed effects (absorbing the former by long-differencing) and uses Driscoll-Kraay standard errors as T->infinity or a wild cluster bootstrap as N->infinity for small T; LP-DiD (Dube et al. 2023) adapts this to staggered-adoption settings by restricting the estimation sample to newly treated or not-yet-treated (“clean control”) observations, which the authors note is “trivial” to impose in LP’s forward-looking regression framework compared with algorithmic pairwise comparisons across control units in other DiD estimators.
Key terms in this paper
Definitions below follow the paper's own usage.
- local projection (LP)
- a single-equation regression method that estimates the impulse response at horizon h directly as the coefficient beta_h in y_{t+h} = alpha_h + beta_h s_t + gamma'_h x_t + v_{t+h} (Eq. 2) -- one separate regression per horizon -- rather than reading impulse responses off an iterated, fitted VAR.
- LP-VAR equivalence
- the result (Plagborg-Møller and Wolf 2021) that LP-estimated impulse responses, and LP-implemented Cholesky and long-run identification schemes, are asymptotically equivalent to their VAR counterparts whenever the data truly follow a VAR of the assumed order -- distinguished in this survey from the "non-invertible" settings in which that equivalence, and VAR identification itself, breaks down.
- LP-IV / non-invertibility
- instrument-based LP identification (Assumption 1: relevance and lead-lag exogeneity) that "achieves identification even in non-invertible settings," i.e., when the number of underlying structural shocks exceeds the number of observables, unlike VAR identification methods that rely on the reduced-form residual covariance matrix.
- projection minimum distance (PMD)
- the two-step estimator of Jordà and Kozicki (2011) that treats LP-estimated coefficients as auxiliary parameters pi and recovers structural parameters theta = g(pi) by minimum distance, explicitly not requiring that the LPs used to estimate pi be structural themselves.
- significance bands
- confidence bands for the LP impulse-response path constructed under the null hypothesis beta_h = 0 via the Lagrange-multiplier principle (Inoue, Jordà, and Kuersteiner, forthcoming), so that they straddle zero rather than the point estimate -- a deliberate complement to ordinary pointwise error bands and sup-t simultaneous confidence bands, which straddle the estimate itself.