Macro Paper Warehouse
Published Classic [Journal of Econometrics] doi:10.1016/j.jeconom.2024.105722 Vol. 244, No. 2, pp. 105722

Local Projections vs. VARs: Lessons from Thousands of DGPs

Dake Li

Mikkel Plagborg-Møller

Christian K. Wolf

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Two standard ways of tracing how the economy responds to a shock — running a separate regression for each future date, or fitting one dynamic system and projecting it forward — disagree, and real data cannot settle which is better because the truth is unknown. So this 2024 paper generates 6,000 artificial economies calibrated to 207 United States series. The separate-regression route carries less bias at distant horizons but far more noise, roughly double by five years out, so the system route usually wins overall. Why it matters: applied researchers get a defensible default, though these are simulation lessons rather than theorems.

What this paper finds — and why it matters

This 2024 Journal of Econometrics paper by Dake Li, Mikkel Plagborg-Møller, and Christian K. Wolf asks a purely practical question rather than proposing a new estimator or identification scheme: when researchers estimate structural impulse responses, does the local-projection (LP) estimator or the vector-autoregression (VAR) estimator perform better in realistic macroeconomic settings, and under what conditions does the ranking flip? Because no single real-world dataset can answer this — the true DGP is never known — the authors build an “encompassing model,” a non-stationary dynamic factor model with six latent factors estimated on the 207-series Stock and Watson (2016) quarterly U.S. dataset (1959Q1-2014Q4), and use it to generate 6,000 simulated economies (3,000 built around a monetary policy shock with the federal funds rate as instrument, 3,000 around a fiscal policy shock with government spending as instrument), each simulated for T=200 quarters with 5,000 Monte Carlo draws. Across this population of realistically calibrated DGPs, comparing least-squares, bias-corrected, and penalized LP against least-squares, bias-corrected, Bayesian, and model-averaged VARs (plus SVAR-IV for the instrumented case), the paper documents a clear and pervasive bias-variance trade-off: at short horizons (h ≤ the p=4 lag length) LP and VAR have similar bias, but at longer horizons VAR bias grows substantially larger than LP bias, while LP’s standard deviation rises steeply with horizon — by h=20 roughly double the VAR’s. Bias-corrected LP removes only about a third of LP’s bias while adding variance, so it is preferred over uncorrected LP only when a researcher places very high weight (ω ≥ 0.9 in the paper’s bias-variance loss function) on bias at intermediate horizons; otherwise VAR-type methods, especially a Bayesian VAR with a Minnesota-type prior, dominate essentially throughout. A further headline finding concerns SVAR-IV: because roughly 90% of the simulated DGPs exhibit a degree of shock invertibility below 49%, the external-instrument SVAR-IV estimator carries substantially higher bias than internal-instrument alternatives at every horizon, though it also has notably lower dispersion. The authors are explicit that these are simulation-based lessons conditional on the choice of encompassing model and DGP class (quarterly, five-variable systems, no sign or long-run restrictions), not universal theorems.

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What question does the paper ask, and why does it need thousands of simulated DGPs rather than a single empirical application to answer it?

The paper asks a purely practical, “horse race” question — does the local-projection (LP) or the VAR estimator deliver more accurate structural impulse responses in realistic macroeconomic settings, and under what conditions does the answer change — rather than proposing a new identification scheme or theoretical result. Answering this with a single real dataset is impossible because the true data-generating process (DGP) is never known, so any one empirical comparison confounds estimator performance with the idiosyncrasies of that particular sample. The authors instead build a large, realistically calibrated population of DGPs and measure each estimator’s bias and variance against the known true impulse response in every one of them (Abstract; Introduction, p. 1).

Q2. How are the thousands of DGPs actually constructed?

The authors estimate one “encompassing” non-stationary dynamic factor model (DFM) on the Stock and Watson (2016) quarterly U.S. dataset (207 series, 1959Q1-2014Q4) — six latent factors following a non-stationary VECM with a VAR(4) representation, plus idiosyncratic AR(4) components for each series — and then draw 6,000 five-variable sub-systems from it (Section 3.1, p. 5; Section 3.3, p. 7). Half (3,000) always include the federal funds rate as the policy instrument (“monetary policy DGPs”); the other half always include federal government spending (“fiscal policy DGPs”). Each simulated economy runs for T=200 quarters with 5,000 Monte Carlo draws (Section 5, p. 10). The resulting DGPs are deliberately heterogeneous: the median IV first-stage F-statistic is 21.13 (10th-90th percentile range 7.91-33.29), and impulse-response shapes are highly varied (10th/90th percentile R² of a quadratic fit is 0.46/0.98) (Table 1, p. 8).

Q3. What identification schemes and estimators does the comparison cover?

Three identification schemes are studied — the shock directly observed, a noisy IV/proxy for the shock, and recursive (Cholesky) identification as a robustness check — and for each, several LP and VAR variants are compared: least-squares LP, bias-corrected LP (Herbst and Johannsen 2023’s analytical small-sample correction), and penalized LP (Barnichon and Brownlees 2019’s B-spline-smoothed LP) on one side; least-squares VAR, bias-corrected VAR (Kilian 1998/Pope 1990), a Bayesian VAR with a Minnesota-type prior (Giannone, Lenza, and Primiceri 2015), and a data-driven VAR model average (Hansen 2016, over AR(1)-AR(20) and VAR(1)-VAR(20)) on the other (Section 3.2, pp. 6-7; Section 4, pp. 9-10; Appendix B, pp. 19-20). For the IV-identified case, SVAR-IV (the external-instrument estimator of Stock 2008, Mertens and Ravn 2013, and Gertler and Karadi 2015) is added as a comparison. All baseline estimators use p=4 lags.

Q4. What is the central “Lesson 1” bias-variance trade-off, and how large is it?

Across the 6,000 DGPs, LP and VAR have similar bias at short horizons (h ≤ the p=4 lag length), but at longer horizons VAR bias grows substantially larger than LP bias — while LP’s standard deviation rises steeply with horizon, reaching roughly double the VAR’s standard deviation by horizon 20 (Section 5.1, pp. 10-11; Figures 2-3, p. 11). Bias-corrected LP removes about one-third of least-squares LP’s bias at all horizons, but bias correction is “not free”: both BC LP and BC VAR have higher standard deviation than their uncorrected counterparts. Bayesian VAR achieves the uniformly lowest median standard deviation of all estimators studied. The authors note these statements hold after applying the relevant small-sample bias corrections (Pope 1990, Kilian 1998, Herbst and Johannsen 2023) (Section 5.1, p. 10).

Q5. When, if ever, is bias-corrected LP actually the better choice in practice?

Bias-corrected LP is optimal only in a narrow region of the paper’s loss-function space: very high weight on bias (ω ≥ 0.9-0.95, where ω=0.5 corresponds to conventional MSE) combined with intermediate horizons (roughly h = 6-14); for essentially any weight below that threshold, VAR-based methods dominate throughout the horizon range, and no region favors uncorrected least-squares LP at all (Section 5.2, pp. 11-13; Figure 6, p. 13). Under MSE loss specifically, least-squares LP usually dominates bias-corrected LP head-to-head, because BC LP’s bias reduction comes at a substantial variance cost (Figure 4, p. 12); the authors describe uncorrected LP as “essentially dominated” — it has more bias than BC LP yet materially higher variance than VAR (p. 12).

Q6. Under what conditions do VAR-based methods do best, and does that depend on which VAR variant?

A Bayesian VAR is preferred over least-squares VAR in a majority of DGPs at short horizons (h ≤ 4); least-squares VAR overtakes it at intermediate horizons (h ∈ [5,12]); and at long horizons (h ≥ 13) the two are comparable and jointly outperform all other methods unless the researcher places very high weight on bias (Section 5.3, pp. 13-15; Figure 9, p. 15). BVAR’s relatively higher bias at intermediate horizons is attributed to its Minnesota-type prior being designed to help one-step-ahead and long-run forecasting, not intermediate-horizon impulse responses (footnote 19, p. 15). By contrast, VAR model averaging performs poorly regardless of loss function or horizon because of a fat-tailed sampling distribution, and bias-corrected VAR is rarely the optimal choice (Section 5.3).

Q7. What does the paper find about the external-instrument SVAR-IV estimator, and why?

SVAR-IV is the most heavily biased estimator among the IV-based methods at every horizon studied, with bias growing steeply as the horizon lengthens, yet it also has substantially lower dispersion than the other IV estimators (Section 5.4, pp. 15-17; Figures 10-11, pp. 16-17). The paper traces the bias to non-invertibility: 90% of the simulated DGPs have a degree of shock invertibility below 49% (Table 1, p. 8), and SVAR-IV — unlike internal-IV local-projection or VAR approaches — is asymptotically biased when the structural shock is not invertible from the observables (citing Plagborg-Møller and Wolf 2021, 2022). The authors nonetheless flag that “its low dispersion is intriguing and may in some cases trump the bias concerns” (p. 16), since excluding the instrument from the reduced-form system reduces the estimated system’s dimension and hence its sampling variance.

Q8. Do conventional model-selection tools (lag-length criteria, specification tests) help researchers detect which situation they are in?

No — the paper finds that standard tools essentially fail to flag VAR mis-specification in this DGP population: the 90th percentile of the AIC-selected lag length never exceeds 2 in any of the 6,000 DGPs (and equals exactly 2 in only 68.3% of them), and a Lagrange-Multiplier serial-correlation test never achieves a rejection probability above 50% in any DGP. The authors conclude that “conventional model selection tools cannot detect substantial mis-specification of VAR(4) in the vast majority of DGPs, despite the fact that many of our DGPs are in fact not well approximated by a VAR(4)” (Section 5.6, pp. 17-18) — meaning a practitioner cannot simply run a specification test to decide whether to trust a fixed-lag VAR.

Q9. How robust are these lessons, and what are the paper’s own stated scope limits?

The bias-variance trade-off and the qualitative estimator rankings are described as robust across a range of alternative designs — stationary DGPs, a restricted set of 17 “salient” macro series (1,581 DGPs), monetary versus fiscal shocks analyzed separately, a longer lag length (p=8), a smaller sample (T=100), a monthly calibration (T=720, p=12), and a larger cross-section of observables (Section 5.5, pp. 16-17) — but the authors are explicit that “these conclusions inevitably depend on the choice of encompassing model and the specific implementation of the impulse response estimators” (Conclusion, p. 19). The study is confined to point estimation, not inference (pointing readers to Inoue and Kilian 2020, Montiel Olea and Plagborg-Møller 2021, and Xu 2023 for confidence-interval procedures); it uses only five-variable systems and does not examine sign restrictions, long-run restrictions, or non-recursive structural schemes; and it focuses on average performance across DGPs rather than near-worst-case performance, though 90th-percentile results are reported in a supplemental appendix (Introduction, p. 3; Conclusion, pp. 18-19).

Key terms in this paper

Definitions below follow the paper's own usage.

Encompassing model
the single non-stationary dynamic factor model — six latent factors with a VAR(4)/VECM representation, estimated on 207 quarterly U.S. series (1959Q1-2014Q4) — from which all 6,000 simulated DGPs are drawn; the paper's stated lessons are conditional on this choice of encompassing model, not universal (Section 3.1, p. 5; Conclusion, p. 19).
Degree of shock invertibility
the extent to which a structural shock can be recovered as a function of current and past values of the observed variables in a given DGP; in this paper's simulated population, 90% of DGPs fall below 49% invertibility, which is the paper's proposed explanation for why the external-instrument SVAR-IV estimator is so heavily biased relative to internal-IV alternatives (Table 1, p. 8; Section 5.4, p. 16).
Bias-variance loss function (ω)
the paper's device for ranking estimators, L_ω = ω·(bias)² + (1-ω)·Var(θ̂), where ω=0.5 corresponds to conventional MSE and ω=1 reflects exclusive concern with bias; used throughout to characterize exactly how much a researcher would need to prioritize bias reduction before an estimator like bias-corrected LP becomes preferable (Section 2, Eq. 4, p. 5).
Bias-corrected LP / bias-corrected VAR
LP or VAR estimates adjusted with an analytical T⁻¹ small-sample bias correction (Herbst and Johannsen 2023 for LP; Kilian 1998 following Pope 1990 for VAR) to remove some of the persistence-driven small-sample bias; in this paper's simulations the correction is not free — it reduces bias (by about one-third for LP) but increases variance relative to the uncorrected estimator (Section 5.1, p. 10; Appendix B, p. 20).
SVAR-IV (external instrument)
the identification approach that excludes the instrument z_t from the reduced-form VAR and projects the structural shock onto the VAR's reduced-form residuals (Stock 2008; Mertens and Ravn 2013; Gertler and Karadi 2015); the paper documents that it is asymptotically biased under non-invertibility — a "pervasive and realistic feature" of its simulated DGPs — but that excluding the instrument also lowers the dimensionality of the estimated system and hence its dispersion (Section 4, p. 10; Section 5.4, pp. 16, 19).
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.