Vector Autoregressions
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
Twenty years after macroeconomists began fitting small systems of interacting time series to the economy, how well had the tool done? This 2001 review scores it on four jobs, using inflation, unemployment and the Federal Reserve's policy rate on United States quarterly data from 1960 to 2000. It finds the approach strong for describing data and a solid forecasting benchmark, but much weaker for causal conclusions: omitting variables the Fed actually watched produces the familiar puzzle that tighter policy seems to raise prices, and small changes in the assumed policy rule can double the estimated effects. That matters because such conclusions are only as good as the assumptions identifying them.
What this paper finds — and why it matters
This 2001 Journal of Economic Perspectives paper by James Stock and Mark Watson reviews how vector autoregressions (VARs) have performed at the four core tasks of applied macroeconometrics — data description, forecasting, structural inference, and policy analysis — roughly twenty years after Christopher Sims’s original 1980 proposal, illustrated throughout with a simple three-variable system (inflation, unemployment, and the federal funds rate) estimated on quarterly U.S. data, 1960-2000. The authors find VARs perform strongly at data description (Granger-causality tests, impulse responses, and variance decompositions reveal, for example, that inflation and unemployment shocks jointly account for about 75% of the federal funds rate’s forecast-error variance at a three-year horizon) and provide a solid forecasting benchmark (a small VAR modestly outperforms both a univariate autoregression and a random walk at most horizons in a pseudo out-of-sample exercise), but they are considerably more skeptical about structural inference and policy analysis. Structural VAR identification is criticized on three grounds — omitted-variable bias (illustrated by the “price puzzle,” which arises when variables the Fed actually used to forecast inflation, like commodity prices, are left out of the model), parameter instability in monetary policy rules over long samples, and implausible zero-restriction timing conventions that are sometimes dressed up as “structural” theory without real economic content — and the paper shows that structural impulse responses can be “very sensitive” to seemingly minor changes in the assumed policy rule (switching from a backward-looking to a forward-looking Taylor rule roughly doubles the estimated inflation and unemployment responses to a funds-rate shock). The authors conclude that VARs’ “structural implications are only as sound as their identification schemes,” and that combining good economic theory and institutional detail with flexible statistical methods like VARs remains the central ongoing challenge for the field.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What are the “four tasks” of applied macroeconometrics the paper uses to organize its assessment of VARs?
Stock and Watson organize their review around VARs’ performance at four tasks: data description (summarizing comovements among macroeconomic variables), forecasting, structural inference (estimating the causal effect of a shock, holding other structural relationships fixed), and policy analysis (evaluating the consequences of a change in a policy rule, not just a one-time shock). They use a deliberately simple illustrative three-variable VAR — inflation, the unemployment rate, and the federal funds rate, quarterly U.S. data 1960:I-2000:IV — to walk through each task concretely rather than surveying results abstractly.
Q2. What are the three “varieties” of VAR the paper distinguishes, and how do they differ?
A reduced-form VAR expresses each variable as a function of lagged values of all variables, with contemporaneously correlated (but serially uncorrelated) error terms, estimated equation-by-equation by OLS; a recursive VAR imposes a specific causal ordering (equivalent to a Cholesky factorization of the reduced-form covariance matrix) so that errors become uncorrelated across equations, with results that can depend on which of the n! possible orderings is chosen; and a structural VAR uses economic theory or institutional knowledge — rather than an arbitrary ordering — to impose identifying restrictions on the contemporaneous relationships among variables, either fully (all causal links specified) or partially (a single equation identified via instruments).
Q3. What does the illustrative VAR reveal about data description — Granger causality, variance decomposition, and impulse responses?
In the illustrative system, unemployment Granger-causes inflation (p=0.02) and the federal funds rate Granger-causes unemployment (p=0.01), while the funds rate does not directly Granger-cause inflation (p=0.27); at a 12-quarter horizon, a recursive variance decomposition attributes about 75% of the funds rate’s own forecast-error variance jointly to inflation and unemployment shocks, while inflation’s forecast-error variance is overwhelmingly (82%) explained by its own shocks. An inflation shock is shown to decay slowly, taking roughly 24 quarters to dissipate, and to be associated with persistent increases in both unemployment and the interest rate — illustrating the kind of multivariate comovement information a univariate model cannot capture.
Q4. How sensitive are the paper’s structural impulse responses to the assumed policy rule, and why does this matter?
Using a backward-looking Taylor rule to identify a structural federal-funds-rate shock, the illustrative system implies inflation falls about 0.3 percentage points over roughly three years and unemployment rises about 0.2 points, concentrated mostly in the third year; switching to a forward-looking Taylor rule instead implies a sharper roughly 0.5-point inflation decline within about a year and a roughly 0.5-point unemployment increase peaking around one year before eventually falling about 0.5 points below baseline by year six. The authors describe the responses as “very sensitive” to this seemingly modest change in the identifying assumption, concluding that structural VAR impulse responses in this application “hinge on detailed institutional knowledge of how the Fed sets interest rates” rather than being a robust, assumption-free description of the data.
Q5. What does the pseudo out-of-sample forecasting exercise show about VARs’ forecasting performance?
Comparing root mean squared forecast errors over a 1985:I-2000:IV pseudo out-of-sample period, the small illustrative VAR generally matches or improves on both a univariate autoregression and a naive random-walk forecast across 2-, 4-, and 8-quarter horizons for inflation, unemployment, and the federal funds rate — for example, the VAR’s 8-quarter-ahead RMSE for the funds rate (1.70) is noticeably lower than the random walk’s (2.18). The authors summarize this as “the VAR either does no worse than or improves upon the univariate autoregression, and both improve upon the random walk forecast,” positioning small VARs as a standard, credible benchmark against which more elaborate forecasting methods are judged.
Q6. What are the three criticisms the paper raises against structural VAR identification?
First, omitted-variable bias: factors left out of the model get absorbed into the estimated “shock,” illustrated by the price puzzle, which the paper attributes to the Fed having used forward-looking information (such as commodity prices) that simple VARs omitted, biasing the estimated response of prices to a policy shock — a problem with the standard fix of including commodity prices in the system. Second, parameter instability: several studies (Bernanke and Blinder 1992; Bernanke and Mihov 1998; Clarida, Galí, and Gertler 2000; Boivin 2000) document that monetary policy rules have not been stable over long samples, undermining constant-coefficient VARs estimated over many decades. Third, implausible timing conventions: zero within-period restrictions (e.g., assuming output cannot respond contemporaneously to a policy shock) may be defensible at a daily frequency but become “less plausible” at a monthly or quarterly frequency, and the authors warn that researchers sometimes construct a convenient theoretical story — a “Wold causal chain” — to justify a particular recursive ordering after the fact, cautioning that “rarely does it add value to repackage a recursive VAR and sell it as structural.”
Q7. What constructive alternatives to timing-restriction-based identification does the paper highlight?
The paper points to three constructive approaches for more credible structural VAR identification: exploiting detailed institutional knowledge (Blanchard and Perotti’s use of statutory tax-code features to identify fiscal shocks; Bernanke and Mihov’s reserves-market model for monetary shocks), imposing long-run restrictions grounded in theory (King, Plosser, Stock, and Watson’s use of long-run monetary neutrality), and agnostic or robust identification using inequality or sign restrictions (Faust 1998; Uhlig 1999) that avoid pinning down the exact contemporaneous timing of effects.
Q8. Why does the paper consider policy analysis the most difficult of the four VAR tasks?
Policy analysis — evaluating the consequences of a change in a policy rule, rather than the effect of a single unanticipated shock — runs into the Lucas critique: if the model’s structural equations involve forward-looking expectations, then all of the VAR’s estimated coefficients will in general depend on the policy rule in place, so a VAR estimated under one rule cannot be mechanically used to forecast outcomes under a different rule. The paper notes that “the practical importance of this critique for VAR-based policy analysis is a matter of debate,” and observes that analyzing a one-time policy innovation while holding the rule fixed is comparatively straightforward (a function of the estimated impulse responses), whereas analyzing an actual change in the rule requires a fully and correctly specified structural VAR with every equation properly identified.
Key terms in this paper
Definitions below follow the paper's own usage.
- recursive VAR
- a VAR in which a specific causal ordering of variables is imposed — equivalent to a Cholesky factorization of the reduced-form error covariance matrix — so that structural shocks become uncorrelated across equations; results can depend on which of the n! possible variable orderings is chosen, which the paper treats as a key limitation relative to genuinely structural identification.
- structural VAR
- a VAR in which economic theory or institutional knowledge, rather than an arbitrary ordering, is used to impose identifying restrictions on the contemporaneous relationships among variables — fully specified if every causal link is pinned down, or partially specified if only a subset of structural equations is identified (e.g., via instruments).
- price puzzle
- the finding that a structural VAR can imply prices rise following a contractionary monetary policy shock, attributed here to omitted-variable bias — the Fed's forward-looking use of information (such as commodity prices) to forecast inflation was left out of the estimated model, biasing the recovered price response.
- Wold causal chain
- the paper's term (used critically) for a recursive ordering that has been given a post-hoc economic justification, cautioning that constructing a story to rationalize a convenient ordering does not make a recursive VAR genuinely structural.
- Lucas critique (applied to VARs)
- the point that if structural relationships embed forward-looking expectations, a VAR's estimated coefficients are not policy-invariant — they depend on the prevailing policy rule — so a VAR estimated under one policy regime cannot be mechanically used to predict outcomes under a genuinely different regime, distinguishing "policy analysis" (rule changes) from simple structural shock analysis (one-time innovations).