Macro Paper Warehouse
Published Classic [Journal of Econometrics] doi:10.1016/j.jeconom.2019.04.024 Vol. 212, No. 1, pp. 137-154

Large Bayesian vector autoregressions with stochastic volatility and non-conjugate priors

Andrea Carriero

Todd E. Clark

Massimiliano Marcellino

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Statistical systems that track many economic series at once become far more useful once they allow the size of shocks to vary over time, but that combination used to be too slow to compute. This 2019 paper rearranges the calculation so it can proceed one equation at a time, fast enough to estimate a 125-variable monthly system in about seven hours. On United States data it finds volatility falling from the 1970s into the mid-1980s, and that breadth and time-varying volatility together improve forecasts by more than the sum of their separate gains. Why it matters: better uncertainty estimates, not merely better central forecasts.

What this paper finds — and why it matters

This 2019 Journal of Econometrics paper by Andrea Carriero, Todd Clark, and Massimiliano Marcellino develops a computationally efficient Markov Chain Monte Carlo (MCMC) algorithm for estimating large Bayesian vector autoregressions (VARs) – models with many included variables – that feature stochastic volatility (time-varying error variances) together with non-conjugate priors, meaning economically motivated shrinkage priors (Minnesota, Sims-Zha, cross-variable shrinkage) whose independent Normal-Wishart form breaks the Kronecker structure that the standard closed-form conjugate posterior requires. The key device is a triangular (Cholesky-type) factorization of the time-varying error covariance matrix, used purely as a computational tool rather than as a structural identification assumption – the authors argue (Section 3.1, “The role of variable ordering”) that the draw of the VAR coefficients from their conditional posterior via this triangular recursion is invariant to the ordering of variables used to build A – which lets the VAR coefficients be drawn equation-by-equation with standard GLS formulas instead of requiring the full system-wide posterior precision matrix to be formed and inverted. This lowers the computational complexity of the coefficient draw from O(N^6) to O(N^4): for a 20-variable VAR the paper reports that estimation using the traditional system-wide algorithm was about 356 times slower than the triangular algorithm (Section 4.1), with the advantage growing further as the number of variables rises. The algorithm makes estimation of very large systems tractable: the paper estimates a 125-variable monthly VAR with 13 lags (roughly 203,250 mean coefficients) on the FRED-MD dataset, drawing 5,000 posterior draws in about seven hours, and documents substantial heterogeneity in stochastic volatility across financial, real, and price variable groups; evidence consistent with the Great Moderation, with posterior median volatilities declining from the 1970s/early 1980s into the mid-1980s; a first principal component of log-volatilities that accounts for roughly 45% of total volatility variance; and recursively identified monetary-policy-shock responses (with the federal funds rate ordered last) consistent with Banbura, Giannone, and Reichlin (2010): a contractionary shock lowers output and inflation with no price puzzle. In a second, forecasting application using four macroeconomic variables over 1960:3-2014:5 with 531 recursive out-of-sample forecasting exercises, the paper finds that a large cross-section improves point forecasts (RMSE) while stochastic volatility adds little further gain there, but stochastic volatility matters for density-forecast accuracy (log scores), and that combining a large cross-section with stochastic volatility produces gains that are super-additive – larger than the sum of the gains from each ingredient taken separately – holding across variables and forecast horizons of 1, 3, 6, and 12 months.

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What computational bottleneck does the paper solve, and how large is the resulting speedup?

Estimating a Bayesian VAR with many variables (N) under a non-conjugate prior requires forming and inverting a posterior precision matrix that scales as N·K x N·K (K being the regressors per equation), which costs O(N^6) operations – prohibitive once N grows beyond a handful of variables. The paper’s triangular algorithm reduces this to O(N^4) by drawing the coefficients equation-by-equation instead of system-wide. For a 20-variable VAR, the authors report that estimating the model with the conventional system-wide algorithm was about 356 times slower than with the triangular algorithm (Section 4.1, p. 146) – and the gap widens further as N increases, since the system-wide cost scales as N^6 while the triangular cost scales as N^4.

Q2. How does the triangular algorithm achieve this speedup, and why doesn’t it impose an identification restriction?

The algorithm factors the time-varying reduced-form covariance matrix as Sigma_t = A^{-1} Lambda_t (A^{-1})’, where A is lower-triangular with ones on the diagonal and Lambda_t is a diagonal time-varying variance matrix; this Cholesky-type structure decouples the posterior of each equation’s coefficients from the equations ordered after it, so all N coefficient blocks can be drawn independently with standard GLS formulas rather than jointly. Because A is triangular, transforming the system by ỹ_t = A y_t and X̃_t = A ⊗ X_t makes equation j depend only on the rows of A associated with equations 1,…,j-1. Crucially, the authors argue in Section 3.1 (“The role of variable ordering,” p. 144) that the ordering of the equations within this triangular recursion is inconsequential to the resulting draw of Pi – so this triangularization is described as an “estimation device only,” not a structural identifying assumption about the ordering of shocks.

Q3. What prior class does the algorithm accommodate, and why is it “non-conjugate”?

The paper accommodates priors of the independent Normal-Wishart form, in which the prior on the VAR coefficients (Pi | Sigma ~ N(Pi^0, Omega_Pi)) does not depend on the error covariance Sigma – this independence is what makes the prior “non-conjugate” and is also what breaks the Kronecker structure the standard conjugate Normal-Wishart posterior relies on, so the usual closed-form posterior is unavailable. This independent-prior form is what allows economically motivated shrinkage schemes – the Minnesota prior, Sims-Zha priors, and cross-variable shrinkage penalties – to be imposed; a purely conjugate prior cannot accommodate these in general (Section 2.3, pp. 139-140).

Q4. What does the 125-variable structural VAR application find about volatility and monetary policy transmission?

Estimating a 13-lag VAR on 125 monthly FRED-MD variables (roughly 203,250 mean coefficients, 5,000 draws in about seven hours), the paper finds substantial heterogeneity in stochastic volatility across financial, real, and price variable groups, together with evidence consistent with the Great Moderation: posterior median volatilities decline from the 1970s/early 1980s into the mid-1980s. A first principal component of the estimated log-volatilities explains roughly 45% of total volatility variance, indicating a single common factor drives much of the co-movement in volatility across the 125 series (Fig. 2, p. 148). Recursively identified monetary policy shocks – with the federal funds rate ordered last – produce impulse responses consistent with Banbura, Giannone, and Reichlin (2010): a contractionary shock reduces output and inflation, with no price puzzle (Section 5.2, pp. 148-149).

Q5. What does the forecasting application show about the separate and combined contributions of a large cross-section and stochastic volatility?

Using four macroeconomic variables (matched to prior BVAR forecasting literature) over 1960:3-2014:5 with 531 recursive out-of-sample forecasts, the paper finds that expanding the cross-section improves point forecasts (lower RMSE) while adding stochastic volatility contributes little further to point-forecast accuracy; stochastic volatility instead matters for density forecasts, where it materially improves log scores, and a large cross-section also improves density forecasts. Most notably, the gains from combining a large cross-section with stochastic volatility are super-additive: the jointly large-and-heteroskedastic model outperforms the sum of the individual gains from a large-only model and an SV-only model (Section 6.2, p. 152), and this pattern holds across variables and across forecast horizons of 1, 3, 6, and 12 months.

Q6. Why does the paper describe the two ingredients – large cross-section and stochastic volatility – as complementary rather than substitutes for forecasting?

The paper’s proposed mechanism is that stochastic volatility captures time-varying forecast uncertainty that matters for density calibration, especially in high-volatility episodes, while a larger cross-section improves signal extraction about the conditional mean – because these two channels address different sources of forecast error, their combined effect exceeds the sum of their separate effects (Section 6.2, pp. 151-152). This is offered as the explanation for the super-additivity documented in Q5, rather than as an independently tested causal claim.

Q7. What are the main limitations the authors themselves note, and what happened to this paper after publication?

The authors flag that for the independent Normal-Wishart prior the marginal likelihood is not available in closed form, which rules out formal Bayes-factor model comparison and requires prior-sensitivity analysis instead (Section 2.3, pp. 139-140); they also note that density forecasts remain imperfect during periods of large, rapid volatility swings such as the 2008-2009 financial crisis, since stochastic volatility helps but does not fully capture jump-like transitions in volatility (Section 6.2, pp. 151-152). A further modeling limitation is that the triangular A matrix is held fixed over time – only the diagonal Lambda_t varies – so the algorithm cannot capture time-varying contemporaneous transmission of shocks across variables, which could matter if financial linkages shift over time; and the recursive ordering used to identify the monetary policy shock in the 125-variable application is acknowledged by the authors as a simplification for that large-scale exercise (Section 5.2, p. 148). Separately, and outside the paper’s own text: this paper has a published corrigendum (Journal of Econometrics 227(2), 2022, pp. 506-512) correcting an error in the original algorithm; the corrigendum is not filed in this library and its contents are not summarized here, so any reader implementing the method described in this 2019 paper should consult the corrigendum before doing so.

Key terms in this paper

Definitions below follow the paper's own usage.

Triangular (Cholesky-type) factorization as an estimation device
the paper's decomposition of the time-varying error covariance as Sigma_t = A^{-1} Lambda_t (A^{-1})', where A is lower-triangular; used here purely to make equation-by-equation Gibbs sampling possible, and argued in Section 3.1 ("The role of variable ordering") to be an estimation device only -- not an ordering-dependent structural interpretation -- since any identification scheme can be applied afterward to the resulting reduced-form draws.
Non-conjugate (independent Normal-Wishart) prior
a prior in which the VAR coefficients' distribution does not depend on the error covariance matrix, unlike the standard conjugate Normal-Wishart prior; this independence is what permits Minnesota, Sims-Zha, and cross-variable shrinkage but also removes the closed-form posterior, motivating the paper's MCMC algorithm.
Stochastic volatility (SV)
time variation in the diagonal elements of Lambda_t, modeled as each log-volatility following a random walk with innovations that are correlated across equations (via an unrestricted covariance matrix Phi); estimated with the Kim-Shepherd-Chib (1998) mixture-of-normals approximation as one step of the Gibbs sampler.
Super-additivity (in the forecasting application)
the empirical finding that the joint improvement in forecast accuracy from combining a large cross-section with stochastic volatility exceeds the sum of the improvements each ingredient delivers on its own, interpreted as reflecting that the two ingredients address distinct sources of forecast error (signal extraction versus time-varying uncertainty).
Ordering invariance
the property, argued in Section 3.1 ("The role of variable ordering") as a prose discussion rather than a numbered formal result, that the draw of the VAR coefficients Pi from their conditional posterior via the triangular recursion is unaffected by which variable is placed first, second, etc. in constructing the triangular matrix A -- distinguishing the paper's use of a triangular factorization from a structural (e.g., Cholesky-identified SVAR) ordering assumption.
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.