Simultaneous Confidence Bands: Theory, Implementation, and an Application to SVARs
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
When economists report a whole sequence of estimates at once — say how output responds over three years to an interest-rate change — the usual uncertainty bands are built one estimate at a time, so reading the sequence as a whole overstates confidence. This 2019 paper shows the narrowest band that still covers the entire sequence simply scales every estimate's margin of error by one common number. In the authors' United States application it runs roughly 35 percent narrower than standard alternatives in one specification. Why it matters: honest but tighter bands let researchers see effects that overly cautious methods hide.
What this paper finds — and why it matters
This 2019 Journal of Applied Econometrics paper by José Luis Montiel Olea and Mikkel Plagborg-Møller addresses a practical problem in applied time-series econometrics: which “simultaneous” (joint, across-parameter) confidence band should researchers use by default when reporting an entire vector of estimated quantities – such as impulse responses across horizons from a structural VAR – rather than one interval per parameter reported separately. The authors set up a general framework in which a possibly nonlinear transformation theta = h(mu) of an asymptotically normal estimator mu-hat is to be covered jointly, and show that a wide class of popular bands (pointwise, Bonferroni, Sidak, projection, and the “sup-t” band) can all be written as members of a single “one-parameter class” that scales every pointwise standard error by one common critical value c. Within this class, the sup-t band – whose critical value is the quantile of the maximum absolute studentized draw from the estimator’s joint asymptotic distribution – is the narrowest band that still achieves exact asymptotic simultaneous coverage, and the paper adds a decision-theoretic result (Proposition 1) showing it uniquely minimizes worst-case regret across all degree-one-homogeneous loss functions, which makes it a defensible default when the researcher does not know which feature of the band matters most to different readers. The paper gives three computationally convenient ways to construct it – a plug-in (delta-method) simulation algorithm, a bootstrap algorithm, and a Bayesian algorithm delivering exact finite-sample simultaneous credibility – and notes all three are first-order asymptotically equivalent. In an empirical application to a monthly U.S. SVAR (July 1979-June 2012, 12 lags; identified two ways, via a recursive/Cholesky scheme and via a Gertler-Karadi (2015)-style external instrument using federal-funds-futures surprises from January 1990) with industrial production, CPI, a one-year bond yield, and the excess bond premium, the sup-t band is substantially narrower than the Bonferroni or Sidak bands – around 35% narrower in the external-instrument specification at 68% confidence – and the narrowing is not merely cosmetic: at the 68% simultaneous level the plug-in sup-t band excludes zero for the industrial-production response at horizons of roughly 13-36 months, letting the authors reject the no-effect null at some horizon in that range, whereas the Bonferroni band does not permit that rejection; conversely, an output response that looks pointwise significant at the 2-month horizon is no longer simultaneously significant once the sup-t multiple-comparison adjustment is applied. A companion Monte Carlo study of bivariate VARs finds the sup-t band 20-25% narrower than Bonferroni/Sidak at 68% confidence and 10-20% narrower at 90% confidence, though for highly persistent data only the Bayesian sup-t implementation is reported to achieve satisfactory finite-sample coverage. The theory is developed for point-identified models with a continuously differentiable transformation h(.); partially identified (e.g., sign-restricted) SVARs require additional considerations, though the authors suggest the Bayesian sup-t band may still be usable there for subjective Bayesian analysis.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What practical decision does this paper help applied researchers make?
It answers the question of which confidence band to report by default when a researcher needs to cover an entire vector of estimated quantities simultaneously – for example, a structural impulse-response function across many horizons – rather than reporting separate pointwise intervals that only cover each parameter one at a time. The paper works in a general framework where a possibly nonlinear transformation theta = h(mu) of an asymptotically normal underlying estimator mu-hat (mu-hat satisfying root-n(mu-hat - mu) converging to a normal distribution) is the object to be covered; a simultaneous 1 - alpha confidence band is a Cartesian product of intervals, one per component of theta, that jointly covers the whole vector with asymptotic probability at least 1 - alpha. The framework nests VAR impulse-response analysis as a leading case, since impulse responses at many horizons are nonlinear transformations of a finite set of underlying VAR coefficients, and the transformation’s Jacobian can even make the asymptotic covariance singular when there are more horizons of interest (k) than underlying parameters (p).
Q2. What is the “one-parameter class” of bands, and how do familiar procedures fit into it?
The authors show that pointwise, Bonferroni, Sidak, projection, and sup-t bands are all special cases of a single “one-parameter class,” in which every component interval is the pointwise estimate plus-or-minus its own standard error times one common scalar critical value c. Formally, the band is the product of intervals [theta-hat_j - sigma-hat_j * c, theta-hat_j + sigma-hat_j * c] for j = 1,…,k, where sigma-hat_j is the delta-method pointwise standard error of theta-hat_j; the only thing that distinguishes the named procedures is which value of c they use (e.g., a chi-square pointwise critical value, a Bonferroni-adjusted one using alpha/k, or a Sidak-adjusted one using (1-alpha)^(1/k)). Because every candidate in Table 1 fits this one template, the choice among them reduces to the choice of critical value, which the paper’s ordering result (Section 3.3) ranks as sup-t <= Sidak <= Bonferroni, with the theta-projection and mu-projection bands typically wider still for alpha <= 0.5.
Q3. What is the sup-t band, and why is it the narrowest band with correct simultaneous coverage?
The sup-t band uses, as its critical value c, the (1 - alpha) quantile of the maximum absolute studentized coordinate of the estimator’s joint asymptotic normal distribution – the smallest c consistent with exact asymptotic simultaneous coverage – which by construction makes it the narrowest member of the one-parameter class that still achieves the nominal coverage. Its asymptotic coverage probability equals P(max_j |Sigma_jj^(-1/2) V_j| <= c) for V jointly normal with the estimator’s asymptotic covariance Sigma, evaluated at c = q_{1-alpha}(Sigma), the (1-alpha) quantile of that maximum. The gain over Bonferroni/Sidak comes specifically from exploiting the correlation across the theta-hat_j’s – the paper notes the sup-t advantage is largest exactly when components are highly correlated, “as is typical for impulse responses at adjacent horizons” – and the sup-t band’s width grows only at the slow rate of the square root of log(k) in the number of parameters k, whereas the projection bands’ width is far more sensitive to k.
Q4. How can the sup-t band actually be computed in practice?
The paper gives two implementations that are asymptotically equivalent but suit different settings: a “plug-in” simulation algorithm using the estimated asymptotic covariance, and a bootstrap-or-Bayes algorithm that draws directly from the bootstrap or posterior distribution of the underlying parameter. In the plug-in version (Algorithm 1), the researcher computes the delta-method covariance estimate Sigma-hat = h-dot(mu-hat) * Omega-hat * h-dot(mu-hat)’, simulates many draws from N(0, Sigma-hat), and takes the empirical (1-alpha) quantile of the maximum studentized draw as the critical value. In the bootstrap/Bayes version (Algorithm 2), the researcher draws many mu-hat replications from the bootstrap or posterior distribution, maps each through h(.) to get a theta-hat replication, and calibrates a common quantile-trimming level so that the resulting rectangle of equal-tailed intervals contains at least fraction 1-alpha of the draws; the resulting Bayesian band has exact finite-sample simultaneous Bayesian credibility (up to simulation error), while the plug-in band is not guaranteed to deliver an asymptotic refinement over it (footnote 11).
Q5. Why does the paper argue the sup-t band is a good default even when the researcher does not know what “narrow” means to different readers?
The paper offers a decision-theoretic justification (Section 4, Proposition 1): among all translation-equivariant simultaneous confidence bands, the sup-t band is the unique minimizer of worst-case regret, where regret is measured relative to the best possible band under any degree-one-homogeneous loss function applied to the vector of interval lengths. Because different readers of a published impulse-response figure might care about different features of the band – its width at one particular horizon, its width averaged across horizons, and so on – a band that is optimal only for one specific loss function could perform poorly for another reader’s implicit preferences. The sup-t band’s worst-case-regret optimality means it never performs arbitrarily worse than the best band tailored to any such loss function, which the authors offer as the argument for treating it as the default choice absent more specific information about how the band will be used; for a fully specified known loss function, the authors note that Freyberger and Rai’s (2018) numerically optimal band could instead be preferred (Section 4.2.4).
Q6. How does the paper illustrate all of this with a real structural VAR?
Section 5 replicates a Gertler-Karadi (2015)-style monthly U.S. SVAR – industrial production, CPI, a one-year government bond yield, and the excess bond premium (Gilchrist-Zakrajsek 2012), estimated in levels with 12 lags over July 1979-June 2012 (external-instrument data from January 1990), with p = 206 underlying VAR/instrument parameters and k = 37 horizons of interest (0-36 months) – and constructs confidence bands for the impulse responses to a monetary policy shock under two identification schemes. The first is a recursive/Cholesky scheme (industrial production and CPI do not respond to the monetary shock on impact; the bond yield does not respond to the excess bond premium on impact); the second is an external-instrument scheme using FF4, a three-month-ahead federal funds futures surprise measured in a short window around FOMC announcements. This design lets the paper compare sup-t, Bonferroni, Sidak, and other bands on the exact kind of object – a multi-horizon impulse-response function – that motivated the theory.
Q7. What do the empirical results show, and do they change any substantive conclusion about the effects of monetary policy?
Yes – the choice of band changes which horizons show a statistically significant output response. The sup-t band is roughly 35% narrower than the Sidak and Bonferroni bands in the external-instrument specification at 68% confidence. At that 68% simultaneous confidence level, the plug-in sup-t band excludes zero for the industrial-production response at horizons of roughly 13-36 months, which lets the authors reject, at some horizon in that range, the null that the monetary shock has no effect on output; the Bonferroni band, being wider, does not permit that rejection at the same nominal level. The comparison also runs the other way: a positive industrial-production response that looks pointwise statistically significant at the 2-month horizon loses its significance once the sup-t band’s simultaneous multiple-comparison adjustment is applied, illustrating that pointwise significance and simultaneous significance are genuinely different standards. Comparing the three sup-t implementations (plug-in, bootstrap, Bayes) shows they are similar at short horizons but diverge somewhat at longer horizons under the recursive scheme, with the authors suggesting the Bayesian band may be preferable when the data are highly persistent.
Q8. What does the companion Monte Carlo simulation show about how general the sup-t advantage is?
A simulation study of bivariate VARs under recursive or external-instrument identification finds the sup-t band 20-25% narrower than the Bonferroni/Sidak bands at 68% confidence, and 10-20% narrower at 90% confidence, with projection bands described as unnecessarily wide throughout. The one qualification concerns persistence: for highly persistent simulated data, only the Bayesian sup-t implementation is reported to deliver satisfactory finite-sample simultaneous coverage, suggesting the plug-in and bootstrap versions can be less reliable specifically in that regime even though all three are asymptotically equivalent.
Q9. What extensions and limitations does the paper itself flag?
The authors note an extension to “generalized familywise error rate control,” in which the band is only required to cover at least k - m of the k parameters (rather than all k), obtained by replacing the maximum operator in Algorithm 2 with the (m+1)-th order statistic – a less stringent, multiple-testing-style coverage criterion (Section 6.1). On limitations: the core theory is developed for point-identified models with a continuously differentiable transformation h(.); it does not directly cover partially identified models such as sign-restriction SVARs, though the authors suggest the Bayesian sup-t band may still be applied there for subjective Bayesian analysis. They also flag that the sup-t band’s width, while growing only at the slow root-log(k) rate, does still grow with the number of parameters, and that the bootstrap implementation is not guaranteed to offer an asymptotic refinement over the plug-in version. Finally, for a researcher who does know the specific loss function relevant to their application, the worst-case-regret optimality of the sup-t band is not the same as being optimal for that one loss function, and a tailored alternative (Freyberger and Rai 2018) may do better.
Key terms in this paper
Definitions below follow the paper's own usage.
- Sup-t band
- in this paper, the member of the "one-parameter class" of simultaneous confidence bands whose common critical value equals the (1-alpha) quantile of the maximum absolute studentized coordinate of the estimator's joint asymptotic (or bootstrap/posterior) distribution; by construction it is the narrowest band in that class achieving exact asymptotic (or, in the Bayesian version, exact finite-sample) simultaneous coverage.
- One-parameter class of bands
- the paper's organizing device for comparing simultaneous confidence procedures -- any band formed by scaling every component's own pointwise standard error by a single common critical value c; pointwise, Bonferroni, Sidak, projection, and sup-t bands are all instances that differ only in their choice of c.
- Worst-case regret / decision-theoretic optimality
- the criterion behind Proposition 1 -- for a candidate band R, regret is R's loss (applied to its vector of interval lengths) divided by the best attainable loss among bands with correct coverage, maximized over all degree-one-homogeneous loss functions; the sup-t band is shown to be the unique minimizer of this worst-case regret among translation-equivariant bands, which the paper offers as the argument for using it as a default when the relevant loss function to the audience is unknown.
- Simultaneous vs. pointwise coverage/significance
- in this paper's empirical illustration, a "pointwise" confidence interval or significance test evaluates one horizon of an impulse response in isolation, while "simultaneous" coverage requires the entire band to contain the true impulse-response vector at once; the paper shows these standards can disagree in both directions (a pointwise-significant 2-month response becomes insignificant simultaneously; horizons 13-36 months become significant simultaneously under sup-t but not under Bonferroni).
- Generalized familywise error rate control
- the paper's Section 6.1 extension of the sup-t construction that relaxes full simultaneous coverage (covering all k parameters) to covering at least k - m of the k parameters, implemented by replacing the "max" operator in the bootstrap/Bayes algorithm with the (m+1)-th order statistic of the studentized draws.