DSGE pileups
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
Why do estimates of the deep parameters in large macroeconomic models so often stack up against the edge of what theory allows? This 2017 paper argues the cause is weak identification rather than a wrongly specified model, and traces three routes: different parameter values that fit the data identically, boundaries created by the model's own mathematics, and multiple valid solutions. In simulations a discount factor whose true value is 0.99 piles up at one, and a widely used medium-sized model shows all three problems together. Why it matters: the diagnostic tools offered help, but the paper is clear they cannot prevent the problem.
What this paper finds — and why it matters
This 2017 Journal of Economic Dynamics & Control paper by Stephen D. Morris is a methodological contribution explaining why maximum-likelihood or Bayesian estimates of DSGE structural parameters often “pile up” – concentrate at or near a boundary of the theoretically admissible parameter space, or otherwise depart from the asymptotic normal distribution – even when the model is correctly specified. Morris argues the underlying cause is weak identification of structural parameters by the reduced-form data, operating through three specific channels in DSGE models: observational equivalence (the map from structural parameters theta to reduced-form parameters pi sends multiple theta values, including economically implausible ones, to the same pi), functional boundaries created by the theta-to-pi mapping itself (rather than by the stated theoretical restrictions), and multiplicity of stable rational-expectations solutions. To study these mechanisms formally, the paper builds on the author’s own prior result (Morris 2016b) that, under stated regularity, stationarity, and left-invertibility assumptions, the ABCD state-space representation of a DSGE model has an exact finite-order VARMA(p, p-1) representation, and proposes a minimum chi-square estimator (MCSE) – asymptotically equivalent to MLE but numerically simpler and bootstrappable – together with an F-test and an overidentification chi-squared test to diagnose pileups. The paper contains no empirical application to real data; all results come from Monte Carlo experiments (T=225 quarters, N=1000 replications, mimicking a typical postwar quarterly sample) on three small illustrative models plus the Smets-Wouters (2007) medium-scale model. In the Brock-Mirman stochastic growth calibration, the discount-factor estimate beta-hat piles up at its upper boundary of 1 even when the true value is beta_0 = 0.99, with Kolmogorov-Smirnov tests rejecting Gaussianity at the 1%, 5%, and 10% levels; in the Krause-Lubik search-and-matching calibration, the match-elasticity parameter xi piles up near its lower boundary of 0.1 with a long left tail, driven by a “natural” functional boundary rather than the stated theoretical restriction; in the An-Schorfheide New Keynesian model, the CRRA coefficient tau and shock-persistence parameter rho_z show multimodal sampling distributions traceable to an observationally equivalent solution point with an economically infeasible tau of -45.4, a problem that Jeffreys priors only partly resolve and that informative conjugate priors can worsen for other parameters (Morris calls informative priors “a double-edged sword”); and in the Smets-Wouters model (41 structural parameters, VARMA(3,2) representation), all three pileup types recur simultaneously – a boundary pileup at 1 for TFP-shock persistence rho_z, a skewed distribution for the investment adjustment-cost parameter, and bimodality in the trend-growth parameter gamma – and persist under an alternative observable set that replaces hours worked with the labor income share. Morris stresses throughout that these phenomena reflect identification weakness intrinsic to DSGE models, not misspecification, and that while the MCSE framework helps diagnose pileups and enables valid bootstrap inference, it does not offer a general method to prevent them.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What is a “DSGE pileup,” and why does it undermine standard inference?
A pileup occurs when the sampling distribution of a DSGE structural-parameter estimator concentrates at or near a boundary of the admissible parameter space, is multimodal, or is otherwise non-Gaussian – even when the model is correctly specified – which calls into question standard asymptotic-normal confidence intervals and tests. Morris frames the moving-average unit-root problem (|theta-hat| = 1 when theta and 1/theta are observationally equivalent under an invertibility restriction) as the canonical example of this phenomenon, then argues DSGE models face an analogous but potentially worse problem: DSGE pileup regions can be entire subsets of the parameter space rather than isolated points, and can appear as multimodal rather than purely boundary distributions.
Q2. What two structural configurations generate a pileup, and how does the paper depict them?
Morris depicts three overlapping sets in parameter space – Theory, Determinacy, and Identification – and locates two distinct pileup mechanisms in their overlap: (i) the true parameter is theoretically admissible but its identification-equivalent counterpart lies outside the theoretical restrictions, pushing the MLE toward the boundary; and (ii) the true parameter and an observationally equivalent point both lie inside the admissible region, leaving the MLE indeterminate between them. These two configurations correspond, respectively, to the boundary-estimate and multimodal pileup types documented in the paper’s later Monte Carlo examples.
Q3. What analytical machinery does the paper build to formalize pileups?
Building on the author’s companion paper (Morris 2016b), Morris shows that under a set of regularity, stationarity, and left-invertibility assumptions on a DSGE model’s ABCD state-space representation, the model’s observables have an exact finite-order VARMA(p, p-1) representation whose reduced-form parameters pi are identifiable and can be consistently estimated. This VARMA representation is the vehicle for everything that follows: it lets Morris write the DSGE structural-to-reduced-form mapping g: Theta -> pi explicitly and study where g creates observational equivalence or boundary behavior.
Q4. What is the minimum chi-square estimator (MCSE), and why does Morris use it instead of directly maximizing the DSGE likelihood?
The MCSE minimizes a quadratic form T(pi-hat_MLE - g(theta))‘I-hat(pi-hat_MLE - g(theta)) over theta, where pi-hat_MLE is the maximum-likelihood estimate of the VARMA reduced-form parameters and I-hat is their estimated asymptotic information matrix; it is asymptotically equivalent to direct MLE of the structural model but has two practical advantages Morris exploits throughout the paper. First, it is numerically simpler because it separates estimation into a first-stage (reduced-form VARMA) MLE and a second-stage minimum-distance step; second, its bootstrap distribution remains valid even when a parameter is subject to pileup, whereas the asymptotic theory underlying conventional standard errors is unreliable in that case – a point Morris qualifies by noting that global identification (assumed throughout the paper’s examples) is what keeps the bootstrap valid, since bootstrap inference can otherwise fail under weak instruments.
Q5. In the simplest example (the Brock-Mirman stochastic growth model), what does the Monte Carlo evidence show?
Calibrating the three-parameter Brock-Mirman model at alpha_0 = 1/3, beta_0 = 0.99, sigma_0 = 0.1 and simulating 1,000 replications of T=225 quarterly observations, the discount-factor estimate beta-hat piles up at its theoretical upper boundary of 1 even though the true value is comfortably interior at 0.99, while the underlying ARMA(1,1) reduced-form estimates remain Gaussian; Kolmogorov-Smirnov tests reject normality of beta-hat at the 1%, 5%, and 10% levels. Because the model is just-identified here (the MCSE coincides with the MLE), Morris uses this example to establish the “boundary estimate” pileup type in its cleanest form and to show the associated small-sample bias of alpha and beta toward 1.
Q6. What does the Krause-Lubik search-and-matching example add – the “skew” pileup type?
In the six-parameter Krause-Lubik (2010) search-and-matching calibration, the match-elasticity parameter xi develops a long left tail and piles up near 0.1 – not the theoretical boundary of Θ (0 < xi < 1), but a “natural” boundary created by the model’s own functional form – a “skew” pileup distinct from the pure boundary case. Morris traces this to a “natural boundary” created by the functional correspondence between two other parameters (A_x and C_y, jointly identifiable but not individually identifiable) rather than by the stated theoretical restriction on xi itself – so xi piles up even in replications where A_x and C_y are estimated well, meaning the pileup originates in the shape of the structural-to-reduced-form mapping, not just in the stated parameter bounds.
Q7. What does the An-Schorfheide New Keynesian example show about multimodality, and can priors fix it?
In the 14-parameter An-Schorfheide (2007) New Keynesian model, Morris finds an observationally equivalent solution point theta_0 implying an economically infeasible CRRA risk-aversion coefficient of tau = -45.4; without bounds on the parameter space this produces a genuinely bimodal sampling distribution in tau and the shock-persistence parameter rho_z (with rho_r and psi_pi also affected), and even after imposing the theoretical bounds that exclude theta_0, the proximity of the true parameter to the boundary keeps multimodality alive in rho_z and rho_gz.** On the Bayesian side, Jeffreys priors reduce multimodality in tau but leave it “very much apparent” in rho_z; informative conjugate priors can eliminate the rho_z multimodality when only mildly informative (prior concentration c=1000) but, when made more informative (c=100), instead push the posterior for rho_rz to accumulate on the wrong side of zero – leading Morris to describe informative priors as “a double-edged sword” rather than a general fix.
Q8. Do these pileup phenomena survive in a realistic, medium-scale DSGE model?
Yes: applying the same Monte Carlo design (T=225, N=1000) to the Smets-Wouters (2007) model – 41 structural parameters mapped to a VARMA(3,2) representation with 252 identifiable reduced-form parameters, after fixing five parameters to preserve identifiability of the rest – reproduces all three pileup types simultaneously: a boundary pileup at 1 for TFP-shock persistence rho_z, a long right-tailed skew for the investment adjustment-cost parameter, and clear bimodality for the average trend-growth parameter gamma. Replacing hours worked with a labor-income-share proxy among the observables changes some pileups but does not eliminate those in beta, the price-indexation parameter, or gamma, leading Morris to conclude that the choice of observables can matter but is not, by itself, an easy fix.
Q9. Does any of this mean the estimated models are misspecified, and does the paper offer a general remedy?
No – Morris is explicit that the boundary estimates, skew, and multimodality found in the Smets-Wouters exercise are “intimately related to the identifiability properties of the model” rather than evidence of misspecification, and the paper does not claim to have found a way to prevent pileups in general. What the MCSE/VARMA framework does provide is diagnosis: an F-test on exclusion restrictions in the VARMA autoregressive polynomial (for example, testing the rho_z boundary pileup) and a chi-squared overidentification test from the MCSE criterion, both of which can help distinguish a genuine pileup from misspecification; Morris describes the approach as “useful in the very common situation that such pileups should arise,” not as a cure.
Key terms in this paper
Definitions below follow the paper's own usage.
- pileup
- the sampling distribution of a structural parameter estimator concentrating at or near a boundary of the admissible parameter space, being multimodal, or otherwise deviating from asymptotic normality -- even under correct model specification -- so that standard asymptotic inference becomes unreliable.
- observational equivalence
- the situation in which the structural-to-reduced-form mapping g: Theta -> pi sends more than one structural parameter vector to the same reduced-form parameters pi; whether the observationally equivalent point lies inside or outside the theoretically admissible set Theta determines, in this paper’s framework, whether the resulting pileup shows up as a boundary estimate or as multimodality.
- minimum chi-square estimator (MCSE)
- the two-step, information-weighted minimum-distance estimator that matches first-stage VARMA maximum-likelihood estimates of the reduced-form parameters to a fitted structural mapping g(theta); used in this paper as a numerically simpler and bootstrappable substitute for directly maximizing the DSGE likelihood, and as the basis for an overidentification test.
- functional (natural) boundary
- a boundary in the sampling distribution of a structural estimator that arises from the shape of the structural-to-reduced-form mapping itself, rather than from an economic-theory restriction on the parameter -- illustrated by the match-elasticity parameter in the Krause-Lubik example, which piles up at a boundary created by its joint functional relationship with another parameter.
- VARMA(p, p-1) representation
- the finite-order vector ARMA representation that, per Morris (2016b) and this paper, any DSGE model satisfying stated regularity, stationarity, left-invertibility, and state-dimension assumptions admits, with p determined by the number of lags needed for a state-dependent regression to attain full column rank; this representation is what makes the structural-to-reduced-form mapping and the MCSE machinery computable in the paper’s examples.