A method for solving and estimating heterogeneous agent macro models
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
Macro models with many different firms or households are hard to compute because the whole distribution is a state variable. This method fits a flexible curve to that distribution, keeps only its moments, and then hands the resulting finite-sized model to Dynare, the standard toolbox macroeconomists already use. Solving takes under a minute, and a full Bayesian estimation under an hour. The estimation makes a substantive point: how lumpy investment is at the firm level changes what the data say about the shocks driving the aggregate economy.
What this paper finds — and why it matters
This is a computational method for solving and estimating heterogeneous-agent macro models with aggregate shocks, built around one substitution: the infinite-dimensional cross-sectional distribution is replaced by a flexible parametric family, so its moments become finite-dimensional endogenous state variables, and the resulting system is then perturbed around its stationary equilibrium. The stated motivating gap is adoption rather than accuracy: of the many existing algorithms, “none are as general, efficient, or easy-to-use as the standard perturbation methods routinely employed to solve representative agent models using prepackaged toolboxes like Dynare. Due to this challenge, heterogeneous agent models have yet to reach widespread adoption, particularly among central banks and policy institutions.” The method has three steps — globally accurate projection approximations in the individual states (Chebyshev collocation for the firm’s value function, an exponential-polynomial density adapted from Algan, Allais and Den Haan for the distribution), computation of the stationary equilibrium without aggregate shocks, and a locally accurate Taylor expansion in the aggregate state, implemented directly in Dynare with a published code template. Demonstrated on the Khan and Thomas (2008) model of heterogeneous firms with fixed capital adjustment costs, a first-order solution takes 35 seconds for a degree-2 approximation of the distribution and 47 seconds for degree 4, most of that in the stationary equilibrium. Accuracy is reported by degree: degree 2 already gives steady-state aggregates “virtually indistinguishable” from a fine histogram (output 0.498 against 0.498, capital 1.006 against 1.007), while degree 1 is visibly wrong (capital 1.200), and higher degrees are needed only for the shape of the distribution, not the aggregates — cross-sectional covariance and variance of log capital do shift with the degree of approximation, though “these differences are small and do not generate differences in the dynamics of aggregate variables.” Business cycle moments are conventional: output volatility 2.14 percent, consumption 0.48 times as volatile as output, investment 3.86 times, all highly correlated because TFP is the only shock. Three generalizations establish the claimed breadth. A second-order approximation is a one-line Dynare change and shows quantitatively small sign- and state-dependence, consistent with Khan and Thomas; mass points from occasionally binding constraints are handled by adding the mass at the constraint as an extra distributional parameter, which is how the Krusell and Smith (1998) household model is solved; and approximate aggregation is shown not to be required, with a volatile-enough investment-specific shock breaking it. The efficiency claim is externally sourced: in Terry’s comparison project, an implementation of this method “solves and simulates the model in 0.098% of the time of the Krusell and Smith (1998) method.” That speed is what makes the paper’s estimation exercise possible: 10,000 Metropolis-Hastings draws in 44 minutes and 32 seconds, estimating the parameters of neutral and investment-specific productivity shocks conditional on three values of the fixed-cost upper bound. As frictions rise from 0.01 to 1, the estimated volatility of investment-specific shocks rises from 0.0056 to 0.0077 while the other parameters barely move, so “matching the aggregate investment data with larger adjustment frictions requires more volatile shocks.” The limits are stated rather than implied: the local approximation is “not well suited for capturing nonlinearities that are global in nature,” the parametric family fails for fat-tailed distributions, the estimation conditions on frictions rather than estimating them jointly, and the resulting variance decompositions differ only slightly across the friction values considered.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What problem is the method for, and why is ease of use part of the problem?
The computational problem is that the aggregate state vector contains the distribution of agents, which is typically infinite-dimensional; the adoption problem is that existing solutions are not usable the way Dynare is. The paper states both in sequence. “The heterogeneous agent models studied in this literature are computationally challenging because the aggregate state vector contains the distribution of microeconomic agents, which is typically an infinite-dimensional object. Although many numerical algorithms have been developed to overcome this challenge, none are as general, efficient, or easy-to-use as the standard perturbation methods routinely employed to solve representative agent models using prepackaged toolboxes like Dynare. Due to this challenge, heterogeneous agent models have yet to reach widespread adoption, particularly among central banks and policy institutions.” The stated goal is correspondingly threefold — “general, efficient, and easy-to-use” — and the paper treats providing “a detailed online code template which explains how to do so” as part of the contribution rather than an appendix.
Q2. What is the fixed point that makes these models hard?
Each agent’s decision depends on its expectation of how the distribution will evolve, and the distribution evolves according to those decisions — an infinite-dimensional fixed point. In the demonstration model, “the aggregate state vector of the model contains the distribution of firms over idiosyncratic productivity and capital, which evolves over time in response to aggregate productivity shocks. The dynamics of the distribution must satisfy a complicated fixed-point problem: each firm’s investment decision depends on its expectation of the dynamics of the distribution, and the dynamics of the distribution depend on firms’ investment decisions. This infinite-dimensional fixed-point problem is at the heart of the computational challenges faced by the heterogeneous agent literature.”
Q3. What are the three steps of the method?
Approximate the equilibrium objects in the individual states with globally accurate projections; solve the stationary equilibrium without aggregate shocks; then take a locally accurate Taylor expansion in the aggregate state. The division of labour between global and local approximation is the design choice: the value function and distribution are approximated “with respect to individual state (ε, k) using globally accurate projection methods” and “with respect to the aggregate state (z, g) using locally accurate perturbation methods in Step 3.” The reason accuracy is affordable in the first dimension and only local in the second is stated directly: “An accurate approximation of the distribution may require a large number of parameters, so I solve for the aggregate dynamics of the model using locally accurate perturbation techniques (which are computationally efficient even when the aggregate state vector is large).” Once the approximations are in place, the equilibrium conditions reduce to a residual function that “is exactly the canonical form studied by Schmitt-Grohe and Uribe (2004),” so “the remainder of the computational method simply follows the steps outlined” there — and Dynare “implements this perturbation procedure completely automatically.”
Q4. How is the distribution approximated, and what does the degree of approximation buy?
By an exponential-polynomial density adapted from Algan, Allais and Den Haan, whose degree controls how far from normal the approximated distribution can be. The family is an exponential of a polynomial in deviations of productivity and log capital from their means, with centralized moments as the parameters; “the key difference from Algan, Allais, and Haan (2008) is that the distribution is over a two-dimensional vector rather than a univariate one.” Degree maps to shape: “a n_g = 2 degree polynomial in this family corresponds to a multivariate normal distribution; n_g > 2 polynomials allow for nonnormal features, such as skewness or excess kurtosis.” A footnote records a suggestion the paper declines to pursue — reducing dimensionality further using copulas to link two univariate distributions — and credits an anonymous referee for it.
Q5. Why do moments end up being the state variable?
Because the family’s parameters can be recovered from its moments, so the moments alone characterize the approximated density, and their law of motion can be computed by integration. Given a moment vector, Algan, Allais and Den Haan’s procedure solves for the associated density parameters, so “the vector of moments m completely characterizes the approximated density. I therefore use the moments m as the characterization of the distribution, and approximate the infinite-dimensional aggregate state (z, g) with the finite-dimensional representation (z, m).” The law of motion follows: “The system (9) provides a mapping from the current aggregate state into next period’s moments m’(z, m) by integrating decision rules against the implied density.” Integrals are computed by two-dimensional Gauss-Legendre quadrature, with a practical warning that “higher degree approximations of the distribution n_g require more points in the quadrature in order to accurately compute the integrals.” Steady-state moments come from iterating the mapping, and the paper reports the status of that step honestly: “Since the mapping is nonlinear there is no guarantee that this iteration converges, but I have found that it does in practice.”
Q6. What model is the method demonstrated on, and how is it calibrated?
The Khan and Thomas (2008) model — a real business cycle economy with firm heterogeneity in productivity and capital plus fixed capital adjustment costs — at Khan and Thomas’s own parameterization, with an annual period. The firm’s value function is approximated by Chebyshev polynomials solved by collocation on a grid, with the expectation over idiosyncratic shocks taken explicitly by Gauss-Hermite quadrature and the expectation over aggregate shocks left to the perturbation step. Parameter values follow Khan and Thomas’s Table 1 “adjusted for the fact that my model does not contain trend growth,” with the labour disutility parameter chosen so steady-state labour supply is one third, and with “the firm-level adjustment costs and idiosyncratic shock process… chosen by Khan and Thomas (2008) to match features of the investment rate distribution reported in Cooper and Haltiwanger (2006).” The paper also notes that its choices of orthogonal and tensor-product polynomials are conveniences rather than requirements — linear splines, complete polynomials and Smolyak polynomials have all been used instead.
Q7. How accurate is the approximation, and does the required degree depend on what you want to measure?
Yes, and the paper separates the two cases: degree 2 suffices for aggregates, degree 4 is needed for the shape of the distribution. On shapes, “the marginal distribution of productivity in Panel (a) is normal, so a n_g = 2 degree approximation gives an exact match to the histogram. In contrast, the marginal distribution of capital in Panel (b) features positive skewness and excess kurtosis, and the conditional distribution of capital varies in both location and shape as a function of productivity. A n_g = 4 degree approximation captures these complicated shapes. Even a n_g = 2 approximation also does quite well, indicating that the true distribution is close to log-normal.” On steady-state aggregates, degree 2 gives output 0.498, consumption 0.412, capital 1.006, wage 0.958 and marginal utility 2.426 against histogram values of 0.498, 0.413, 1.007, 0.958 and 2.421 — “virtually indistinguishable from higher degree approximation or the histogram” — whereas degree 1 is clearly inadequate, giving output 0.509, consumption 0.555, capital 1.200 and marginal utility 1.803. The introduction states the distinction as a headline: “Although degrees of approximation on the low end of this range are sufficient to capture the dynamics of aggregate variables, degrees on the high end are necessary to capture changes in the shape of the distribution over time.”
Q8. How long does it take, and where does the time go?
Between 35 and 47 seconds for a first-order solution, mostly in the stationary equilibrium. “Using Dynare, computing a first-order approximation of the model’s dynamics takes between 35 seconds for a n_g = 2 degree approximation of the distribution and 47 seconds for a n_g = 4 degree approximation. The majority of time is spent computing the stationary equilibrium of the model. The remaining time is spent evaluating the derivatives of the model’s equilibrium conditions f at steady state and solving the resulting dynamic system.” Runtimes are on Dynare 4.5.3 in Matlab R2016a on a 3.10 GHz Windows desktop with 32 GB of RAM. A footnote adds an important qualification for estimation use: “This running time overstates the time it takes to solve the model in each step of an estimation procedure because it includes the time spend processing the model and taking symbolic derivatives. Once these tasks are complete, they do not need to be performed again for different parameter values.”
Q9. Are the solved dynamics right?
They match what other algorithms produce on the same model, and the paper leans on independent verification for the claim. “The dynamics of aggregate variables are well in line with what Khan and Thomas (2008) and Terry (2017b) have reported using different algorithms to solve the model,” and separately, “Terry (2017a) finds that the aggregate dynamics implied by my method are quantitatively close to the results of other methods used in the literature, such as Krusell and Smith (1998).” The reported moments are standard real-business-cycle values: output volatility 2.14 percent at degree 2 and 2.16 percent at degree 4, with consumption 0.48 times as volatile as output, investment 3.86 times, hours 0.61, the real wage 0.48 and the real interest rate 0.08; correlations with output are 0.90 for consumption, 0.97 for investment, 0.95 for hours, 0.90 for the real wage and 0.80 for the real interest rate. All variables are HP-filtered with smoothing parameter 100. The paper explains the uniformly high correlations mechanically rather than substantively: “All variables are highly correlated with each other because aggregate TFP is the only shock driving fluctuations in the model.”
Q10. Do distributional dynamics behave as well as aggregate dynamics?
Not quite — they are more sensitive to the degree of approximation, though not enough to move the aggregates. “The dynamics of cross-sectional features of the distribution are more sensitive to the degree of approximation n_g than are the dynamics of aggregate variables… Although the response of mean log capital is virtually identical across the degrees of approximation n_g, the responses of the other two statistics change with n_g. However, these differences are small and do not generate differences in the dynamics of aggregate variables described above.” The reason aggregates are insulated is identified precisely: what firms need to forecast is the level of marginal utility, and “the impulse response of marginal utility is virtually identical for the different degrees of approximation.” The tabulated cross-sectional statistics confirm it — the correlation of Cov(ε, log k) with output rises from 0.7157 at degree 2 to 0.8432 at degree 4, and that of Var(log k) from 0.5752 to 0.6539, but marginal utility’s correlation with output is −0.8999, −0.9001 and −0.8999 across degrees 2, 3 and 4.
Q11. What does the second-order approximation show, and where does the local approach break down?
Aggregate nonlinearities are quantitatively small in this model, and the paper names three specific classes of problem the method cannot handle. Second order is “changing a Dynare option from order = 1 to order = 2,” and the result is that “the dynamics of the second-order approximation closely resemble those of the first-order approximation,” with “quantitatively small sign-dependence and state-dependence” in impulse responses computed nonlinearly as in Koop, Pesaran and Potter — “consistent with Khan and Thomas (2008), who find little evidence of such nonlinearities using the nonlinear Krusell and Smith (1998) algorithm.” The limitations are stated in the same subsection rather than deferred: the method “is not well suited for capturing nonlinearities that are global in nature. For example, portfolio choice problems in which assets differ only in their aggregate risk characteristics cannot be directly solved using the Dynare code template because the stationary distribution of portfolios is not unique. In addition, the method is ill-suited to solve models which feature the economy transitioning between multiple steady states, such as Brunnermeier and Sannikov (2014).” One scope clarification matters here and the paper supplies it: “these limitations apply to aggregate dynamics only; the method does compute a fully global approximation of individual behavior. Hence, the method can be used to solve models in which aggregate shocks affect the distribution of idiosyncratic shocks, such as the uncertainty shocks in Bloom, Floetotto, Jaimovich, Saporta-Eksten, and Terry (2018).” A further restriction attaches to the parametric family itself: it “requires that moments of the distribution exist up to the order of the approximation, which is restrictive for models in which the distribution features a fat tail.”
Q12. In what sense does the method not require approximate aggregation, and why does that matter?
Because the distribution enters the aggregate state vector directly rather than being summarized by a forecasting rule, so the method works whether or not a few moments predict prices well. For the demonstration model approximate aggregation does hold — Khan and Thomas “solve the model by approximating the distribution with the aggregate capital stock and find that their solution is extremely accurate. Hence, for this particular model, my method and the Krusell and Smith (1998) method are both viable.” The efficiency comparison is externally sourced and stark: “in a comparison project, Terry (2017a) shows that his implementation of my method solves and simulates the model in 0.098% of the time of the Krusell and Smith (1998) method. This gain in speed makes the full-information Bayesian estimation in Section 5 feasible.” The generality claim is then demonstrated by construction: adding volatile enough aggregate investment-specific shocks makes it so that “the aggregate capital stock does not accurately approximate how the distribution affects aggregate dynamics. Extending Krusell and Smith (1998)’s algorithm would therefore require adding more moments to the forecasting rule, which quickly becomes infeasible since each additional moment adds another state variable.” The trade-off is stated symmetrically earlier on: Krusell and Smith’s advantage “is that it is globally accurate with respect to this reduced aggregate state and, therefore, can capture global nonlinearities more easily than the locally accurate approach pursued in this paper.”
Q13. How are mass points handled?
By adding the mass at the constraint as one extra distributional parameter and applying the polynomial family only away from the constraint. The demonstration model’s distribution has a density, which makes smooth polynomial approximation easy, but “in other models, the distribution may not be characterized by its density; for example, in the Krusell and Smith (1998) model, the occasionally binding borrowing constraint introduces mass points into the distribution of households.” The fix: “I simply add an additional parameter to the distribution approximation — the mass of households at the borrowing constraint — and approximate the distribution away from the constraint with the polynomial family. In principle, one could extend this procedure to incorporate multiple mass points as well.” A footnote is careful about what is being approximated: “strictly speaking, the distribution in the Krusell and Smith (1998) model is a collection of mass points, and I approximate the collection of points away from the borrowing constraints with a smooth polynomial.”
Q14. What exactly is estimated, and what is held fixed?
The five parameters of two aggregate shock processes, estimated separately conditional on each of three values of the fixed-cost upper bound — not jointly with it. The model is extended with investment-specific productivity shocks, which enter only the capital accumulation equation, following a joint process with a loading of neutral innovations onto the investment-specific shock. That loading is not incidental: “Without this loading factor, investment-specific shocks would induce a counterfactually negative comovement between consumption and investment and, therefore, be immediately rejected by the data.” The fixed-cost upper bound varies from 0.01 through 0.1 to 1, spanning micro-level investment behaviour “from frictionless to extreme frictions,” with the remaining parameters fixed at standard values and the model frequency switched to quarterly to match the data. The observables are log-linearly detrended real output and real investment. Posterior sampling is by Metropolis-Hastings, and the paper reports that “Dynare computes 10,000 draws from the posterior in 44 minutes, 32 seconds” — the introduction states the figure as “less than 42 minutes,” a discrepancy within the paper worth noting. A footnote explains why estimation is cheap: “the parameters of the shock processes do not affect the stationary equilibrium of the model. Hence, the stationary equilibrium does not need to be recomputed at each point in the estimation process.”
Q15. What does the estimation find?
Larger micro-level frictions require a more volatile investment-specific shock process to match the same aggregate data, with the other estimated parameters essentially unchanged. “As the upper bound of the fixed costs increases from ξ = 0.01 to ξ = 1, the estimated variance of investment-specific shocks significantly increases from σ_q = 0.0056 to σ_q = 0.0077; hence, matching the aggregate investment data with larger adjustment frictions requires more volatile shocks.” The 90 percent highest-posterior-density sets are [0.0047, 0.0067] and [0.0059, 0.0093] respectively — overlapping, which is worth carrying alongside the word “significantly.” The loading also shifts, from −0.0045 to −0.0037, “because larger frictions reduce the negative comovement between consumption and investment.” Everything else moves negligibly: the neutral TFP persistence is 0.9793, 0.9797 and 0.9789 across the three cases and its volatility 0.0080 in all three, so “micro-level adjustment frictions mainly matter for the inference of the investment-specific shock process.”
Q16. Does the shift in estimated shock volatility change the substantive picture?
Only slightly, and the paper says so. The variance decomposition with flexible adjustment attributes 84.73 percent of output variance, 91.21 percent of consumption variance and 35.08 percent of investment variance to neutral shocks; with extreme frictions the corresponding figures are 84.08, 90.56 and 34.27 percent. “Although the investment-specific shock accounts for a slightly larger share of fluctuations with large adjustment frictions (ξ = 1), the differences are quantitatively small.” And the exercise’s own limitation is flagged twice, in Section 5 and in the introduction: “Of course, the ideal estimation exercise would jointly estimate the capital adjustment frictions and aggregate shock processes using both micro- and macro-level data. The illustrative results in this section show that, using my method, such exercises are now feasible.” The word carrying the weight in the paper’s abstract is correspondingly hedged — the micro behaviour “quantitatively shapes inference about the aggregate shock processes, suggesting an important role for micro data.”
Q17. How does the paper place itself relative to the closest methods?
As a parametric middle path: it borrows the density family from Algan, Allais and Den Haan but perturbs instead of projecting, and it borrows the local-in-aggregates idea from Reiter but replaces his histogram with a parametric family. On the first: Algan, Allais and Den Haan “solve for the dynamics of the parameters of the distribution using a globally accurate projection technique, which is computationally slower than the locally accurate perturbation method that I use.” On the second: Reiter (2009), “himself building on an idea of Campbell (1998),” solves the Krusell-Smith model with local approximations in the aggregate state but “approximates the distribution with a fine histogram, which requires many parameters to achieve acceptable accuracy. This choice limits Reiter (2009)’s approach to problems with a low-dimensional individual state space because the size of the histogram grows exponentially in the number of individual states.” Against the continuous-time adaptation of Ahn, Kaplan, Moll, Winberry and Wolf, which reduces the distribution nonparametrically: “My method reduces the distribution in a parametric way, which allows for a more straightforward extension to nonlinear approximations.” The paper also credits precursors in the (S,s) literature, Childers’s formalization in function space, and Veracierto’s alternative of approximating the history of decision rules without any direct approximation of the distribution.
Q18. What is the paper’s closing argument?
That the method’s real payoff is bringing micro data into DSGE estimation, where its restrictions are currently either absent or smuggled in through priors. The conclusion restates the methodological claim — “In contrast to much of the existing literature, the method does not rely on the dynamics of the distribution being well-approximated by a small number of moments, expanding the class of models which can be feasibly computed” — and states the ambition behind releasing code and a user guide: “with the hope that it will bring heterogeneous agent models into the fold of standard quantitative macroeconomic analysis.” The forward-looking claim is where the paper’s interest lies: “A particularly promising avenue for future research is incorporating micro data into the estimation of DSGE models. As I showed in Section 5, micro-level behavior places important restrictions on estimated model parameters. In the current representative agent DSGE literature, these restrictions are either completely absent or imposed only indirectly through prior beliefs. The computational method I developed in this paper instead allows the micro data to formally place these restrictions itself.”
Key terms in this paper
Definitions below follow the paper's own usage.
- Parametric family for the distribution
- The functional form the method substitutes for the true cross-sectional density: an exponential-polynomial family adapted from Algan, Allais and Den Haan (2008), here extended from a univariate to a bivariate state (idiosyncratic productivity and capital). Its degree n_g sets the flexibility -- "a n_g = 2 degree polynomial in this family corresponds to a multivariate normal distribution; n_g > 2 polynomials allow for nonnormal features, such as skewness or excess kurtosis." The approximation carries a stated restriction: it "requires that moments of the distribution exist up to the order of the approximation, which is restrictive for models in which the distribution features a fat tail. In those cases, a different parametric family should be used."
- Moments as endogenous state variables
- The consequence of the parametric approximation that makes the rest of the method work: because the family's parameters are recovered from its centralized moments by Algan, Allais and Den Haan's procedure, "the vector of moments m completely characterizes the approximated density," so the infinite-dimensional aggregate state (z, g) can be replaced by the finite (z, m) and the law of motion for the distribution by a law of motion for the moments, obtained by integrating decision rules against the implied density. Steady-state moments are found by iterating this mapping, with a caveat the author states rather than hides: "Since the mapping is nonlinear there is no guarantee that this iteration converges, but I have found that it does in practice."
- Global in individual states, local in aggregate states
- The method's deliberate split of approximation strategies: globally accurate projection methods (Chebyshev collocation for the value function, the parametric family for the distribution) with respect to individual state variables, and locally accurate Taylor expansion with respect to the aggregate state. This is what makes a large aggregate state vector affordable, since perturbation is "computationally efficient even when the aggregate state vector is large." The author is explicit about where the local half fails -- global nonlinearities such as portfolio choice with non-unique stationary portfolio distributions, or economies transitioning between multiple steady states -- while noting the limitation "appl[ies] to aggregate dynamics only; the method does compute a fully global approximation of individual behavior."
- Approximate aggregation (and independence from it)
- Khan and Thomas's property, following Krusell and Smith, that the aggregate capital stock alone almost completely captures how the firm distribution influences aggregate dynamics -- and the assumption the paper's method does not need. The paper's position is comparative rather than dismissive: where approximate aggregation holds "my method and the Krusell and Smith (1998) method are both viable," but the method is faster, and where it fails "extending Krusell and Smith (1998)'s algorithm would therefore require adding more moments to the forecasting rule, which quickly becomes infeasible since each additional moment adds another state variable." The paper demonstrates the failure case by making investment-specific shocks volatile enough.
- Micro frictions shaping inference about aggregate shocks
- The paper's substantive economic finding, obtained by estimating the aggregate shock processes conditional on three values of the fixed-cost upper bound rather than jointly with it: as micro-level adjustment frictions rise, matching the same aggregate investment data requires a more volatile investment-specific shock process. The paper frames this as a restriction that micro data can impose on DSGE estimation, noting that in the representative-agent literature such restrictions "are either completely absent or imposed only indirectly through prior beliefs," while conceding that "the ideal estimation exercise would jointly estimate the capital adjustment frictions and aggregate shock processes using both micro- and macro-level data."