Solving heterogeneous-agent models by projection and perturbation
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
Models where households differ in wealth and income are hard to solve because the whole cross-sectional distribution of wealth is, in principle, a state variable of infinite dimension. The standard fix, following Krusell and Smith (1998), summarizes it by a few statistics, like its mean. This paper instead solves precisely for the distribution's full shape in the no-aggregate-shock steady state, then takes only a first-order approximation of how that high-dimensional object moves when small aggregate shocks hit. In a test model, the few-moment shortcut works well for productivity shocks alone but breaks down once redistributive tax shocks matter, while the new method stays accurate either way.
What this paper finds — and why it matters
This paper proposes a numerical method for solving stochastic general-equilibrium models with incomplete markets and a continuum of heterogeneous agents – a class of problems where, as the paper puts it, “the state vector includes the whole cross-sectional distribution of wealth,” an infinite-dimensional object in principle. The dominant approach at the time, pioneered by Krusell and Smith (1998), represents that distribution with only a small number of statistics (typically the mean) and works very well for the models it was built for; but the paper argues this cannot serve as a general solution, since in models where the shape of the distribution itself drives the dynamics – such as (S,s) pricing or inventory models, or the redistributive-shock example this paper constructs – a low-dimensional summary can miss essential dynamics. The proposed method instead computes a solution that is fully nonlinear in the idiosyncratic (individual) shocks but only linear in the aggregate shocks: it first solves precisely for the steady-state cross-sectional distribution and consumption function (with no aggregate shocks but the full idiosyncratic shock process), using cubic splines for the consumption function and a fine histogram (up to 1000, and as a robustness check 5000, intervals) for the wealth distribution; it then computes a first-order perturbation of that whole high-dimensional representation with respect to small aggregate shocks, using Sims (2001)’s solver for linear rational-expectations systems. Applied to a test model of household saving with uninsurable income risk, liquidity constraints, an aggregate technology shock, and an i.i.d. redistributive capital-tax shock, the method reproduces the Krusell-Smith “approximate aggregation” finding when only the technology shock is active (a one-moment forecast of future aggregate capital is nearly exact), but shows that this breaks down once the tax shock is introduced, in which case even a four-moment forecast leaves sizable error while the paper’s high-dimensional, spline-based solution remains accurate to roughly 10^-6 in absolute forecast error. The method is explicitly a linear approximation in the aggregate dimension – suited to cases “where individual shocks are much bigger than aggregate shocks” – and the paper proposes it as a first step that can be combined with state-space reduction (via spline or principal-component bases) before, in future work, attempting higher-order perturbations in a reduced aggregate state space.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What computational problem is the paper trying to solve, and why do standard approaches fall short for some models?
The paper targets heterogeneous-agent general-equilibrium models in which “the state vector includes the whole cross-sectional distribution of wealth,” an infinite-dimensional object, and argues that reducing this to a handful of summary statistics – the dominant strategy following Krusell and Smith (1998) – “can hardly serve as a general approach” (Introduction, pp. 1-2). It notes that in models “where firms follow (S,s)-policies (price setting, inventory holdings etc.), we expect that the cross-sectional distribution of the relevant variables enters the solution in an essential way, and the solution algorithm may have to keep track of a medium- or high-dimensional representation of this distribution” (p. 2), motivating a method that does not presuppose a low-dimensional summary is adequate.
Q2. What is the two-step “projection and perturbation” method, in outline?
The method “combines features of projection methods and of perturbation methods,” computing “a solution that is fully nonlinear in the idiosyncratic shocks, but only linear in the aggregate shocks” (Introduction, p. 2; Section 3, p. 10). Step one solves the stationary equilibrium of the economy with the aggregate shocks turned off (constant technology and tax rate) but the full idiosyncratic productivity-shock process retained, using a finite (“projection”) parameterization of the consumption function and the wealth distribution. Step two computes a first-order perturbation of that finite-dimensional system’s dynamics with respect to small aggregate shocks, exploiting the fact that “with our approach we can include a very detailed representation of the cross-sectional distribution in the state vector” – up to 1000 state variables in the paper’s examples (Introduction, p. 2).
Q3. How, concretely, is the steady state approximated – the consumption function and the distribution?
The consumption/savings function is represented by a cubic spline through np+1 “knot points,” with extra points concentrated near the liquidity-constraint kink because “the consumption function has high curvature and is more difficult to approximate with splines there” (Section 3.1.1, pp. 7-9, esp. footnote 3). The cross-sectional wealth distribution is represented as a histogram: its support is truncated at a maximum level and split into nd small intervals (up to nd=1000 in the baseline, nd=5000 for an accuracy check), with the probability mass in each interval stacked into a vector p_t (Section 3.1.2, pp. 8-10). The steady state is found by (1) solving the household’s Euler equation by collocation at the spline knots for a guessed capital stock, (2) computing the stationary histogram as the fixed point p* = Π* p* of the resulting Markov transition matrix via inverse iteration, and (3) using one-dimensional root-finding (bisection or Brent’s method) to find the capital stock consistent with that histogram (Section 3.2, pp. 11-12).
Q4. How is the response to aggregate shocks computed, and what tool does it rely on?
Collecting the full finite-dimensional system – consumption-function spline parameters, distribution histogram, transfers, and the exogenous shocks – into a state vector Θ_t, the paper writes the model’s equilibrium conditions as H(Θ_{t-1}, Θ_t, η_t, ε_t)=0 and linearizes this system around the steady state Θ, computing the needed partial derivatives “by forward differencing” rather than analytic or automatic differentiation, which the paper calls “clearly not computationally efficient… but it turns out to be good enough”* (Section 3.3, pp. 12-13). The resulting linear system is solved using Christopher Sims’s (2001) method for linear rational-expectations models; with nd=1000 state variables “Sims’ package then needs about 30 minutes to solve the model” (p. 13), with most of the time spent reordering eigenvectors.
Q5. What test model is the method applied to, and why does it include a tax shock alongside the usual technology shock?
The test economy is a standard incomplete-markets saving model – a continuum of ex ante identical households facing uninsurable idiosyncratic labor-productivity shocks and a borrowing (liquidity) constraint, with Cobb-Douglas production and CRRA utility – augmented with a government that levies a tax on capital purely “to create some random redistribution of wealth” (Section 2, pp. 4-5; Section 2.2, p. 4). The tax rate follows an AR(1) process chosen i.i.d. in the exercises specifically “to create unpredictable short-run redistributions” (Section 4, p. 15), so that, unlike the technology shock, it shifts the shape of the wealth distribution itself rather than just its scale – deliberately constructing “a laboratory in which one can test a method that uses a high-dimensional representation of the distribution” (Section 5.3.1, p. 19).
Q6. Does the paper confirm or overturn the Krusell-Smith “approximate aggregation” result?
Both, depending on the shock: with only the technology shock active, the paper confirms approximate aggregation almost exactly – “the one-step R2 is 0.9999998” for a forecast of future capital using only its first moment – but once the redistributive tax shock is introduced, “a forecast based on one moment explains very little of the variation in aggregate capital,” and “even if we include 4 moments into the VAR, the forecast error is not always small in relative terms” (Section 5.3.1, pp. 18-19). The paper is explicit that this is a property of the exercise, not a claim that the specific tax-shock design is empirically realistic: “I do not claim that the specification with the big tax shocks is realistic; the purpose of the exercise is only to create a laboratory” for testing a high-dimensional method (p. 19). It also cautions that its moment-forecast numbers are “an upper bound” on what a Krusell-Smith-style solution could achieve, not a direct measurement of that method’s own solution error (p. 19).
Q7. How accurate is the proposed method itself, by the paper’s own error checks?
Two error sources are checked separately: discretization error and state-space-reduction error. The average absolute Euler-equation residual (expressed as a relative consumption error) is on the order of 10^-7, though the maximum residual – concentrated near the liquidity-constraint kink where the consumption function is most curved – is about three orders of magnitude higher, around 10^-4 (Section 5.2, Table 3, pp. 17-18). Comparing the invariant distribution computed with nd=1000 versus nd=5000 intervals, the mean of the wealth distribution “is pinned down quite precisely” while higher moments (skewness, kurtosis) are harder to pin down exactly, though the paper judges accuracy “quite good” given that even minimal transition-matrix differences can compound over an infinite horizon (p. 17). For the reduced, spline-basis solution, the absolute 10-year-ahead forecast error of aggregate capital is “always around 10^-6 or lower” for spline bases of dimension 100, “better than the 4-moment forecast by at least one order of magnitude, often more” (Section 5.3.2, p. 20).
Q8. What is the state-space reduction technique, and how well does it work?
Rather than carrying the full nd-dimensional histogram as a state variable, the paper approximates deviations of the distribution from steady state by p_t ≈ p + B b_t for a chosen basis matrix B (e.g., splines) and a much smaller coefficient vector b_t, replacing the exact transition equation with a weaker projected version (Section 3.4.1, pp. 14-15).* For a further reduction, the paper simulates the model, applies principal-components analysis to the b_t coefficients that actually arise, and uses the leading principal components as a new, even lower-dimensional basis; it reports that “using 5 components gives somewhat better results than 10; using only 3, accuracy would drop by an order of magnitude,” while “using 10 or 30 principal components makes essentially no difference” (Section 5.3.2, p. 20), and cautions that this PCA-based shortcut “will not always work” and is offered mainly as a proof of concept that the linear solution can itself be used to search for good low-dimensional state aggregations (p. 20).
Q9. What limitations and scope conditions does the paper itself attach to the method?
The method is explicitly a first-order (linear) approximation in the aggregate shocks only – “if we are satisfied with a first-order perturbation in the aggregate state variables, as we are in this paper, we can stop at this point” – and the paper flags this as best suited to settings “where individual shocks are much bigger than aggregate shocks, a typical situation in economic applications,” while noting that “even in cases where nonlinearity in aggregate shocks is essential, the linear solution is a useful first step” (Section 3.4.2, p. 15; Introduction, p. 3). It also notes unresolved theoretical ground: existence of the underlying continuum-of-agents stationary equilibrium is not proven for this exact model (“It has not yet been proven that a stationary equilibrium in Θ_t exists”), so the numerical procedure is offered as something that “can probably be seen as a ‘constructive proof’” only for the discretized, approximate economy (Section 2.5, p. 7). The paper’s own suggested next steps – testing the method on asset pricing, and computing genuinely nonlinear-in-aggregate-shocks solutions once a low-dimensional state representation is found adequate – are left to future work (Section 6, p. 20).
Q10. How does the method relate to the other numerical approaches circulating at the time?
The paper positions itself between two existing strands: moment-reduction methods in the tradition of den Haan (1997) and Krusell and Smith (1998), which “try to represent the cross-sectional distribution of wealth by a very small number of statistics,” and full local-perturbation methods such as Kim, Kim, and Kollmann (2005) and Preston and Roca (2006), which “perturb the solution around the deterministic steady state, with no aggregate or idiosyncratic shocks and a degenerate cross-sectional distribution,” a purely local approach the paper says has trouble “handling of inequality constraints, for example liquidity constraints” (Introduction, pp. 1-2). It also notes a close contemporary, Algan, Allais, and Haan (2006), which “uses a flexible parameterization of the cross-sectional distribution and makes heavy use of projection methods,” and credits Bohacek and Kejak (2005) as similar in spirit to its own steady-state step, though “many details of the implementation differ” (Introduction, p. 2, footnote 1). The paper’s distinguishing feature relative to all of these is retaining full nonlinearity in the idiosyncratic dimension while accepting linearity only in the aggregate dimension, which is what permits the high-dimensional (nd up to 1000-5000) distributional state space.
Key terms in this paper
Definitions below follow the paper's own usage.
- Approximate aggregation (Krusell-Smith)
- the finding, due to Krusell and Smith (1998, 2006) and confirmed in this paper's own technology-shock exercise, that in some heterogeneous-agent models "everything one has to know about today's distribution of capital is the mean" for forecasting future aggregate capital with reasonable precision -- "higher moments of the distribution matter very little." The paper treats this as an empirical property of a particular model and shock structure, not a general law: its own tax-shock exercise (Section 5.3.1) shows a one-moment forecast "explains very little of the variation in aggregate capital" once redistributive shocks are added, so the paper's numbers should be read as "an upper bound" -- if a time-series forecast using n moments is already imprecise for the true model, a solution method that restricts households to forecasting with only those n moments cannot be expected to be precise either.
- Projection-and-perturbation method
- the paper's own two-step solution technique, described in the Introduction as computing "a solution that is fully nonlinear in the idiosyncratic shocks, but only linear in the aggregate shocks": first solve the stationary (steady-state) equilibrium without aggregate shocks but with the full idiosyncratic shock process, approximating the consumption function by cubic splines and the cross-sectional wealth distribution by a fine histogram (up to nd=1000 or 5000 intervals); then compute "a first-order perturbation of the solution in the aggregate variables," i.e. linearize the dynamics of that entire finite-dimensional representation around the steady state using Sims (2001)'s method for linear rational-expectations systems.
- Histogram (finite-interval) representation of the distribution
- the paper's discretized description of the cross-sectional wealth distribution, obtained by splitting its support into nd small intervals and stacking the probability mass in each interval into a vector p_t (Section 3.1.2); this histogram, together with the finite set of spline knot points c_t describing the consumption function, becomes literally a (very high-dimensional) state variable of the linearized aggregate model -- "the distribution is characterized by up to 1000 state variables" in the paper's numerical examples, in contrast to the few-moment state vectors used by Krusell-Smith-style methods.
- State-space reduction via basis functions
- the paper's technique, developed as an alternative to carrying the full nd-dimensional histogram as a state variable, of approximating deviations of the distribution from its steady state by a lower-dimensional linear combination of smooth basis functions, p_t ~ p* + B b_t (Section 3.4), where B is a chosen basis (splines, or a set of basis vectors chosen by principal-components analysis of the b_t that actually arise in simulation) and b_t is a much smaller vector of time-varying coefficients; the paper shows that even reducing to a handful of principal components can match the accuracy of the full spline basis if the basis is chosen using the model's own simulated dynamics.