Macro Paper Warehouse
Published Classic [NBER Macroeconomics Annual] doi:10.1086/696046

When inequality matters for macro and macro matters for inequality

SeHyoun Ahn — Princeton University

Greg Kaplan — University of Chicago; NBER

Benjamin Moll — Princeton University; NBER

Thomas Winberry — University of Chicago Booth School of Business; NBER

Christian Wolf — Princeton University

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Macroeconomists have long defended single-household models with two excuses: models with many households are too hard to solve, and inequality does not matter much for aggregates anyway. This 2018 paper attacks both. It presents a fast continuous-time method -- solve the steady state exactly, then linearize around it, shrinking the resulting system -- that solves the standard benchmark model roughly 1,500 times faster and more accurately than earlier algorithms, though as a local approximation its accuracy degrades as aggregate shocks grow. Applied to a two-asset model matched to U.S. household balance sheets, inequality can matter greatly for aggregates, and aggregate shocks can substantially reshape inequality.

What this paper finds — and why it matters

This paper argues that two standard excuses for relying on representative-agent macro models – that heterogeneous-agent models with aggregate shocks are computationally intractable, and that realistic household heterogeneity does not matter much for aggregate dynamics anyway – are both weaker than commonly believed. To address the first, the authors extend Michael Reiter’s discrete-time linearization approach to continuous time: they solve a model’s stationary equilibrium fully nonlinearly using the finite-difference methods of Achdou et al. (2015), take a first-order Taylor expansion of the full discretized equilibrium system around that steady state using automatic differentiation, and solve the resulting large linear system of stochastic differential equations by standard techniques, exploiting continuous time’s tendency to generate sparse transition matrices; a companion model-free dimensionality-reduction method, adapted from the engineering “model reduction” literature, lets the computer – rather than the researcher – identify the low-dimensional information in the cross-sectional distribution needed to forecast prices accurately, generalizing the “approximate aggregation” logic of Krusell and Smith (1998) beyond cases where a hand-picked set of moments happens to work. On the standard Krusell-Smith (1998) business-cycle model, the method is roughly 1,500 times faster and about three times more accurate (by Den Haan’s 2010 error metric) than the best-performing algorithm in the JEDC comparison project, though its accuracy – being a local, linear approximation – degrades as the size of aggregate shocks grows. To address the second excuse, the authors apply their (open-sourced) toolbox to a two-asset incomplete-markets model, calibrated to match the U.S. joint distribution of income, wealth, and marginal propensities to consume, in which “wealthy hand-to-mouth” households arise endogenously from a costly-to-access illiquid asset. This richer model jointly reproduces two features of aggregate consumption dynamics – sensitivity to predictable income changes and relative smoothness – that have long challenged representative-agent and simple spender-saver benchmarks, and, in an extension with capital-skill complementarity, shows that aggregate productivity shocks can generate substantial, shock-specific swings in income and consumption inequality, providing what the authors call “a striking counterexample to the main result of Krusell and Smith (1998).”

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What two “excuses” for representative-agent modeling does the paper set out to challenge?

The paper states its main message directly: “both of these excuses are less valid than commonly thought” – the claim that heterogeneous-agent models with aggregate shocks are computationally intractable, and the claim, attributed in part to a reading of Krusell and Smith (1998), that realistic heterogeneity “generates only limited additional explanatory power for aggregate phenomena” (Introduction, pp. 1-2). The paper quotes Lucas’s (2003) summary of the conventional wisdom – “For determining the behavior of aggregates, [Krusell and Smith] discovered, realistically modeled household heterogeneity just does not matter very much” – and notes a “discrepancy between this perception and the results in Krusell and Smith (1998)” themselves, since their own extension with preference heterogeneity “features aggregate time series that depart significantly from permanent income behavior” (p. 1, fn. 3).

Q2. What are the three steps of the paper’s core linearization method?

The method (Section 2.2) proceeds in three steps: first, solve for the stationary equilibrium of the model without aggregate shocks, using the finite-difference method of Achdou et al. (2015) to discretize the household’s Hamilton-Jacobi-Bellman equation and the associated Kolmogorov Forward equation for the cross-sectional distribution; second, take a first-order Taylor expansion of the full discretized equilibrium conditions – value function, distribution, and prices – around this steady state, using automatic differentiation because “the size of the system is large” enough that derivatives “by hand” are infeasible; third, solve the resulting linear system of stochastic differential equations via a standard Schur decomposition, checking the Blanchard and Kahn (1980) condition for a unique stable solution. The paper explicitly frames the underlying logic as a re-derivation of Reiter’s (2009) insight – “our method builds on the ideas of Dotsey, King and Wolman (1999), Campbell (1998), and Reiter (2009)” – cast in continuous time (Section 2, pp. 9-10).

Q3. Why does the paper work in continuous time rather than discrete time, given that Reiter’s original method was discrete-time?

The paper gives three reasons continuous time is “heavily exploited” (Section 2, “Continuous Time,” p. 6): it is easier to represent occasionally binding constraints – the borrowing constraint “is absorbed into a simple boundary condition on the value function,” so the consumption first-order condition “holds with equality everywhere in the interior of the state space”; first-order conditions for optimal policy are typically simpler, “often solved by hand”; and, “most importantly in practice, continuous time naturally generates sparsity in the matrices characterizing the model’s equilibrium conditions,” because continuously moving state variables only drift infinitesimally in infinitesimal time, so households can reach only neighboring grid states. The paper notes that discrete-time models can also generate sparsity (citing Reiter’s 2010 discussion), but typically only with very short time periods or with denser/higher-bandwidth matrices than the continuous-time approach produces, and their own two-asset model is “so large that sparsity is necessary to store and manipulate these matrices” at all (Section 2, footnote 11).

Q4. What does the linearized solution capture about heterogeneity, and what does it explicitly rule out?

The method preserves full nonlinearity at the individual level and “distributional dependence” – the impulse response of an aggregate depends on the initial cross-sectional distribution because individual policy elasticities to the aggregate shock differ across the state space (Section 2.3, pp. 13-14) – but, being a first-order approximation in the aggregate state, it imposes certainty equivalence with respect to aggregate shocks (the standard deviation of the aggregate shock does not enter decision rules) and rules out any size- or sign-dependence in the aggregate response: “our linearization method eliminates any potential sign- and size-dependence” (Section 2.3, pp. 14-15). The paper is explicit that this makes the method “less suitable for various asset-pricing applications in which the direct effect of aggregate uncertainty on individual decision rules is key,” and flags extending the perturbation to higher orders, or allowing nonlinear dependence on low-dimensional (but not the full high-dimensional) aggregate states, as directions for future work (Section 2.3, p. 14).

Q5. How fast and accurate is the method on the standard Krusell-Smith (1998) test case?

Solving the Krusell-Smith model takes about a quarter of a second (0.27 seconds total: steady state, derivatives, linear-system solution, and IRF simulation combined), versus “over seven minutes” for the fastest algorithm in the JEDC comparison project of Den Haan (2010) – “more than 1500 times slower” (Section 2.4, Table 1, pp. 16-17). Using Den Haan’s (2010) error metric – the maximum log difference between aggregate capital simulated from the linearized solution versus from the model’s true nonlinear dynamics – the method achieves a maximum error of 0.049% at the standard calibration (0.7% shock standard deviation), which the paper reports is “three times as accurate as the Krusell and Smith (1998) method,” the most accurate algorithm in that comparison, at 0.16% (Section 2.4, Table 2, p. 17). Because the method is a local approximation, this accuracy “decreases in the size of the shocks,” rising to a 3.282% maximum error at a 5% shock standard deviation (Table 2).

Q6. What is the model-free dimensionality-reduction method, and why is it needed?

For larger models – such as the paper’s two-asset application, whose unreduced system exceeds 120,000 equations – solving the full linear system is “prohibitively expensive,” so the paper develops a method to project the high-dimensional distribution (and value function) onto a low-dimensional basis while preserving the accuracy of the resulting forecasts for prices (Section 3, pp. 18-21). Rather than positing, as Krusell and Smith (1998) did on economic grounds, that a specific small set of moments (the mean) suffices, the method borrows techniques from the engineering “model reduction” literature (citing Antoulas, 2005, and Amsallem and Farhat, 2011) to let the algorithm discover an appropriate basis directly, which the paper describes as generalizing “Krusell and Smith (1998)’s insight that only a small subset of the information contained in the cross-sectional distribution… is required to accurately forecast the variables agents need to know… in a completely model-free way” (Section 3.1, p. 19). This reduction, applied to the Krusell-Smith model itself, further cuts solve time to roughly 0.1 seconds while preserving accuracy (Section 2.4, p. 17).

Q7. What is the two-asset model used to demonstrate that “inequality matters for macro,” and what does it match?

The two-asset model (Section 4, building on Kaplan and Violante, 2014, and Kaplan, Moll, and Violante, 2016) lets households save in a low-return liquid asset and a higher-return illiquid asset subject to a transaction cost, calibrated to five targeted moments of the U.S. Survey of Consumer Finances 2004 – mean liquid and illiquid wealth, and the fractions of poor hand-to-mouth, wealthy hand-to-mouth, and negative-liquid-wealth households (Table 6, p. 41). The calibrated model matches these targets closely and, as a result, reproduces a bimodal distribution of marginal propensities to consume – an average quarterly MPC out of a $500 windfall of 22.5%, “composed of high MPCs for hand-to-mouth households (around 0.4) and small MPCs for non-hand-to-mouth households” – consistent with the empirical evidence in Johnson, Parker, and Souleles (2006), Parker et al. (2013), and Fagereng, Holm, and Natvik (2016) (Section 4.2, pp. 42-43). Solving this much larger model (N=60,000 individual grid points, reduced from over 120,000 to roughly 2,145 knot points for the value function) still takes well under five minutes (Table 7, p. 46) and, the paper states, “to the best of our knowledge, this model cannot be solved using any existing methods” without the reduction step (Introduction, p. 4).

Q8. What two features of aggregate consumption dynamics does the two-asset model jointly match, and why can’t a representative-agent model do the same?

The model jointly matches “sensitivity” – the tendency of aggregate consumption growth to respond to predictable changes in income growth – and “smoothness” – the fact that consumption growth is markedly less volatile than income growth – two facts that “have proven to be a challenge for representative agent models” in the literature (Campbell and Mankiw, 1989; Christiano, 1989; Ludvigson and Michaelides, 2001) (Section 5.2, Table 9, pp. 50-52). The mechanism is the wealthy-hand-to-mouth households: because their consumption tracks current income (which reacts less on impact but more persistently than permanent income to the productivity growth shock), their response is smaller on impact but longer-lasting than a representative household’s, “generat[ing] autocorrelation in consumption which allows the model to match the fact that consumption responds even to predictable changes in income” – while the representative-agent version of the same model instead generates “too little sensitivity once we condition on the real interest rate” (Section 5.2, pp. 50-51).

Q9. What does the “wealthy hand-to-mouth” impulse response show at the individual level?

In response to a positive aggregate productivity shock, the consumption of hand-to-mouth households “responds twice as much as average consumption upon impact, but dies out more quickly,” because these households respond mainly to the (larger but less persistent) change in current income, while non-hand-to-mouth households respond mainly to the (smaller but more persistent) change in permanent income driven by the dynamics of the capital stock (Section 4.4, Figure 10, pp. 47-48). The paper states this differential response is precisely “[d]ue to the presence of these hand-to-mouth households” that “the impulse response of aggregate consumption to a productivity shock is very different in the two-asset model compare[d] with a representative agent model” (p. 47).

Q10. How does the paper’s second application show that macro shocks reshape inequality, and how does this relate to Krusell and Smith (1998)?

Extending the production side to include imperfectly substitutable high- and low-skill labor with capital-skill complementarity (Section 6), the paper shows that a negative shock to unskilled-labor productivity produces a consumption trough “more than twice as low in the two-asset model than in the representative agent model,” in part because the shock is “concentrated among low-skill workers who are more likely to be hand-to-mouth,” and that a positive shock to capital-specific productivity raises skilled wages by more than unskilled wages, “increas[ing] labor income inequality” through capital-skill complementarity (Sections 6.3-6.4, pp. 57-58). The paper explicitly frames these results as running counter to a common gloss on Krusell and Smith (1998): “the response of aggregate consumption to both of these aggregate shocks differs dramatically from that in the representative agent counterpart, thereby providing a striking counterexample to the main result of Krusell and Smith (1998)” (Introduction, p. 5).

Q11. What limitations or open questions does the paper flag about the linearization approach itself?

Beyond the certainty-equivalence and no-size/sign-dependence limitations already noted (Q4), the paper is explicit that its solution’s accuracy is inherently local: it “breaks down for very large shocks” (Section 2.3, p. 15), and the reported Den Haan error rises sharply as shock size increases (Table 2). It also notes that the linearized dynamics rule out “nonlinear amplification effects that result in a bimodal ergodic distribution of aggregate states” of the kind studied by He and Krishnamurthy (2013) and Brunnermeier and Sannikov (2014) (Section 2.3, footnote 27). Methodologically, the paper flags as future work extending the perturbation to higher orders or otherwise allowing nonlinear dependence on low-dimensional aggregate states, and – in its conclusion – observes that the speed gains from the method open the door to bringing heterogeneous-agent models to micro panel data on distributional dynamics, an exercise complicated by the fact that “micro data, especially from surveys, are often inconsistent with national accounts data on macroeconomic aggregates” (Conclusion, p. 59).

Key terms in this paper

Definitions below follow the paper's own usage.

Continuous-time linearization method
the paper's three-step solution procedure for heterogeneous-agent models with aggregate shocks, cast in continuous time: (1) solve the model's stationary equilibrium -- with idiosyncratic but no aggregate shocks -- fully nonlinearly, using the finite-difference method of Achdou et al. (2015) to discretize the household's Hamilton-Jacobi-Bellman equation and the stationary Kolmogorov Forward equation describing the cross-sectional distribution; (2) take a first-order Taylor expansion of the full discretized equilibrium conditions (value function, distribution, prices) around that steady state, computed via automatic differentiation to exploit the sparsity of the discretized transition matrix; (3) solve the resulting large linear system of stochastic differential equations by a standard Schur-decomposition/Blanchard-Kahn approach. The paper frames this as a direct continuous-time extension of Reiter's (2009) discrete-time linearization idea, chosen because continuous time handles occasionally binding constraints more easily, yields simpler first-order conditions, and -- most importantly in practice -- generates sparse transition matrices that make very large models tractable.
Model-free dimensionality reduction
the paper's technique for shrinking the linear system in step (2) above before solving it, developed because the full system becomes "prohibitively expensive" in large models such as the paper's own two-asset application (over 120,000 equations before reduction). Rather than hand-picking a small set of distributional moments the way Krusell and Smith (1998) picked the mean, the method borrows tools from the engineering "model reduction" literature (Antoulas, 2005; Amsallem and Farhat, 2011) to let the computer find a low-dimensional basis that accurately spans the part of the distribution needed to forecast prices, in a way the paper describes as generalizing Krusell-Smith's "approximate aggregation" insight "in a completely model-free way"; the paper credits Reiter (2010) as the first to apply model-reduction ideas of this kind to a linearized heterogeneous-agent model.
Den Haan accuracy/speed benchmark
the paper's finding, using the error metric of Den Haan (2010) -- the maximum log difference between aggregate capital simulated under the linearized solution versus under the model's actual nonlinear dynamics -- that at the standard Krusell-Smith (1998) calibration (0.7% productivity shock standard deviation) their method attains a maximum error of 0.049%, "three times as accurate as the Krusell and Smith (1998) method, which is the most accurate algorithm in Den Haan (2010) and gives eps_DH = 0.16%," while solving the model in about a quarter of a second versus the roughly seven minutes taken by the fastest algorithm in that comparison project -- "more than 1500 times slower" than the paper's method. The paper is explicit that, because the method is a local (linear) approximation, "its accuracy decreases in the size of the shocks," reporting a maximum error of 3.282% once the shock standard deviation is raised to 5%.
Wealthy hand-to-mouth households
households in the paper's two-asset model who hold zero liquid wealth but positive illiquid wealth, and who therefore behave like hand-to-mouth consumers -- setting consumption equal to disposable income -- despite holding substantial net worth, because withdrawing from the illiquid asset is costly; the paper reports that in the calibrated steady state "roughly two-thirds of the hand-to-mouth households are 'wealthy hand-to-mouth'... while the remaining one-third are 'poor hand-to-mouth,'" and that this group is what generates the model's bimodal distribution of marginal propensities to consume (high MPCs, around 0.4, for hand-to-mouth households; low MPCs for others) and, in turn, its ability to jointly match the sensitivity and smoothness of aggregate consumption growth.
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.