Macro Paper Warehouse
Published Classic [Journal of Economic Dynamics and Control] doi:10.1016/j.jedc.2018.01.002

Exploiting MIT shocks in heterogeneous-agent economies: the impulse response as a numerical derivative

Timo Boppart — Institute for International Economic Studies, Stockholm University

Per Krusell — Institute for International Economic Studies, Stockholm University; NBER

Kurt Mitman — Institute for International Economic Studies, Stockholm University

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Models where households differ in wealth are hard to solve: in principle the whole wealth distribution is a state variable a linearization must track. Rather than build a law of motion for that distribution, this 2018 paper computes a single deterministic path, the economy's perfect-foresight response to one small, one-time shock, using only methods already needed for the steady state. Under the working assumption that the true response is close to linear, that path is the model's numerical derivative, so dynamics for any shock sequence follow by scaling and summing copies. The assumption holds up well in their business-cycle model, but the method offers no separate test.

What this paper finds — and why it matters

This paper proposes a new way to compute the equilibrium of heterogeneous-agent models with aggregate uncertainty: rather than building a recursive, linearized law of motion for a high-dimensional state such as the cross-sectional wealth distribution – the strategy behind existing linearization methods including Reiter (2009, 2010), Childers (2017), and Ahn, Kaplan, Moll, Winberry, and Wolf (2017) – the authors solve, using entirely standard nonlinear methods, a single deterministic perfect-foresight transition path that follows a one-time, unexpected (“MIT”) shock away from the model’s non-stochastic steady state. Under the working assumption that the true stochastic equilibrium is well approximated by a linear system, that one nonlinear transition path is literally the model’s numerical derivative at each future time horizon, so the response to any sequence of aggregate shocks – and to any number of independent shocks, with computation time rising only linearly in the number of shocks – can be recovered by scaling and summing copies of this single impulse response, with no analytical differentiation and no explicit treatment of the distribution as a linearized state. The only nontrivial numerical tool the method requires is value-function iteration: once to solve the model’s non-stochastic steady state, and again, backward over a finite horizon, to solve for the transition path, with the passage of calendar time serving as the sole state variable added along the way. Applied to a standard Aiyagari-style economy with valued leisure and two aggregate technology shocks, the method reproduces conditional-moment results close to Dynare’s linearization- and second-order-perturbation-based solutions in a representative-agent benchmark, delivers accurate impulse responses (including for distributional statistics like the Gini coefficient and the hand-to-mouth share) in the heterogeneous-agent case, and, in an extension with a consumption externality and deficit-financed transfers, shows that Ricardian equivalence breaks down and fiscal transfers have real stabilizing effects – all without added numerical difficulty. The paper reports that its own scalability and additivity checks are satisfied “with flying colors” in this application, but is explicit that the whole approach rests on the assumption that linearization is in fact a good approximation for the model at hand, offers no formal stability or determinacy characterization of the transition path it computes, and, unlike recursive methods, provides no separate goodness-of-fit metric for that assumption.

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What is the paper’s central methodological claim?

The paper’s core idea is that a single, nonlinearly computed transition path following a small, one-time aggregate shock can be treated directly as “a numerical derivative in sequence space,” which alone supplies the model’s linearized solution to recurring stochastic shocks (Abstract; Introduction, p. 6). As the paper puts it: “we regard this impulse response path as a numerical derivative in sequence space and hence provide our linearized solution directly using this path.” Because a linear system’s response to any sequence of innovations is the scaled, additive superposition of its response to each individual innovation, the sequence {x0, x1, x2, …} obtained from one nonlinear “MIT shock” transition is, under the assumption that linearization works well, already the complete linear solution – so that “our model with aggregate shocks can be obtained by mere simulation based on the one deterministic path” (p. 7).

Q2. How does this differ from standard linearization methods for heterogeneous-agent models, such as Reiter’s?

Existing linearization methods construct a recursive representation – functions G and H mapping the (typically infinite-dimensional) state, including the wealth distribution Γ, to outcomes and its own law of motion – and then linearize G and H, usually via analytically or automatically computed Taylor terms; the present method “does not rely on direct derivation of first-order Taylor terms” and “does not use recursive methods… whereby aggregates and prices would be expressed as linear functions of the state” (Abstract). The paper explains that the two approaches are mathematically equivalent in what they deliver – “those techniques construct (i) a linear mapping from (log) z and some summary description of the aggregate state… to x and then (ii) a resulting linear law of motion for the summary of the aggregate state. But this delivers precisely a linear system that, in reduced form, gives our simple [additive superposition] … so in this sense the intermediate steps (i) and (ii) are not needed if one can obtain {x0, x1, x2, …}” directly, which the nonlinear transition-path computation does (Section 1.2, p. 9).

Q3. What exactly is an “MIT shock,” and how is the transition path computed?

An MIT shock, a term the paper attributes to Tom Sargent, is “an unpredictable shock to the steady-state equilibrium of an economy without shocks”: the economy starts at its non-stochastic steady state, is hit once by an unexpected shock, and its perfect-foresight transition back to steady state is traced out under the (literal) assumption that no further shock will occur (Section 2, p. 11). For the heterogeneous-agent model, the transition is found by a shooting algorithm (Section 4, pp. 29-30): guess a path for the capital-labor ratio over a long horizon T (350 quarters in the application); solve the household’s value function backward from the terminal steady-state value function; starting from the steady-state distribution, simulate the cross-sectional distribution forward using the resulting policy functions and the productivity Markov chain; compute the capital and labor implied by that simulated distribution; and update the guessed capital-labor path (by a relaxation rule) until the maximum discrepancy is below tolerance. The paper reports that seeding the guess with the representative-agent transition path yields “roughly first-order convergence” and typically fewer than 15 iterations.

Q4. What is the only nontrivial numerical tool the method requires, and how is the steady state itself solved?

“The key numerical tool required to implement it is value-function iteration, using a very limited set of state variables” (Abstract; Section 6, p. 48-49). The steady state is solved using the endogenous grid-point method of Carroll (2006) on a 50-point, non-linearly spaced asset grid, with the idiosyncratic income process discretized to a 7-state Markov chain via Rouwenhorst’s method, and the stationary cross-sectional distribution represented on a finer 1,000-point histogram, recovered as the eigenvector of the transition matrix associated with eigenvalue one (Section 4, p. 29). Solving for the transition path reuses the same value-function-iteration machinery, “now backwards over time, so time is the only additional state variable here,” and the paper stresses “there is no additional conceptual difference between the two cases” of solving a steady state versus a transition (Section 1.2, p. 8).

Q5. How does the method handle multiple aggregate shocks, and what does that cost computationally?

Additional independent shocks are handled by computing one additional deterministic transition path per shock, after which “given the linearity, the effects of the two shocks are simply additive”; the paper states that “computation time rises linearly in the number of independent shocks” (Abstract; Section 1.2, p. 8). In the concluding remarks the authors argue this scales well beyond their own two-shock application: extending the method to a model with, e.g., seven shocks (citing Smets and Wouters, 2003, as an example with price/wage frictions and habits) “is not cumbersome: computing impulse responses to seven shocks is not much more time-consuming t[han] the case of two shocks that we considered here” (Section 6, p. 47-48).

Q6. What checks does the paper use to test whether the linearity assumption actually holds, and what do they find?

Two diagnostics are proposed and run on the heterogeneous-agent benchmark model (Section 5.2.3, pp. 39-40): first, comparing normalized impulse responses to shocks of different sign and size (+/-0.01 and +/-2 standard deviations) for evidence of asymmetry or curvature; second, comparing the impulse response to two shocks hitting simultaneously against the linear sum of their individually computed impulse responses. The paper reports “a striking absence of asymmetry or non-linearity” for the neutral technology shock, and “virtually the same result” for the investment-specific shock, “though here there is a visible deviation with a slightly smaller response for the larger positive shock”; for the joint-shock additivity test, it finds “very little sign of departure from linearity.” The authors are explicit about what a failed check would mean: “if one finds that scalability is not satisfied, the model’s behavior is fundamentally more non-linear and other methods will have to be used, such as those offered in Krusell and Smith (1998)” (Section 6, p. 48).

Q7. How is the method validated in the simpler representative-agent case?

The representative-agent transition path is computed in Dynare using its nonlinear solver, then treated as an impulse-response-based linearization and compared against Dynare’s own standard linearized (first-order) solution and a second-order perturbation solution (Section 5.2.1, pp. 32-34). The paper reports that “our method produces nearly identical correlation statistics to those using Dynare linearization and they are also very close to the values using second-order perturbation” (Table 1). A robustness check (Table 2) varying the length of impulse response used (10, 100, or 500 periods) shows accuracy degrading somewhat with highly persistent shocks when the impulse response is truncated too early, “suggesting that a longer time horizon is necessary for an accurate approximation when using highly persistent shocks,” though a 100-period truncation “still does a decent job.”

Q8. What does the heterogeneous-agent application show, including for distributional outcomes?

Applied to an Aiyagari-style economy with valued leisure and two technology shocks (neutral and investment-specific), the method’s impulse responses for output and aggregates are broadly similar to the representative-agent case for the neutral shock but show substantially more propagation for the investment-specific shock, “with all series noticeably different” (Section 5.2.2, pp. 36-38). A stated advantage of the method is that “we are able to directly simulate interesting statistics from the distribution”: the paper reports steady-state values of 0.77 for the wealth Gini coefficient and 26% for the fraction of borrowing-constrained households, and computes impulse responses and correlations for the Gini coefficient and the hand-to-mouth share alongside the usual aggregates (Table 3, Figure 4).

Q9. What does the policy extension show about Ricardian equivalence, and how does the paper relate this to a common misreading of Krusell and Smith (1998)?

Adding a demand externality (aggregate consumption raises TFP) and a deficit-financed lump-sum transfer rule, the paper finds that “under the deficit-financed transfer policy, aggregate consumption is significantly stabilized relative to laissez-faire” (Section 5.2.4, p. 44-45), even though “the representative-agent version of the present economy would make lump-sum transfers completely ineffective because of Ricardian Equivalence.” The paper explicitly pushes back on the idea that Krusell and Smith (1998) show distribution is generally irrelevant to aggregates: “it is often claimed that the model in Krusell and Smith (1998) ‘shows’ that the distribution (almost) does not matter for outcomes… the opposite is in fact shown in the same paper,” since Krusell-Smith’s own result that the distribution barely matters holds only in the case of common discount factors, whereas with discount-factor heterogeneity – the case that “matches the empirical wealth distribution well” – the distribution does matter (Section 5.2.4, p. 41). Under incomplete markets, transfers redistribute across households with different marginal propensities to consume, which is what breaks Ricardian equivalence in this model.

Q10. What limitations and scope conditions does the paper itself attach to the method?

The method’s central limitation is that it presumes, rather than derives, that the true equilibrium is close to linear: “our approach of using numerical linearization to solve for equilibria relies on nonlinearities not being important in the model at hand,” which the paper calls “a clear weakness,” offset only by the fact that the scalability checks make this assumption “an integral part of the procedure” rather than an unexamined premise (Section 6, p. 48). Two further limitations are flagged explicitly: the method offers no analytical stability or determinacy characterization comparable to Blanchard and Kahn (1980) – “we do not characterize its uniqueness nor whether there are local explosive paths,” leaving “determinacy issues… for future research” (Section 6, p. 49) – and, because the approach is “sequential in nature, as opposed to recursive,” it cannot produce the kind of internal goodness-of-fit statistic (an R^2 on forecasting rules) that recursive methods can, so “a disadvantage of the present method is that it does not offer a metric of fit” beyond the scalability checks themselves (Section 6, p. 49).

Key terms in this paper

Definitions below follow the paper's own usage.

MIT shock
a term the paper attributes to Tom Sargent, referring to "an unpredictable shock to the steady-state equilibrium of an economy without shocks" -- the economy starts at its non-stochastic steady state, is hit once by an unexpected shock, and its perfect-foresight transition back to that same steady state is computed under the assumption that no further shock will ever occur, even though, strictly, further shocks are expected under rational expectations; the paper notes this apparent tension is resolved because, in a (approximately) linear model, the perfect-foresight response to a single unexpected shock coincides with the conditionally expected path of the variable under recurring stochastic shocks.
Impulse response as a numerical derivative
the paper's central methodological device: the nonlinearly computed transition path of any aggregate variable x, in response to a single small period-0 innovation to a shock, is treated directly as "a numerically computed derivative of the initial shock" at each future horizon -- the sequence {x0, x1, x2, ...}. Because a linear system's response to any sequence of shocks is just the scaled, additive superposition of its response to each individual innovation, this single nonlinear impulse response is, under the assumption that linearization is a good approximation, already the complete linearized solution; no separate step of deriving analytical Taylor terms or a recursive linear law of motion for the distribution is needed.
Sequential (shooting) solution versus recursive linearization
the paper's description of its own solution strategy as "sequential in nature, as opposed to recursive, like most other approaches" -- rather than expressing aggregates and prices as functions G and H of a (typically infinite-dimensional) state that must itself be linearized, the method solves a finite sequence of deterministic prices via a shooting algorithm: guess a path for the capital-labor ratio over a long but finite horizon T, solve the household's dynamic program backward from the terminal steady state, simulate the distribution forward from the initial steady state using the resulting policy functions, check whether implied and guessed price paths match, and update the guess (by relaxation) until they do. The only state variable added relative to solving the ordinary stationary equilibrium is time itself.
Scalability / linearity checks
the paper's proposed diagnostics for whether treating the impulse response as a linear derivative is a good approximation: (i) compare impulse responses to shocks of different signs and sizes (e.g. +/-0.01 versus +/-2 standard deviations) normalized by shock size, checking for asymmetry or curvature; and (ii) compare the impulse response to two shocks hitting simultaneously against the sum of the two shocks' individually computed impulse responses, checking additivity. In the paper's own RBC-style heterogeneous-agent economy both checks are "passed... with flying colors," though a "visible deviation" appears for a large positive investment-specific shock; if such checks fail, the paper says the model is "fundamentally more non-linear" and other methods (e.g. Krusell and Smith, 1998) are needed instead.
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.