Macro Paper Warehouse
Published Classic Online 21 Dec 2025

Structural Reinforcement Learning for Heterogeneous Agent Macroeconomics

Yucheng Yang — University of Zurich and SFI

Chiyuan Wang — Peking University, School of Computer Science and BIGAI

Andreas Schaab — UC Berkeley

Benjamin Moll — London School of Economics

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Standard models with many households and shocks hitting everyone make each person's decisions depend on the whole population's makeup -- even though people only care about a few market prices. This paper lets people react only to current prices, learning how prices move by simulating the economy forward, much as an animal learns from trial and error -- except each already knows their own budget and income rules, so only price behavior needs learning. That matters: the method solves a savings economy with market-wide shocks -- so hard a research group once gave up on it -- in about a minute, and a similarly hard interest-rate economy in about three.

What this paper finds — and why it matters

Standard recursive formulations of heterogeneous-agent models with aggregate risk force the entire cross-sectional distribution of agents into the Bellman equation characterizing individual decisions – the “Master equation” – purely because low-dimensional equilibrium prices, unlike the distribution itself, do not follow a Markov process, so rational agents forecasting prices end up needing to forecast the whole distribution. This extreme curse of dimensionality remains the central computational bottleneck for global solutions of heterogeneous-agent models, so severe that even a Huggett (1993) model with aggregate risk – despite looking simple – proved impossible for any team to solve in an influential benchmarking exercise and was dropped from the project altogether. This paper sidesteps the Master equation entirely using ideas from reinforcement learning (RL): agents learn equilibrium price dynamics directly from simulated paths, as standard RL would, but the paper’s “structural reinforcement learning” (SRL) approach departs from standard RL by assuming agents have structural knowledge of their own individual-state dynamics (their budget constraint and idiosyncratic income process), letting the authors compute exact policy gradients by differentiating through these known dynamics rather than relying on the noisy, approximate policy gradients standard RL methods estimate; only the equilibrium price process itself is treated as unknown and learned from simulation. By further restricting agents to condition their policies only on current (or briefly lagged) prices rather than the full price history or the distribution, the paper solves for a low-dimensional “restricted perceptions equilibrium” in the sense of Sargent (1991) rather than the full rational-expectations equilibrium – expectations are restricted in functional form but remain statistically consistent with actual outcomes. Because policy functions depend only on prices, they double as individual supply/demand schedules that can be integrated across the distribution and market-cleared period-by-period along a simulation, treating market clearing as part of the “environment” (in RL parlance) rather than something solved inside an optimization loop – which is what lets the method efficiently handle nontrivial market-clearing conditions that have historically been very hard. Implemented in JAX on a single GPU, the resulting structural policy gradient (SPG) algorithm solves the Krusell and Smith (1998) model in about 55 seconds, the previously-unsolved Huggett (1993) model with aggregate risk in around one minute, and a one-asset HANK model with a forward-looking New Keynesian Phillips curve in around three minutes – with the Krusell-Smith solution closely matching alternative global solutions of the rational-expectations equilibrium, and allowing agents a longer history of lagged prices barely moving the solution, indicating most of the information relevant for forecasting prices is already contained in current prices. The paper is explicit that its algorithm, as presented, is not itself intended as an empirically realistic theory of how real economic agents form expectations, though it suggests the “sampling”-based logic behind SRL could in principle be developed into one.

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. Why does the cross-sectional distribution become a state variable in standard heterogeneous-agent models, even though agents don’t directly care about it?

“This difficulty arises even though agents do not directly ‘care about’ the distribution, i.e. it does not enter their objective functions; instead, as is standard in competitive equilibrium models, they only care about prices. Intuitively, low-dimensional equilibrium prices do not follow a Markov process but the extremely high-dimensional distribution does. Therefore, agents with rational expectations forecast prices by forecasting distributions” (Introduction, p. 1). This indirect channel – needing the distribution only to forecast prices, which are themselves Markov only jointly with the distribution – is what the paper’s method is designed to break.

Q2. What, precisely, is “structural” about structural reinforcement learning?

“Our approach differs from standard RL in that we assume that agents have structural knowledge about the dynamics of their own individual states (e.g. their budget constraint and idiosyncratic income process). We term this hybrid approach structural reinforcement learning (SRL)… In contrast to standard policy gradient methods which estimate approximate policy gradients, this assumption allows us to compute exact policy gradients by differentiating through these individual dynamics” (Introduction, pp. 1, 4). “The only part of the environment that is treated as unknown is the process for general-equilibrium prices and aggregate shocks, which is learned from simulated data” – the analogy to structural VARs is explicit: agents’ own laws of motion are structurally known, while the aggregate price process is estimated (Introduction, p. 2, fn. 4).

Q3. What is a “restricted perceptions equilibrium,” and how does the paper justify restricting agents to observe only current prices?

“By the Wold representation theorem, this assumption does not, by itself, imply any departure from full information rational expectations” when agents observe the full history of prices; the further restriction that policy functions depend “only on current prices or, perhaps, a short price history” means the model solves “a restricted perceptions equilibrium (RPE) in the sense of Sargent (1991), Evans and Honkapohja (2001), or Branch (2006): while agents’ expectations are restricted, they are nevertheless statistically consistent with actual equilibrium outcomes, thereby fulfilling one of the desiderata of rational expectations”** (Introduction, p. 2).

Q4. Why was the Huggett model with aggregate risk historically considered so hard to solve, and what does that history reveal about the state of the field?

“According to Maliar and Maliar (2020), in the influential JEDC special issue on numerical solution methods for HA models with aggregate risk (Den Haan et al., 2010), all participants were initially asked to solve two benchmark models globally – the Krusell and Smith (1998) model and the Huggett (1993) model with aggregate shocks… However, as they report, no single team was able to successfully solve the latter model, and it was ultimately dropped from the JEDC project” (Introduction, p. 2, fn. 6). The paper’s ability to solve this exact model in about a minute is presented as direct evidence of the SRL method’s power relative to prior global solution approaches.

Q5. How does the paper’s treatment of market clearing differ from standard practice, and why does that matter for computational efficiency?

“Our treatment of market clearing mirrors the RL literature’s distinction between agents and environment: agents can interact with their environment under any given policy including suboptimal ones; finding optimal policies is conceptually separate. In line with this dichotomy, we treat market clearing as part of the environment and find equilibrium prices also for suboptimal policies. This approach differs from standard practice in macroeconomics which first finds optimal policies given prices in an inner loop” (Introduction, pp. 3-4). Because policy functions depend only on current prices, they “double as individual supply schedules,” so integrating over the distribution to get the aggregate supply schedule makes solving for the period-by-period market-clearing price straightforward along any simulated path – “efficiently handl[ing] non-trivial market clearing conditions like in Huggett (1993)” that other methods have found intractable.

Q6. How does the SPG algorithm’s dimensionality reduction differ from the Krusell-Smith “perceived law of motion” approach?

“Almost all existing global solution methods for heterogeneous agent models use dynamic programming… [either directly on] the high-dimensional Master equation… or, as in Krusell and Smith (1998), they use model-generated data to estimate a low-dimensional Markovian ‘perceived law of motion’ (PLM) for moments of the distribution and then apply dynamic programming to this lower-dimensional approximate problem. Our SPG algorithm instead works directly with the sequential formulation of the problem and never attempts to force it into the standard Markovian structure required for applying dynamic programming… we never estimate a PLM” (Introduction, p. 2). Instead, the method exploits the basic RL insight that value functions are expected discounted lifetime utilities that “can… be approximated by averaging across simulated trajectories,” then finds the low-dimensional policy maximizing expected lifetime utility via stochastic gradient ascent.

Q7. What accuracy checks does the paper report for the HANK application’s market-clearing condition?

Plotting the bond market’s aggregate demand against its (fixed, zero-net) supply over a simulated path, “asset demand fluctuates tightly around the supply level, and deviations from the horizontal supply line measure residuals in the bond market clearing condition. On average, the relative absolute deviation of aggregate demand from supply is 0.22%,” which the authors attribute primarily to the use of linear interpolation when solving for market-clearing prices (section 4, pre-Conclusion). In the Krusell-Smith model, the paper separately reports that “the resulting policies are close to those obtained from alternative global solutions of the rational expectations equilibrium,” and that letting agents condition on a longer lagged-price history “hardly moves the solution,” indicating most forecast-relevant information is already in current prices (Conclusion).

Q8. What does the paper suggest as a longer-run implication of its approach, beyond serving as a fast solution method?

“The core idea of our approach – that agents form price expectations by sampling – could, in principle, serve as a building block for an empirically realistic theory of expectations formation in macroeconomics (which the algorithm in this paper is not),” noting that “the plausibility of RL-based approaches is supported by evidence that reinforcement learning underpins a substantial share of human and animal learning” (Conclusion). The authors are explicit, however, that developing such a theory would require converting the algorithm to a fully online, incremental RL procedure in which agents update policies continuously while interacting with their environment, rather than only after observing many pre-simulated price trajectories as the current SPG algorithm does.

Q9. How does this paper’s approach relate to the Master-equation and other global-solution methods elsewhere in this reading list?

The paper explicitly frames itself as sidestepping the Master equation that underlies Den Haan (1996), Krusell and Smith (1998), Schaab (2020), and Bilal (2023) – three of which appear elsewhere in this reading list – by never representing the distribution as a state variable at all, in contrast to Schaab’s (2020) global finite-dimensional distribution representation and Bilal’s (2023) analytic Master-Equation perturbation (Introduction, p. 1, fn. 1). Relative to the other machine-learning-based global methods on this list, DeepHAM (Han, Yang and E) and Kase, Melosi and Rottner (2022), the key methodological distinction is that this paper’s agents differentiate exactly through their own known individual dynamics rather than learning a full policy or value-function approximation purely from data, and its market-clearing-as-environment design is specifically built to handle the nontrivial market-clearing conditions (like Huggett’s) that those other neural-network approaches do not target as their central test case.

Key terms in this paper

Definitions below follow the paper's own usage.

Structural reinforcement learning (SRL)
The paper's core method: agents learn equilibrium price dynamics from simulated paths, as in standard reinforcement learning (RL), but -- unlike standard RL, which estimates approximate policy gradients -- they exploit known structural knowledge of their own individual state dynamics (e.g. their budget constraint and idiosyncratic income process) to compute exact policy gradients by differentiating directly through those dynamics. Only the process for equilibrium prices and aggregate shocks is treated as unknown and learned from simulation. The specific algorithm is called a structural policy gradient (SPG) algorithm (Introduction, pp. 1-2).
Restricted perceptions equilibrium (RPE) with price-only state variables
The paper's device for keeping the state space low-dimensional: agents are assumed to observe the history of equilibrium prices but not the cross-sectional distribution itself (which, by the Wold representation theorem, need not by itself imply any departure from full-information rational expectations), and their policy functions are further restricted to depend only on current prices or a short price history. This yields a "restricted perceptions equilibrium" (RPE) in the sense of Sargent (1991) and Evans and Honkapohja (2001): expectations are restricted in form, but remain statistically consistent with actual equilibrium outcomes (Introduction, p. 2).
The historically unsolved Huggett-with-aggregate-risk benchmark
The observation, central to the paper's design, that the Huggett (1993) model with aggregate risk (which the authors call "the HANC model with savings through bonds," following Maliar and Maliar 2020) proved so difficult that in an influential JEDC special issue benchmarking global solution methods, "no single team was able to successfully solve the latter model, and it was ultimately dropped from the JEDC project" -- making it, alongside a one-asset HANK model with a forward-looking Phillips curve, one of the two headline test cases the paper's method is shown to solve (Introduction, p. 2, fn. 6).
Market clearing as part of the environment
The paper's approach to handling equilibrium price determination: rather than first finding optimal individual policies given prices in an inner loop and only then imposing market clearing, the method "treat[s] market clearing as part of the environment and find[s] equilibrium prices also for suboptimal policies," mirroring the standard reinforcement-learning distinction between agents (who can act suboptimally while learning) and the environment (which responds to whatever policy is currently in force), which is what allows the method to efficiently handle nontrivial market-clearing conditions like Huggett's (Introduction, pp. 3-4).
Solving three benchmark models within minutes
The paper's headline computational results: the Krusell-Smith (1998) model solves in about 55 seconds, the Huggett (1993) model with aggregate risk in around one minute, and a one-asset HANK model with a forward-looking New Keynesian Phillips curve in around three minutes, all on a single NVIDIA A100 GPU via a JAX implementation -- with average relative deviation of aggregate bond demand from supply of just 0.22% in the HANK application's market-clearing condition (Introduction, p. 3; section 4, Conclusion).
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.