<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Chiyuan Wang | Macro Paper Warehouse</title><link>https://macropaperwarehouse.com/authors/chiyuan-wang/</link><description>Chiyuan Wang</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><atom:link href="https://macropaperwarehouse.com/authors/chiyuan-wang/index.xml" rel="self" type="application/rss+xml"/><item><title>Structural Reinforcement Learning for Heterogeneous Agent Macroeconomics</title><link>https://macropaperwarehouse.com/papers/structural-reinforcement-learning-for-heterogeneous-agent-macroeconomics/</link><guid>https://macropaperwarehouse.com/papers/structural-reinforcement-learning-for-heterogeneous-agent-macroeconomics/</guid><description>&lt;p&gt;Standard recursive formulations of heterogeneous-agent models with aggregate risk force the entire cross-sectional distribution of agents into the Bellman equation characterizing individual decisions &amp;ndash; the &amp;ldquo;Master equation&amp;rdquo; &amp;ndash; purely because low-dimensional equilibrium prices, unlike the distribution itself, do not follow a Markov process, so rational agents forecasting prices end up needing to forecast the whole distribution. This extreme curse of dimensionality remains the central computational bottleneck for global solutions of heterogeneous-agent models, so severe that even a Huggett (1993) model with aggregate risk &amp;ndash; despite looking simple &amp;ndash; proved impossible for any team to solve in an influential benchmarking exercise and was dropped from the project altogether. This paper sidesteps the Master equation entirely using ideas from reinforcement learning (RL): agents learn equilibrium price dynamics directly from simulated paths, as standard RL would, but the paper&amp;rsquo;s &amp;ldquo;structural reinforcement learning&amp;rdquo; (SRL) approach departs from standard RL by assuming agents have structural knowledge of their own individual-state dynamics (their budget constraint and idiosyncratic income process), letting the authors compute &lt;em&gt;exact&lt;/em&gt; policy gradients by differentiating through these known dynamics rather than relying on the noisy, approximate policy gradients standard RL methods estimate; only the equilibrium price process itself is treated as unknown and learned from simulation. By further restricting agents to condition their policies only on current (or briefly lagged) prices rather than the full price history or the distribution, the paper solves for a low-dimensional &amp;ldquo;restricted perceptions equilibrium&amp;rdquo; in the sense of Sargent (1991) rather than the full rational-expectations equilibrium &amp;ndash; expectations are restricted in functional form but remain statistically consistent with actual outcomes. Because policy functions depend only on prices, they double as individual supply/demand schedules that can be integrated across the distribution and market-cleared period-by-period along a simulation, treating market clearing as part of the &amp;ldquo;environment&amp;rdquo; (in RL parlance) rather than something solved inside an optimization loop &amp;ndash; which is what lets the method efficiently handle nontrivial market-clearing conditions that have historically been very hard. Implemented in JAX on a single GPU, the resulting structural policy gradient (SPG) algorithm solves the Krusell and Smith (1998) model in about 55 seconds, the previously-unsolved Huggett (1993) model with aggregate risk in around one minute, and a one-asset HANK model with a forward-looking New Keynesian Phillips curve in around three minutes &amp;ndash; with the Krusell-Smith solution closely matching alternative global solutions of the rational-expectations equilibrium, and allowing agents a longer history of lagged prices barely moving the solution, indicating most of the information relevant for forecasting prices is already contained in current prices. The paper is explicit that its algorithm, as presented, is not itself intended as an empirically realistic theory of how real economic agents form expectations, though it suggests the &amp;ldquo;sampling&amp;rdquo;-based logic behind SRL could in principle be developed into one.&lt;/p&gt;</description></item></channel></rss>