Estimating Nonlinear Heterogeneous Agents Models with Neural Networks
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
Estimating a model of many households in its full, nonlinear form -- with a genuine floor on interest rates and real uncertainty -- has been out of reach, since fitting it to data means resolving the model at hundreds of thousands of trial values. This paper feeds the parameters into the network that solves the model, so one trained network covers every value, and trains a second, cheaper network to mimic a proper likelihood calculation. That matters: these let the authors fit a 100-household nonlinear model with 12 unknowns in under two days, recovering true values from test data -- though heterogeneity turns out barely identified by ordinary aggregate data.
What this paper finds — and why it matters
Economists routinely approximate away features of their models – nonlinear dynamics, aggregate uncertainty, agent heterogeneity – to make estimation feasible, but it is often unclear how much these simplifications distort a model’s predictions. This paper develops a neural-network-based solution and estimation method designed to avoid such approximations entirely. The key device is to treat a model’s structural parameters as additional “pseudo state variables” fed into the neural networks that approximate its policy functions, so that a single (more expensive) training run yields the model’s entire solution mapping across the whole parameter space, rather than the solution at one parameter point – exploiting the fact that neural networks scale cheaply to extra inputs. Because likelihood-based estimation of nonlinear models also requires a computationally costly Monte Carlo (particle) filter at every parameter draw, the paper trains a second “surrogate” neural network – the neural network particle filter – on a modest sample of particle-filter-evaluated likelihoods, giving a near-instant approximate mapping from parameters to likelihood that can be plugged into a standard Metropolis-Hastings sampler. After validating the approach on a linearized New Keynesian model (where the true solution is known analytically) and on a tractable representative-agent model with an aggregate zero-lower-bound nonlinearity (where results closely match a conventional particle-filter estimation), the paper applies its method to a fully nonlinear Heterogeneous Agent New Keynesian (HANK) model with 100 households, idiosyncratic labor-productivity risk, an individual borrowing limit, and a zero lower bound on the nominal interest rate – a model with hundreds of state and pseudo-state variables and 12 estimated structural parameters, including parameters that directly govern the degree of household heterogeneity. Using the calibrated model as the true data-generating process and 500 simulated periods of output growth, inflation, and the interest rate, a 1-million-draw Bayesian estimation recovers all 12 parameters, with the true value falling inside the 90% credible interval in every case, completed in under two days on a modern desktop computer – what the authors describe as the first estimation of a HANK model in its fully nonlinear specification. The exercise also reveals a strong asymmetry in identification: parameters governing aggregate dynamics (habit formation, price-adjustment costs, monetary-policy responses, shock persistences) are estimated precisely, while the posteriors for parameters governing idiosyncratic risk and the borrowing limit are “rather flat,” suggesting standard macro aggregate data contains little information about the underlying degree of heterogeneity and that richer, distributional data would likely be needed to pin these down.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What computational bottleneck does the paper’s method specifically target?
“Likelihood estimation adds an additional important layer of complication in that it requires solving the model at hundred of thousands of different points in the parameter space. As a result, only extremely fast and efficient solution methods are compatible with likelihood analysis” (Literature review, p. 4). Neural networks can approximate high-dimensional solution mappings efficiently once trained, but naively re-solving (re-training) the model at every candidate parameter value during estimation would still be prohibitively slow; the paper’s contribution is a way to avoid that repeated re-solving altogether.
Q2. How does treating parameters as “pseudo state variables” solve this problem?
“A critical step we propose in this paper is to use the model’s parameter vector as direct input in the neural network. In other words, we treat the parameters as pseudo state variables by exploiting the scalability of neural networks. We then train this extended neural network with the parameters as inputs over the entire parameter space. This step… yields the model’s solution mapping… instead of just one point in this solution mapping” (Introduction, pp. 2-3). Although training becomes harder with the extra inputs, the computational gain is large because the resulting network can be evaluated at any new parameter draw almost instantly, compared to repeatedly solving the model with conventional global methods.
Q3. What is the “neural network particle filter,” and why is it needed in addition to the policy-function network?
Nonlinear models require a Monte Carlo filter (e.g., a particle filter) to evaluate the likelihood, which is “computationally more costly than linear filters (e.g., the Kalman filter),” limiting “the amount of parameter combinations we can evaluate.” The paper’s solution is “an additional neural network that provides a direct mapping from the model parameters to the value of the likelihood function – via the particle filter… trained only on some selected data points to approximate the outcome of the particle filter in an efficient manner,” requiring “only thousands of data points” rather than the millions of draws a conventional estimation would need, so that “it takes us a few milliseconds to evaluate the likelihood of models with hundreds of state variables at one parameter value” (Introduction, pp. 3-4).
Q4. What is the estimated HANK model’s structure, and which 12 parameters are estimated?
The model has 100 households facing idiosyncratic labor-productivity risk and an individual borrowing limit B, Rotemberg price adjustment, and a Taylor-rule-governed nominal interest rate subject to a zero lower bound (section 4, pp. 28-31). The 12 estimated parameters split into those “related to the idiosyncratic risk” – the borrowing limit B and the persistence ρ_s and standard deviation σ_s of labor productivity – and those “related to aggregate risk”: habit h, Rotemberg pricing cost φ, Taylor-rule persistence ρ_r and responses to inflation and output θ_Π, θ_Y, the preference-shock persistence ρ_ζ, and the standard deviations of the preference, growth, and monetary-policy shocks (σ_ζ, σ_g, σ_mp) (Table 3, p. 32).
Q5. How is the estimation validated, and what is the headline result?
Using the calibrated model itself as the true data-generating process, the authors simulate 500 periods of quarterly output growth, inflation, and the nominal interest rate, then run a Random Walk Metropolis-Hastings algorithm with 1 million post-burn-in draws using the trained policy and likelihood-surrogate networks. “The results show that the posterior median is very close to the true value. In particular, the true value is contained in the 90% credible interval for all parameters,” and the entire estimation procedure – training both networks and running the sampler – “takes us less than 2 days with a modern day desktop computer” (section 4.4, pp. 32-33).
Q6. What are the two main advantages the authors claim over conventional estimation of HANK models?
First, “we can estimate HANK models in its fully nonlinear specification,” fully accounting for the zero lower bound and the borrowing limit, unlike approaches that “either restrict the aggregate dynamics (by relying on linearisation or perfect foresight) or the heterogeneous setup (by using RANK or TANK models).” Second, “our method does not restrict the choice of parameters that we can include in the estimation,” because the model solves simultaneously for the (stochastic) steady state and the dynamics, allowing estimation of parameters – like the borrowing limit and the moments of idiosyncratic risk – “that affect the (stochastic) [steady state]” directly (section 4.4, p. 33).
Q7. What does the paper find about how well aggregate data identifies the degree of heterogeneity in the model?
“The posterior is rather flat in this area. This indicates that standard aggregate data used for macroeconomic models does not contain too much information for estimating the parameters related to the degree of heterogeneity or inequality. Including distributional data is probably necessary to better pin down such values” (section 4.4, p. 33), referring specifically to the borrowing limit B and the persistence/volatility of idiosyncratic labor productivity, ρ_s and σ_s. By contrast, the parameters governing aggregate dynamics are estimated tightly, illustrating an asymmetry in what standard time-series data (output growth, inflation, interest rates) can and cannot reveal about a heterogeneous-agent model’s deep structure.
Q8. How does this paper’s approach differ from the two other explicitly “global” solution methods on this reading list that also use machine learning – DeepHAM and structural reinforcement learning?
Unlike DeepHAM (Han, Yang and E), which uses simulated-path optimization together with algorithmically-learned “generalized moments” of the distribution to solve a single calibrated model globally, this paper’s central innovation is aimed specifically at estimation: treating parameters themselves as network inputs so the same trained network covers the entire parameter space, paired with a fast likelihood surrogate. The paper explicitly distinguishes itself from prior neural-network solution methods, noting “none of these papers estimate nonlinear, heterogeneous agents models with likelihood methods,” and from Fernández-Villaverde et al. (2019), whose comparable exercise “estimate[s] the critical parameter of a model with a financial sector and heterogeneous households with neural networks using the time series of GDP growth” but at much smaller scale (one parameter versus this paper’s twelve) (Literature review, pp. 4-5). It also explicitly contrasts itself with Schaab (2020), noting that paper “studies a nonlinear HANK model based on a solution method unrelated to machine learning” (p. 5) – underscoring that this paper’s distinguishing contribution is coupling a global, nonlinear solution method to full likelihood-based Bayesian estimation, not the global solution step alone.
Key terms in this paper
Definitions below follow the paper's own usage.
- Parameters as pseudo state variables
- The paper's central computational device: instead of training a neural network to solve the model at one fixed parameter vector, the model's structural parameters are included as additional inputs to the network -- "pseudo state variables" -- so that a single trained network approximates the entire solution mapping from the parameter space to the model's equilibrium law of motion, rather than one point on that mapping. This exploits the "scalability" of neural networks (additional inputs are cheap to add) to make the repeated re-solving that likelihood-based estimation requires unnecessary (Introduction, pp. 2-3).
- Neural network particle filter (surrogate likelihood)
- A second neural network, distinct from the one approximating policy functions, trained on a modest number (15,000 in the paper's application) of particle-filter-evaluated likelihood values at different parameter draws, so as to provide "a direct mapping from the model parameters to the value of the likelihood function... via the particle filter." Once trained, it evaluates the likelihood at a new parameter point in milliseconds, letting a Markov Chain Monte Carlo sampler run without repeatedly invoking the (expensive) particle filter itself (Introduction, pp. 2-3; section 4.3).
- The estimated nonlinear HANK model with ZLB and borrowing limit
- A 100-household nonlinear HANK model combining idiosyncratic labor-productivity risk and an individual borrowing limit with an aggregate zero lower bound on the nominal interest rate, Rotemberg price adjustment, and shocks to preferences, government spending/growth, and monetary policy -- estimated with 12 structural parameters split between those governing idiosyncratic risk (the borrowing limit B, and the persistence/volatility of labor productivity) and those governing aggregate dynamics (habit, price adjustment cost, Taylor-rule coefficients, shock persistences and standard deviations) (section 4, pp. 28-33, Table 3).
- Recovery of the true data-generating process
- The paper's headline validation result: using the calibrated model itself as the data-generating process and 500 simulated periods of output growth, inflation, and the interest rate, a 1-million-draw neural-network-based Metropolis-Hastings estimation recovers all 12 structural parameters with posterior medians close to the true values and the true value contained in the 90% credible interval for every parameter, completed in under two days on a modern desktop computer (section 4.4, p. 32-33, Table 3).
- Flat identification of heterogeneity parameters from aggregate data
- A finding that standard macro aggregates (output growth, inflation, interest rates) carry very little information about the parameters that govern the model's cross-sectional heterogeneity -- the borrowing limit and the persistence/standard deviation of idiosyncratic labor-productivity risk -- whose posteriors "are rather flat," in contrast to the sharply identified aggregate-side parameters, suggesting "that standard aggregate data used for macroeconomic models does not contain too much information for estimating the parameters related to the degree of heterogeneity or inequality" and that distributional data would likely be needed instead (section 4.4, p. 33).