<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Methodology | Macro Paper Warehouse</title><link>https://macropaperwarehouse.com/topics/methodology/</link><atom:link href="https://macropaperwarehouse.com/topics/methodology/index.xml" rel="self" type="application/rss+xml"/><description>Methodology</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Jan 2026 00:00:00 +0000</lastBuildDate><item><title>A Learning Model of Financial Instability</title><link>https://macropaperwarehouse.com/papers/a-learning-model-of-financial-instability/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/a-learning-model-of-financial-instability/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Williams asks whether the recurrent boom-bust dynamics of Minsky&amp;rsquo;s financial instability hypothesis — &amp;ldquo;periods of stability lead to periods of instability&amp;rdquo; — can arise endogenously from a tractable rational-agent model in which investors learn about asset returns. This matters because standard rational-expectations asset-pricing models cannot generate the high, volatile price-dividend ratios, sizeable risk premia, and recurrent crashes seen in data, and because Minsky&amp;rsquo;s narrative has long lacked a clean formal mechanism. The paper&amp;rsquo;s main contribution is theoretical (a new instability/limit-cycle result for adaptive learning), with a secondary quantitative exercise.&lt;/p&gt;
&lt;p&gt;Model setup: A small-open-economy variant of the Lucas (1978) consumption-based asset-pricing model studied under learning by Adam, Marcet and Nicolini (2016). A representative agent with power utility (risk aversion gamma, discount factor beta) can borrow/lend at a fixed risk-free gross return R and holds a unit supply of stock paying an i.i.d.-growth dividend (log dividend growth = d + sigma*W, with centered binomial shocks W in {-1,1}). Adding the risk-free asset creates a portfolio problem and endogenous debt dynamics (the net asset position omega), which the closed-economy literature lacks. Agents wrongly believe log returns are i.i.d. binomial with mean m and standard deviation s, and update (m, s^2) by constant-gain recursive least squares with gain epsilon (the weight on new information). A borrowing/leverage constraint (0 &amp;lt;= v &amp;lt;= vbar on the stock portfolio share) ensures equilibrium exists. The self-confirming equilibrium (SCE) has (m,s)=(mu,sigma), v=1, omega=1, and a constant price-dividend ratio.&lt;/p&gt;
&lt;p&gt;Mechanism: The pricing function is extremely steep near v=1; the derivative at the SCE is delta&amp;rsquo;(1)=delta*(1+delta*), so with a mean P/D near 29 a 1-percentage-point fall in v (to 0.99) implies roughly a 30% drop in P/D (to ~20.3). Tranquil periods lower volatility estimates, raising v and prices; once heavily invested, the economy is fragile. Booms end via two mechanisms: binding leverage constraints (rare in the calibration, driving only one crash in the long simulation) and — the novel and dominant channel — a rapid boom raising perceived variance faster than perceived mean, causing agents to cut v and triggering a crash.&lt;/p&gt;
&lt;p&gt;Main quantitative findings (with magnitudes and scope): Theoretically, the SCE is stable only for gains below a threshold; at epsilon-bar the Jacobian of the averaged system has complex eigenvalues on the unit circle (a Neimark-Sacker / discrete Hopf bifurcation), and above it a stable limit cycle exists (Theorem 1, using Kuznetsov 1998). The threshold is approximately epsilon-bar = 8.9 x 10^-4, far below the calibrated epsilon = 0.0052 (about six times larger), so empirically plausible gains imply instability. Eigenvalues at threshold: 0.512 +/- 0.859i = e^(+/-1.0333i). Calibration uses Shiller (2024) S&amp;amp;P 500 data, 1871-2022 annual: empirical P/D mean 28.97, sd 15.53; log P/D mean 3.25, sd 0.46; 100x log return mean 6.51, sd 16.90; dividend growth 100x(d,sigma)=(1.56, 11.104). Optimizing (beta,gamma,epsilon) the baseline matches log P/D (mean 3.15 vs 3.25, sd 0.46 vs 0.46) and returns (6.44 vs 6.51; sd 16.85 vs 16.90) with beta=0.979, gamma=3.278, epsilon=0.0052, and a low risk-free rate 100xlog R=0.87. Crashes (defined as a 30% P/D drop) occur every ~38 years in the baseline vs ~25 years in data; matching the data frequency would need a larger gain near 0.025. The closed-economy and rational-expectations versions essentially cannot produce such crashes. Drawbacks: consumption growth is too volatile (sd ~16.79 vs 1.27 in data) and return predictability is far stronger than in the data.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-exactly-drives-the-instability-and-how-is-it-established-rather-than-merely-simulated"&gt;Q1. What exactly drives the instability, and how is it established rather than merely simulated?&lt;/h3&gt;
&lt;p&gt;Instability comes from the feedback between beliefs (m, s) and the net asset/debt position omega: beliefs set the portfolio share, which sets prices and returns, which feed back into beliefs. Williams formalizes this by stacking current beliefs, lagged beliefs, and the state omega into a 5-dimensional first-order system X_{t+1}=G(X_t, chi_t), then studies the deterministic averaged system Xbar_{t+1}=Gbar(Xbar_t) (averaging only over the i.i.d. dividend shocks chi, NOT over omega as the small-gain limit does). Linearizing at the SCE fixed point, Theorem 1 shows all Jacobian eigenvalues lie inside the unit circle for gains below a threshold epsilon-bar, a complex pair hits the unit circle at epsilon-bar (Neimark-Sacker bifurcation), and a unique stable closed invariant curve (limit cycle) appears for epsilon just above. He verifies the nondegeneracy and stability conditions numerically.&lt;/p&gt;
&lt;h3 id="q2-why-does-small-gain-analysis-mislead-here-and-what-is-the-methodological-contribution"&gt;Q2. Why does small-gain analysis mislead here, and what is the methodological contribution?&lt;/h3&gt;
&lt;p&gt;Standard learning convergence results take the gain to zero, treating state dynamics as &amp;lsquo;fast&amp;rsquo; relative to beliefs and averaging over the state. Williams shows this is valid only for extremely small gains in his model because the radius of stability is tiny (epsilon-bar ~ 8.9e-4). Averaging over omega destroys the very belief-state feedback that drives cycles. His contribution to the learning literature is applying discrete-time bifurcation theory (Kuznetsov 1998) to show a Neimark-Sacker bifurcation and stable limit cycle in an economic learning model — which he states is novel — relating it to prior cautions by Cho (2018), Chien-Cho-Ravikumar (2020), and instability examples in Evans-Honkapohja (2009) and Honkapohja-McClung (2023).&lt;/p&gt;
&lt;h3 id="q3-what-are-the-two-crash-mechanisms-and-which-dominates"&gt;Q3. What are the two crash mechanisms and which dominates?&lt;/h3&gt;
&lt;p&gt;(1) Binding leverage constraint: if v hits vbar during a boom, inflows stop, generating a negative return surprise that lowers the mean estimate and cuts v. This is rare in the calibration — it drives only the final crash in the long simulation. (2) Endogenous volatility: a rapid boom raises both the estimated mean and variance of returns; when the variance effect dominates, agents cut the risky share even without hitting the constraint. Because the economy is in the steeply sloped pricing region, a tiny cut produces a large crash. This is the dominant, novel mechanism and causes all other crashes, including those in the highlighted closeup. In one example the portfolio share peaks just above one (period 441), and a move from v=1.004 to 1.000 produces about a 48% P/D drop; the cascade bottoms near v=0.47 and P/D around 2, a decline of over 95% from peak.&lt;/p&gt;
&lt;h3 id="q4-what-does-the-representative-boom-bust-cycle-look-like-quantitatively"&gt;Q4. What does the representative boom-bust cycle look like quantitatively?&lt;/h3&gt;
&lt;p&gt;In a &amp;gt;1,000-period simulation, P/D rises 30-50% within a span of years then crashes by a similar or larger amount. In the detailed cycle the P/D rises from 30 to 50 over a few periods before crashing to around 2. After a crash, volatility estimates start high and decline monotonically over roughly 50 periods; agents slowly raise v, prices rise (amplified by the omega multiplier as accumulated bonds are sold), until a rapid boom enters the fragile region and crashes again. Severe crashes of similar magnitude recur at periods 327, 442, 801, and 1067.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-role-of-stochastic-shocks-versus-endogenous-dynamics"&gt;Q5. What is the role of stochastic shocks versus endogenous dynamics?&lt;/h3&gt;
&lt;p&gt;Conditional impulse responses (at periods 432, 438, 440 into a boom) show shocks matter most early: at t=432 a positive shock reinforces the boom while a negative shock dampens fluctuations with little belief change. By t=438 positive/negative impulses are qualitatively similar but differ in magnitude. By t=440 the endogenous dynamics dominate and shock differences are minimal — the boom continues only a couple periods before a severe crash. Shocks govern timing and magnitude, but endogenous belief changes ultimately drive the cycles.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-open-economy-assumption-matter-and-what-is-the-closed-economy-comparison"&gt;Q6. How does the open-economy assumption matter, and what is the closed-economy comparison?&lt;/h3&gt;
&lt;p&gt;The baseline is a small open economy: international trade in bonds (fixed R) but only domestic equity trade, which permits nonzero net debt and asset flows. This debt/portfolio-adjustment channel is essential. In the closed economy (R adjusts each period to clear bonds at zero net supply, v=1), with baseline parameters the fit is much worse: P/D too high (3.70), returns lower (4.08), and far less volatile (sd P/D 0.15). Re-optimizing the closed model improves means but misses volatilities (overshoots return sd at 17.74, undershoots P/D sd at 0.36) and requires very different parameters (beta=0.903, gamma=4.736, epsilon=0.0272); crashes occur only every ~469 years (extremely rare). Intermediate cases with partial interest-rate adjustment keep the closed-economy qualitative features. The empirical justification: foreign investors held 33% of US Treasuries, 27% of corporate debt, but only 17% of US equities in 2023 (vs 46% Treasuries and 9% equities in 2006).&lt;/p&gt;
&lt;h3 id="q7-how-does-the-speed-of-learning-gain-trade-off-against-fit"&gt;Q7. How does the speed of learning (gain) trade off against fit?&lt;/h3&gt;
&lt;p&gt;As the gain falls toward zero, the P/D ratio converges to its SCE value log(P/D)~3.6 and its distribution concentrates there (lower volatility); higher gains raise volatility and crash frequency but lower the mean P/D because more time is spent recovering from crashes (booms are short-lived, crashes slow to recover — an asymmetry). The calibration balances mean and volatility of P/D at epsilon=0.0052, but matching the observed crash frequency would need a larger gain near 0.025. The model can match price level/volatility OR crash frequency but struggles to match the speed of market dynamics simultaneously.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-main-empirical-drawbacks"&gt;Q8. What are the main empirical drawbacks?&lt;/h3&gt;
&lt;p&gt;(1) Consumption growth is far too volatile (model sd ~16.79 vs data 1.27), inherited from using volatile empirical dividend growth as the driving process; treating stocks as levered equity claims (Abel 1999) could break the consumption-dividend link. (2) Return predictability — both autocorrelation and long-term reversal — is much stronger than in the data, where it is weak at best; additional shocks or heterogeneity would dampen it. (3) The subjective excess return is essentially uncorrelated with the P/D ratio, whereas survey expected returns are positively correlated with P/D (Greenwood-Shleifer 2014; Adam-Marcet-Beutel 2017; Barberis et al. 2018); allowing different gains for the mean and variance moves the model closer to survey evidence.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-differ-from-closely-related-prior-work"&gt;Q9. How does this differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Versus Branch and Evans (2011), who also have agents learning about risk and return: their booms/crashes are rare &amp;rsquo;escape&amp;rsquo; events from equilibrium, whereas in Williams&amp;rsquo;s model they are typical outcomes driven by a fundamental instability (a stable limit cycle), not rare escapes. Versus Adam, Marcet and Nicolini (2016): Williams adds a fixed-rate risk-free asset, creating a portfolio problem and debt dynamics (omega) that are crucial for the boom-bust cycles. Versus behavioral/extrapolation and diagnostic-expectations models (Barberis et al. 2018; Bordalo-Gennaioli-Shleifer 2018; Bianchi-Ilut-Saijo 2024), Williams uses standard adaptive learning, and crucially crashes collapse valuations far below fundamentals (not mere reversion to fundamentals), with stability breeding instability as in Minsky.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;A full policy analysis is outside the paper&amp;rsquo;s scope, but Williams notes a higher interest rate lowers excess stock returns and makes boom-bust cycles less frequent — yet potentially more severe (when a boom does occur, larger price/return spikes). This implies policymakers face tradeoffs more complex than simply &amp;rsquo;leaning against the wind&amp;rsquo; of bubbles. The scope conditions: the model has exogenous output growth, a representative agent, a constant risk-free rate, and a constant rational-expectations P/D, so all fluctuations are attributed to learning; relaxing these (e.g., for finance-real interactions) is left for future work.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>A Tractable Income Process for Business Cycle Analysis</title><link>https://macropaperwarehouse.com/papers/a-tractable-income-process-for-business-cycle-analysis/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/a-tractable-income-process-for-business-cycle-analysis/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Guvenen, McKay, and Ryan estimate a stochastic income process for US male workers that simultaneously matches five empirical regularities from Social Security Administration administrative panel data covering 1978–2011: (i) flat and acyclical variance of income growth rates, (ii) volatile and procyclical Kelley skewness, (iii) very high kurtosis — targeted at 20 for one-year changes and 12 for five-year changes — (iv) a near-linear rise in cross-sectional log-income variance from age 25 to 55, and (v) a systematic factor structure in business cycle incidence whereby income losses during recessions are predictably related to a worker&amp;rsquo;s pre-recession income rank. All five facts are drawn from Guvenen et al. (2014) and Guvenen et al. (2021), which document them from SSA records on individual income histories.\n\nThe income process adds three key departures to the workhorse persistent-plus-transitory Gaussian specification. First, transitory &amp;ldquo;nonemployment&amp;rdquo; shocks — arriving annually with approximately 45% probability and drawn from an exponential distribution — create fat tails through their arrival (large income losses) and departure (large income gains), and leave a persistent &amp;ldquo;scarring&amp;rdquo; residue through a passthrough parameter ψ estimated at 9.4% in the baseline nonemployment model. Each year, roughly 8.6% of workers experience income declines of 50% or more from the nonemployment shock alone, and 1.8% fall to effectively zero income. The scarring mechanism makes the left tail of the income growth density fatter than the right tail, consistent with the data (left-tail log-density slope 1.4, right-tail slope –2.2). Second, innovations to the persistent AR(1) component are drawn from a time-varying three-component normal mixture — with the dominant central component realized with about 83% probability and near-zero standard deviation (~1%), flanked by left-tail and right-tail components with probabilities of ~10.9% and ~6.2% and standard deviations of ~16.4% and ~19.2% — whose means shift with contemporaneous aggregate wage income growth (xt = β·Δwt). This mean-shifting mechanism generates procyclical skewness under an acyclical variance, because it redistributes probability mass between the tails without altering mixture probabilities or component variances. Third, a piecewise-linear factor structure makes each individual&amp;rsquo;s income sensitivity to aggregate fluctuations depend on the persistent component of income (γi + zi,t), with a kink separating two slope regimes. In the Great Recession, workers at the 10th percentile of pre-recession income lost approximately 18 percentage points more than workers at the 90th percentile; both the bottom and top deciles were more exposed than the middle of the distribution, producing a V-shaped incidence pattern.\n\nEstimation uses simulated method of moments (SMM) with 360,000 simulated individuals per year, a 1947 burn-in start, and optimization via the TikTak global algorithm. Six models of increasing complexity are estimated, each requiring only one individual state variable (the persistent component z) — matching the parsimony of the standard model. The workhorse Gaussian model (Model 1) understates the variance of one-year log income changes by 60–80%; introducing nonemployment shocks (Model 2) largely resolves this, matching one-year variance exactly and narrowing the five-year shortfall to 30%. Adding the time-varying normal mixture (Model 3) generates procyclical skewness and acyclical variance. Adding the factor structure (Model 4) captures differential recession exposure. Models 5 and 6 introduce Heterogeneous Income Profiles (HIP, σκ = 0.015) and estimate AR(1) persistence freely, obtaining ρ ≈ 0.80, which better captures the right tail of the income growth distribution.\n\nThe paper recommends Model 5 as a general-purpose benchmark (without the factor structure), Model 4 when differential business cycle incidence is central, and Model 3 when maximum parsimony is needed. The richer income dynamics documented here have direct implications for quantifying the welfare cost of business cycles, the value of social insurance, the design of automatic stabilizers, the distribution of marginal propensities to consume, and asset pricing under heterogeneous agents.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-estimation-procedure-and-what-data-does-it-use"&gt;Q1. What is the estimation procedure and what data does it use?&lt;/h3&gt;
&lt;p&gt;The paper uses simulated method of moments (SMM), targeting approximately 120+ moments derived from Social Security Administration administrative panel data on individual income histories of US male workers over 1978–2011 (from Guvenen et al. 2014 and 2021). The simulation panel contains 360,000 individuals per year, initialized in 1947 with a burn-in period. Optimization uses the TikTak global algorithm (Arnoud et al., 2019). Moments targeted include the 10th, 50th, and 90th percentiles of one-, three-, and five-year income growth averaged across 1979–2011 (nine moments); kurtosis at one-year and five-year horizons (two moments); cross-sectional variance of log income at ages 25, 35, 45, and 55 (four moments); left- and right-tail mass and log-density slopes from the 1995–1996 income growth distribution (four moments); the full time series of Kelley skewness for one-, three-, and five-year changes (93 moments); and piecewise-linear slopes of the factor structure for seven business cycle episodes — four recessions and three expansions covering 1979–2010 (14 moments). Moments are weighted approximately equally, with skewness moments down-weighted collectively.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-key-departures-from-the-workhorse-gaussian-model-and-what-feature-does-each-address"&gt;Q2. What are the three key departures from the workhorse Gaussian model and what feature does each address?&lt;/h3&gt;
&lt;p&gt;First, transitory &amp;rsquo;nonemployment&amp;rsquo; shocks drawn from an exponential distribution, arriving with ~45% annual probability, along with a scarring parameter ψ that loads a fraction of the transitory shock onto the persistent state — this generates the high kurtosis, thick tails, and asymmetry (steeper right than left tail) of the income growth distribution. Second, a three-component time-varying normal mixture for persistent innovations — the component means shift with the aggregate wage component xt = β·Δwt — producing procyclical skewness and acyclical variance simultaneously. Third, a piecewise-linear factor structure f(γi + zi,t) mediating each individual&amp;rsquo;s exposure to aggregate fluctuations, capturing the V-shaped relationship between pre-recession income rank and recession income loss.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-scarring-mechanism-and-how-large-is-it-empirically"&gt;Q3. What is the scarring mechanism and how large is it empirically?&lt;/h3&gt;
&lt;p&gt;Transitory nonemployment shocks ζi,t are assigned with probability (1 − pζ) each year and drawn from an exponential distribution with parameter λ, where ℓi,t ∈ [0,1] represents the income fraction lost. A fraction ψ of this transitory shock flows permanently into the persistent state zi,t via ˜ηi,t = ηi,t + ψζi,t. In Model 2, the annual probability of receiving a nonemployment shock is 45% (pζ ≈ 0.55), λ = 3.357 (mean income loss fraction ≈ 0.30), and ψ = 9.4%. Each year, 8.6% of workers experience income declines of 50% or more from the nonemployment shock alone, and 1.8% effectively lose all income (full-year nonemployment). The scarring makes the right tail steeper than the left tail in the income growth distribution, as re-employed workers do not return to their pre-shock income level.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-time-varying-normal-mixture-generate-procyclical-skewness-without-changing-variance"&gt;Q4. How does the time-varying normal mixture generate procyclical skewness without changing variance?&lt;/h3&gt;
&lt;p&gt;The three normal mixture components for the persistent innovation η are: a central component (probability ~83%, standard deviation ~1%), a left-tail component (~10.9%, ~16.4% sd), and a right-tail component (~6.2%, ~19.2% sd). Their means shift via the latent variable xt = β·Δwt: the central and left-tail means move with xt while the right-tail mean does not. A normalization ensures xt has zero mean-income effect. In recessions (xt &amp;lt; 0, Δwt &amp;lt; 0), the left-tail component&amp;rsquo;s mean shifts down and the right-tail component&amp;rsquo;s mean shifts up relative to the central, generating more left-skewed draws without changing the probabilities or variances of the components — hence acyclical variance and procyclical skewness. Alternative designs (cyclical mixture probabilities or variances) did not generate both patterns simultaneously.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-factor-structure-and-how-non-monotonic-is-it"&gt;Q5. What is the factor structure and how non-monotonic is it?&lt;/h3&gt;
&lt;p&gt;In deep recessions the factor structure is broadly monotone decreasing over the bulk of the distribution (lower-income workers lose more), with the 10th percentile losing about 18 percentage points more than the 90th percentile in the Great Recession (2007–2010). However, the pattern reverses for the top 10% of the income distribution: high earners also face large losses in financial-market-driven recessions, producing a V-shape. The piecewise-linear model f(q) with a kink at q-bar and slopes α1 (below) and α2 (above) captures this. The model fits the Great Recession V-shape and the mild 1990–1992 and 2000–2002 recessions (where the pattern is flatter, consistent with smaller drops in wt), but struggles to fit the large top-income losses in 2000–2002 without an additional stock-market-correlated factor.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-levels-vs-differences-puzzle-and-how-is-it-resolved"&gt;Q6. What is the levels-vs-differences puzzle and how is it resolved?&lt;/h3&gt;
&lt;p&gt;The canonical persistent-plus-transitory Gaussian model (Model 1) faces a fundamental tension: it can fit the cross-sectional variance of log income levels at each age, but it then understates the variance of one-year and five-year log income changes by 60–80% (squared standard deviations from Figures 8a and 9a). This tension was documented by Heathcote, Perri, and Violante (2010). Introducing the nonemployment shocks in Model 2 largely resolves it: the one-year variance of log income changes is matched exactly, and the five-year understatement narrows to about 30%. The nonemployment shock contributes high-frequency variance in income changes without requiring a comparably large increase in the variance of the persistent state, because it is mostly transitory.&lt;/p&gt;
&lt;h3 id="q7-what-role-does-hip-play-and-what-tensions-does-it-create"&gt;Q7. What role does HIP play and what tensions does it create?&lt;/h3&gt;
&lt;p&gt;Heterogeneous Income Profiles (HIP, σκ = 0.015 from Baker 1997 and Guvenen et al. 2021) allow AR(1) persistence ρ to be estimated freely rather than restricted to 1. The estimated ρ falls to 0.80 in Models 5 and 6. HIP provides a convex component to the lifecycle variance profile (from dispersion in individual growth-rate slopes κi) that offsets the concave contribution of mean-reverting persistent shocks, maintaining a near-linear age-variance profile at ρ &amp;lt; 1. Lower persistence better fits the right tail of annual income growth and the standard deviation of five-year changes. However, in Model 6 HIP worsens the fit to the factor structure, because mean reversion at ρ &amp;lt; 1 already generates faster income growth for low-income workers in expansions, reducing the work the factor structure needs to do in booms while resisting the factor structure&amp;rsquo;s ability to generate large losses for low-income workers in recessions.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-and-alternative-specifications-are-estimated"&gt;Q8. What robustness checks and alternative specifications are estimated?&lt;/h3&gt;
&lt;p&gt;The paper estimates two supplementary models reported in Appendix B. Model 2&amp;rsquo; removes the scarring component (ψ ≡ 0) from Model 2, finding a worse fit particularly in the histogram, kurtosis, and lifecycle inequality moments. Model 3&amp;rsquo; replaces the time-varying mixture with a static normal mixture (β ≡ 0), still improving over Model 2 (objective falls from 2.44 to 2.26) via better tail fit and average skewness, but without capturing the procyclical skewness time series. Model 4&amp;rsquo; removes time variation from the innovation distribution (β ≡ 0) while retaining the factor structure, showing that the factor structure fit survives without time variation in skewness. Additionally, the paper discusses a special parsimony case: under ρ = 1, homothetic preferences, and no factor structure, z can be normalized away entirely, leaving no individual state variable.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-prior-work-on-non-gaussian-income-processes"&gt;Q9. How does this paper relate to and differ from prior work on non-Gaussian income processes?&lt;/h3&gt;
&lt;p&gt;Kaplan, Moll, and Violante (2018) capture leptokurtic income growth but include no business cycle variation and no factor structure. McKay (2017), McKay and Reis (2021), and Catherine (2021) allow for procyclical skewness in income risk but do not target high kurtosis or a factor structure. Bhandari, Evans, Golosov, and Sargent (2021) allow for a factor structure but do not match higher-moment properties of income risk. Other work documenting the relevant facts includes Guvenen, Ozkan, and Song (2014) for countercyclical skewness in US SSA data; Guvenen, Karahan, Ozkan, and Song (2021) for lifecycle earnings dynamics from the same source; Harmenberg (2021) and Kramarz, Nimier-David, and Delemotte (2021) for related European evidence; and Guvenen, Schulhofer-Wohl, Song, and Yogo (2017) for factor structure evidence labeled &amp;lsquo;worker betas.&amp;rsquo; This paper is the first to jointly target and fit all four properties within a single tractable process that adds only one state variable.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-and-structural-implications-highlighted-by-the-paper"&gt;Q10. What are the policy and structural implications highlighted by the paper?&lt;/h3&gt;
&lt;p&gt;Leptokurtic income risk (high kurtosis, fat tails) has quantitatively important effects on the value of social insurance and optimal redistribution (Saez, 2001; Golosov, Troshkin, and Tsyvinski, 2016) and interacts with borrowing constraints to shape the distribution of wealth and marginal propensities to consume (Kaplan, Moll, and Violante, 2018). Cyclical variation in income risk — the procyclical skewness feature — matters for the welfare cost of business cycles (Storesletten, Telmer, and Yaron, 2001; Krebs, 2003, 2007) and for the optimal design and welfare value of automatic stabilizers (McKay and Reis, 2021; Bhandari et al., 2021). The factor structure is relevant for cyclical variation in income inequality and for asset pricing under household heterogeneity (Mankiw, 1986; Constantinides and Duffie, 1996; Constantinides and Ghosh, 2016). The scope condition throughout is male US workers in the SSA administrative data; no direct results are provided for female workers, self-employed individuals, or other countries, though the modeling framework is general.&lt;/p&gt;
&lt;h3 id="q11-what-practical-guidance-does-the-paper-provide-for-incorporating-the-process-into-dynamic-models"&gt;Q11. What practical guidance does the paper provide for incorporating the process into dynamic models?&lt;/h3&gt;
&lt;p&gt;The paper provides explicit Bellman equation structure: cash on hand m and the persistent income state z are the two endogenous individual state variables (z being the single income-process state variable), with individual parameters γ and κ treated as fixed effects. Income at each node requires evaluating a closed-form expression from Equation 1. Expectations over next-period z and ζ are handled via quadrature, with the time-varying mixture of normals requiring quadrature nodes that shift with the aggregate state S and S′ — following McKay and Reis (2021). Under the special case ρ = 1, homothetic preferences, and no factor structure, all variables can be normalized by exp(z + γ), eliminating z as a state variable and reducing the problem to one with no idiosyncratic income state. The authors note that a perpetual-youth demographic structure avoids tracking age as a state variable.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Procyclical skewness&lt;/strong&gt;: In the paper&amp;rsquo;s sense: the Kelley skewness of the cross-sectional distribution of one-year and five-year income growth rates falls significantly during every NBER recession (distribution shifts left — more large negative shocks, fewer large positive ones) and rises during expansions, while the standard deviation of that distribution shows no discernible cyclical pattern. This is a feature of the income shock distribution itself, not of average income levels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nonemployment shock with scarring&lt;/strong&gt;: A transitory income loss event modeled as an exponential random variable ℓi,t ∈ [0,1] (representing the fraction of income lost) arriving with probability ~45% per year. A fraction ψ of this transitory shock is loaded permanently onto the persistent income state — the &amp;lsquo;scarring&amp;rsquo; effect — so that re-employed workers do not fully return to their pre-shock income trajectory. In the paper&amp;rsquo;s model this single mechanism generates high kurtosis, thick double-Pareto tails, and asymmetric tail slopes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time-varying normal mixture for persistent innovations&lt;/strong&gt;: A three-component mixture of normals for the AR(1) innovation η in which the component means (not probabilities or variances) shift proportionally to contemporaneous aggregate wage income growth via a loading parameter β. A mean-preserving normalization ensures no effect on average income. This mean-shifting mechanism moves probability mass between the central and tail components of the innovation distribution, generating procyclical skewness while keeping income growth variance acyclical.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Factor structure in business cycle incidence&lt;/strong&gt;: A systematic, pre-determined relationship between a worker&amp;rsquo;s position in the persistent income distribution and the magnitude of income change experienced during a given recession or expansion. Modeled as a piecewise-linear function f(γi + zi,t) that multiplies the aggregate income component wt, with slopes that differ below and above an estimated kink point. Empirically, the factor structure produces a V-shaped incidence pattern: income losses in deep recessions are largest at both the bottom and top of the pre-recession income distribution, and smallest in the middle.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Income scarring parameter (ψ)&lt;/strong&gt;: The fraction of a transitory nonemployment shock ζi,t that is permanently loaded onto the persistent income state zi,t via the equation ˜ηi,t = ηi,t + ψζi,t. Estimated at 9.4% in Model 2 and 15.1% in Model 3. Controls the degree to which transitory shocks generate long-lasting income effects and determines the relative steepness of the left versus right tails of the annual income growth distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneous Income Profiles (HIP)&lt;/strong&gt;: Individual-specific linear deterministic growth-rate slopes κi distributed with standard deviation σκ = 0.015 (calibrated from Baker 1997 and Guvenen et al. 2021), representing permanent heterogeneity in the steepness of individual income trajectories over the lifecycle. Introducing HIP allows the AR(1) persistence parameter ρ to be estimated below 1 (≈0.80 in Models 5–6) while preserving the near-linear age-variance profile, because the convex variance contribution of heterogeneous slopes offsets the concavity induced by mean-reverting persistent shocks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Kelley skewness&lt;/strong&gt;: In the paper&amp;rsquo;s use: a robust, percentile-based measure of skewness defined as [(P90 − P50) − (P50 − P10)] / (P90 − P10), which the paper prefers for income growth distributions because it is less sensitive to extreme outliers than moment-based skewness. Used as the primary target for capturing business cycle variation in the shape of the income growth distribution.&lt;/p&gt;</description></item><item><title>An Analytical Model of Behavior and Policy in an Epidemic</title><link>https://macropaperwarehouse.com/papers/an-analytical-model-of-behavior-and-policy-in-an-epidemic/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/an-analytical-model-of-behavior-and-policy-in-an-epidemic/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper builds a tractable, fully analytical version of the workhorse macro-epidemiology (&amp;ldquo;econ-epi&amp;rdquo;) model and uses it to characterize how susceptible individuals behave during a deadly epidemic, how a social planner would have them behave, and the externality that separates the two. The motivation is that prior macro-SIR results came almost entirely from numerical simulation; a closed-form treatment can expose general insights those simulations missed and provide a transparent benchmark for any future epidemic. The model appends the standard Kermack-McKendrick SIR system (susceptible S, infected I, recovered R, deceased D, with transmission rate β, recovery rate γr, death rate γd, and γ := γr + γd) with forward-looking agents who choose an activity level λ ∈ [0,1] that scales transmission via β = βa·λ + βo. The single key modeling departure is LINEAR (rather than convex) costs of mitigation, microfounded by indivisible activity choices in the spirit of Rogerson (1988); this makes the optimal control bang-bang or singular and yields closed-form solutions. Three constants organize the analysis: the herd immunity threshold S̄ := γ/β, the basic reproduction number R0 := 1/S̄, and the infection fatality rate IFR := γd/γ. A central composite statistic is the cost-benefit ratio of mitigation κ := (uW − uL)/(βa·IFR·VSL), where VSL := uW/ρ is the value of statistical life in utility terms.\n\nMain results. (1) Decentralized equilibrium (Proposition 1): there is no mitigation at the very start and the very end of the epidemic; mitigation occurs only over an interval [t0, t1). Susceptibles begin mitigating just below full susceptibility, the infection rate peaks exactly at t0 (when precautions are greatest), and from then on the effective reproduction number sits slightly below one, producing a gently declining infection path — a pattern the author notes is broadly consistent with first-wave Covid-19 data. The equilibrium infection trajectory is approximated by the simple ray I(t) ≈ (S(t)/S̄)·κ, and the equilibrium steady-state susceptibility is S∞ ≈ S̄ − S̄·√(2κR0). A higher κ and lower S̄ both reduce mitigation and raise infections (a &amp;ldquo;fatalism effect&amp;rdquo;). (2) Socially optimal behavior (Propositions 2-3): optimal policy is bang-bang (λ* ∈ {0,1}) — no mitigation at start and end, full mitigation in a single intermediate interval. The planner &amp;ldquo;holds fire,&amp;rdquo; lets infections climb high, then imposes maximal restrictions late, driving the system quickly to herd immunity. The optimal long-run susceptibility is S∞* ≈ S̄ − S̄·2κR0/(κR0 − 1)². (3) The externality: contrary to the conventional view, susceptibles&amp;rsquo; privately optimal behavior is EXCESSIVELY cautious — the equilibrium infection rate lies below the optimal infection rate for any S above herd immunity — yet cumulative deaths are HIGHER in equilibrium than under the planner. Mitigation by susceptibles mostly substitutes infection risk intertemporally (&amp;ldquo;flattening the curve also makes it fatter&amp;rdquo;); beyond eliminating epidemic overshoot it cannot prevent the inevitable share 1 − S̄ from being infected. The planner&amp;rsquo;s late-strong-short lockdown comes close to implementing a lottery that randomly selects who gets sick.\n\nImplications. Because the externality runs in the opposite direction to standard intuition, optimal policy can call for the government to INCREASE interaction (the paper cites the UK&amp;rsquo;s 2020 &amp;ldquo;Eat Out To Help Out&amp;rdquo; subsidy as an analogue). Results are framed as technical/foundational insights, not direct prescriptions: the benchmark abstracts from reinfection, variants, vaccines/cures, healthcare capacity limits, and endogenous IFR, all of which can shift specific recommendations while leaving the underlying forces intact.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-or-solution-strategy-and-what-makes-the-analytical-characterization-possible"&gt;Q1. What is the &amp;lsquo;identification&amp;rsquo; or solution strategy, and what makes the analytical characterization possible?&lt;/h3&gt;
&lt;p&gt;This is a theory paper, so the relevant strategy is solving the dynamic optimization analytically rather than empirically. The enabling assumption is LINEAR costs of mitigation (instantaneous utility u = λ·uW + (1−λ)·uL), microfounded by indivisible activity choices as in Rogerson (1988), where λ is the probability of being active in a mixed-strategy equilibrium. Linearity makes the current-value Hamiltonian linear in the control λ, so the optimal control is bang-bang or singular with switching function ψ(t) := uW − uL − (ηs(t) − ηi)·βa·I(t). This permits closed-form characterization of switching points and trajectories. The main &amp;rsquo;threat&amp;rsquo; the author addresses is generality: does linearity drive the conclusions? Section VI shows numerically that convex costs (U = uL + λ^(1−α)·(uW − uL), with α the convexity degree) merely smooth out the kinks and corners without changing qualitative features — passing what the author calls the &amp;lsquo;Solow test.&amp;rsquo;&lt;/p&gt;
&lt;h3 id="q2-what-is-the-core-economic-mechanism-behind-excessive-caution-and-the-two-ways-the-paper-frames-the-externality"&gt;Q2. What is the core economic mechanism behind &amp;rsquo;excessive caution,&amp;rsquo; and the two ways the paper frames the externality?&lt;/h3&gt;
&lt;p&gt;In equilibrium, the singular-control optimality condition equates a constant marginal cost of mitigation (uW − uL) to a marginal benefit (ηs(t) − ηi)·βa·I(t). The shadow value of being susceptible ηs(t) rises over time (cumulative future infection risk and cumulative future mitigation effort both decline as the epidemic progresses), while ηi is constant. To keep the equation balanced, βa·I(t) must fall, so agents become more cautious over time. First framing of the externality: the planner recognizes that at least 1 − S̄ of the population must eventually be infected (and a share IFR of those die); individuals recognize this too (perfect foresight) but each wants to avoid being in the infected group, so they over-mitigate, merely delaying rather than preventing infections. Second framing: stronger mitigation today lowers near-term infections but raises later infections — &amp;lsquo;flattening the curve also makes it fatter&amp;rsquo; — so beyond removing overshoot, mitigation only substitutes infection risk intertemporally. The planner internalizes the whole time path; individuals take the aggregate infection rate as given.&lt;/p&gt;
&lt;h3 id="q3-why-is-the-optimal-lockdown-late-strong-and-short-rather-than-gradual"&gt;Q3. Why is the optimal lockdown &amp;rsquo;late, strong, and short&amp;rsquo; rather than gradual?&lt;/h3&gt;
&lt;p&gt;From the planner&amp;rsquo;s law of motion, the velocity Ṡ/S is proportional to I. An interior λ would lower instantaneous costs proportionately but increase the duration of mitigation more than proportionately (since both λ and I are lower), so gradualism is dominated. This makes optimal policy bang-bang with a single interval of maximal restriction. The planner therefore holds fire, lets I climb high (where the system moves fast), then imposes λ=0 to drive the trajectory quickly to herd immunity — minimizing cumulative deaths at minimum cost rather than flattening the curve.&lt;/p&gt;
&lt;h3 id="q4-how-do-equilibrium-and-optimal-cumulative-deaths-compare-and-why-does-the-more-cautious-equilibrium-produce-more-deaths"&gt;Q4. How do equilibrium and optimal cumulative deaths compare, and why does the more cautious equilibrium produce MORE deaths?&lt;/h3&gt;
&lt;p&gt;Cumulative deaths equal IFR·(1 − S∞). The equilibrium steady-state susceptibility S∞ ≈ S̄ − S̄·√(2κR0) lies below the planner&amp;rsquo;s S∞* ≈ S̄ − S̄·2κR0/(κR0 − 1)², meaning the equilibrium overshoots herd immunity by more, so 1 − S∞ (cumulative infections) and hence deaths are higher in equilibrium. The equilibrium&amp;rsquo;s caution lowers the infection rate at each S above herd immunity and stretches the epidemic out (raising economic cost), but does not prevent the inevitable infections and in fact allows more overshoot than the planner&amp;rsquo;s quick-to-herd-immunity strategy. Cumulative death toll is increasing in R0 and in κ.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-role-of-the-cost-benefit-ratio-κ-and-the-fatalism-effect"&gt;Q5. What is the role of the cost-benefit ratio κ and the &amp;lsquo;fatalism effect&amp;rsquo;?&lt;/h3&gt;
&lt;p&gt;κ := (uW − uL)/(βa·IFR·VSL) combines preferences, epidemiology, and policy effectiveness: the numerator is the utility cost of mitigation; the denominator is the benefit (lower activity reduces transmission by βa, preventing deaths by IFR, each life worth VSL = uW/ρ). A higher κ lowers mitigation and raises the equilibrium infection rate, starts mitigation later (lower S(t0)), and raises cumulative deaths. The &amp;lsquo;fatalism effect&amp;rsquo; has two parts: a lower S̄ (greater lifetime chance of falling ill) dissuades mitigation today; and the high expected cumulative future mitigation effort at the epidemic&amp;rsquo;s start lowers the value of staying alive, further tempering precaution. The simple approximation I(t) ≈ (S(t)/S̄)·κ captures the first part but omits the second.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-practical-back-of-the-envelope-contribution"&gt;Q6. What is the practical &amp;lsquo;back-of-the-envelope&amp;rsquo; contribution?&lt;/h3&gt;
&lt;p&gt;The paper provides a recipe to trace the equilibrium epidemic path without solving the full dynamic model: (1) compute the thresholds S(t0) ≈ 1 − κ/(√(2κR0)·(1−S̄))·S̄(1−S̄), S(t1) ≈ S̄ − ρ/(βo + βa), and S∞ ≈ S̄ − S̄·√(2κR0); (2) plot the ray I = (S/S̄)·κ between the thresholds; (3) splice it on both sides with the no-mitigation (λ=1) trajectory I = −S + S̄·log S + C0. This rivals running the naive SIR model in simplicity but is grounded in optimizing behavior, giving a more plausible benchmark for human populations. The author intends it for forecasting any future epidemic.&lt;/p&gt;
&lt;h3 id="q7-how-do-the-results-relate-to-and-differ-from-prior-numerical-econ-epi-work"&gt;Q7. How do the results relate to and differ from prior numerical econ-epi work?&lt;/h3&gt;
&lt;p&gt;The equilibrium characterization is qualitatively consistent with Farboodi et al. (2021) — little mitigation at the start, then a jump keeping the effective reproduction number just below 1 — the only difference being their path is smoother due to convex costs. Eichenbaum-Rebelo-Trabandt (2021) get a qualitatively different, still hump-shaped equilibrium infection path because in their calibration mitigation is too weak to push the effective reproduction number below 1 (so βo is not &amp;lsquo;sufficiently low&amp;rsquo;). For the planner, the paper&amp;rsquo;s late-strong-short lockdown differs from work finding early/strong responses (Farboodi et al.) or intermediate restrictions (Alvarez et al. 2021; Eichenbaum et al. 2021), for two reasons: (1) this model rules out suppression/vaccine arrival as a feasible endgame, whereas papers allowing vaccine arrival find early strong suppression optimal; (2) the planner here controls only susceptibles&amp;rsquo; behavior with linear costs, whereas broader instruments and convex costs make intermediate restrictions more attractive. The paper is, to the author&amp;rsquo;s knowledge, the first to derive equilibrium and optimal behavior fully analytically and to show the susceptibles&amp;rsquo; externality makes the infection rate too LOW socially.&lt;/p&gt;
&lt;h3 id="q8-what-do-the-costate-shadow-value-dynamics-reveal"&gt;Q8. What do the costate (shadow-value) dynamics reveal?&lt;/h3&gt;
&lt;p&gt;The private value of infection ηi = (uI + (γr/ρ)·uW)/(ρ+γ) is time-invariant (payoffs while ill/recovered/dead don&amp;rsquo;t depend on timing). The social value of an infected person η&lt;em&gt;i is time-varying because the planner internalizes onward transmission via a (η&lt;/em&gt;i − η&lt;em&gt;s)(βaλ&lt;/em&gt; + βo)S* term. η&lt;em&gt;i is deeply negative at the epidemic&amp;rsquo;s start (diverging as I→0, because an infinitesimal seed inflicts unboundedly large relative damage), rises sharply and roughly tracks the private value during the bulk of the epidemic (e.g. when S ∈ [0.5, 0.9]), and settles just above zero in the long run. In the long run the social value of an additional infected person can even be negative when γd is high, because the value of that person&amp;rsquo;s life is below the welfare loss from infections they spread. The social value of a susceptible η&lt;/em&gt;s is always below the private value (except converging to uW/ρ in the long run), reflecting unpriced future contagion.&lt;/p&gt;
&lt;h3 id="q9-what-robustnessextension-checks-does-the-paper-run"&gt;Q9. What robustness/extension checks does the paper run?&lt;/h3&gt;
&lt;p&gt;Section VI: (1) Convex costs (numerical, α=0.3) smooth kinks but preserve qualitative features. (2) Broader planner instruments — controlling susceptibles AND infected (without distinguishing them), or restricting everyone identically — are &amp;lsquo;double-edged&amp;rsquo;: more costly (especially late when many are recovered) but more effective because they also restrict the infected; effectiveness gains peak at intermediate restrictions (around λ=1/2) due to the quadratic contact function, which makes intermediate restrictions and earlier/longer lockdowns more attractive, moving results toward Alvarez et al. (2021). Section VII discusses healthcare/ICU capacity constraints (optimal to hold infections at the capacity level until near herd immunity; endogenous IFR brings equilibrium and optimal paths closer but doesn&amp;rsquo;t change the externality&amp;rsquo;s nature), feasible suppression (optimal policy becomes a discrete choice between herd-immunity and best suppression strategy; equilibrium behavior is largely insensitive to suppression feasibility), and temporary immunity/endemicity (strengthens the fatalism effect, raising equilibrium infections; optimal policy still rushes to steady state, now also to avoid costly multiple waves).&lt;/p&gt;
&lt;h3 id="q10-what-is-the-calibration-used-for-the-figures-and-is-it-meant-to-be-quantitatively-serious"&gt;Q10. What is the calibration used for the figures, and is it meant to be quantitatively serious?&lt;/h3&gt;
&lt;p&gt;The calibration resembles Covid-19 but is explicitly illustrative, not a serious quantitative calibration. A model period is a week. Epidemiological parameters: βo = 0.7, βa = 1.24, γr = 0.77, γd = 0.0078, implying R0 = 2.5, S̄ = 0.4, IFR = 1%, and average disease duration of 9 days; under full mitigation (λ=0) R0 falls to 0.9. Annual discount rate is 4% (weekly ρ = 0.96^(−1/52) − 1). Utility is logarithmic; weekly consumption is $60,000/52 ≈ $1,250 so uW = log(1250) ≈ 7; full lockdown cuts consumption 20%, giving uL = 6.6, (uW − uL)/uL = 3.2%. With VSL = $10 million, κ = 0.002 (0.2%).&lt;/p&gt;
&lt;h3 id="q11-what-are-the-key-caveats-and-the-scope-of-the-policy-implications"&gt;Q11. What are the key caveats and the scope of the policy implications?&lt;/h3&gt;
&lt;p&gt;The author stresses the model is a stripped-down BENCHMARK: no reinfection, no variants, constant IFR, no cure or vaccine (so herd immunity pins down minimum feasible deaths). Specific results are &amp;rsquo;technical contributions, not direct normative prescriptions.&amp;rsquo; The striking implication that a planner might subsidize interaction (forcing susceptibles to interact, since optimal activity sometimes exceeds equilibrium activity) faces an implementability problem — restricting activity is easier than increasing it. The herd-immunity-quick strategy ceases to be optimal once suppression is feasible (vaccine/cure expected), ICU constraints bind with endogenous IFR, or immunity is only temporary; but the underlying forces (the susceptibles&amp;rsquo; intertemporal infection-substitution externality) continue to operate in all these richer settings.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Herd immunity threshold (S̄)&lt;/strong&gt;: S̄ := γ/β, the level of susceptibility below which the infected pool shrinks; in this model, because there is no cure or vaccine, it pins down the minimum feasible deaths and is the endgame both equilibrium and planner converge toward.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cost-benefit ratio of mitigation (κ)&lt;/strong&gt;: κ := (uW − uL)/(βa·IFR·VSL), a composite statistic combining preferences, epidemiology, and policy effectiveness; the numerator is the utility cost of mitigation and the denominator the benefit (transmission reduction βa times deaths averted IFR times value of statistical life). Higher κ means less mitigation and more infections.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Excessive caution / susceptibles&amp;rsquo; externality&lt;/strong&gt;: The paper&amp;rsquo;s central finding that privately optimal mitigation by susceptibles is too cautious socially — the equilibrium infection rate lies below the optimal rate for any S above herd immunity — because each individual wants to avoid being in the inevitable infected share, merely substituting infection risk intertemporally rather than preventing it; the conventional one-way infected-spreader externality view is therefore incomplete.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Linear costs of mitigation / singular control&lt;/strong&gt;: The assumption (microfounded by indivisible activity choices à la Rogerson 1988) that utility is linear in activity λ, making the Hamiltonian linear in the control so the optimum is bang-bang or singular; this delivers sharp closed-form solutions whose intuitions survive under convex costs (the &amp;lsquo;Solow test&amp;rsquo;).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Late-strong-short lockdown&lt;/strong&gt;: The socially optimal policy in this benchmark: hold fire while infections climb high, then impose maximal restrictions (λ=0) in a single intermediate interval that quickly drives the system to herd immunity — minimizing cumulative deaths at minimum cost rather than flattening the curve.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Costates (ηs, ηi)&lt;/strong&gt;: Shadow values of being in the susceptible and infected states. ηi (private) is constant since the payoffs of being ill are timing-independent; the planner&amp;rsquo;s η*i is time-varying because it internalizes onward transmission and can even be negative in the long run when the death rate is high.&lt;/p&gt;</description></item><item><title>An irrelevance theorem for risk aversion and time-varying risk</title><link>https://macropaperwarehouse.com/papers/an-irrelevance-theorem-for-risk-aversion-and-time-varying-risk/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/an-irrelevance-theorem-for-risk-aversion-and-time-varying-risk/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Chen and Palomino prove a general irrelevance theorem identifying when risk aversion and time-varying risk are irrelevant for key model dynamics in representative-agent macroeconomic models. The central research question is why advances in risk modeling — Epstein-Zin (EZ) recursive preferences, long-run risk, disaster risk — generate rich asset price behavior in endowment economies but fail to produce commensurate effects in standard production economies. The paper resolves this puzzle by characterizing the precise structural conditions under which risk parameters become irrelevant, and provides a taxonomy for how models can escape those conditions.&lt;/p&gt;
&lt;p&gt;The theoretical framework is a representative-agent model with EZ preferences, which separate the elasticity of intertemporal substitution (EIS, parameter psi) from risk aversion (gamma). The remaining economic structure — production technology, resource constraints, government policy, financial sector — is assumed to exhibit an analogous separation: variables that control expected values (&amp;ldquo;first moment states,&amp;rdquo; such as capital and productivity) are separated from variables that control higher central moments (&amp;ldquo;higher moment states,&amp;rdquo; such as stochastic volatility of productivity). The paper proceeds through three settings of increasing generality: a two-period illustrative model, a dynamic stochastic growth model with capital adjustment costs (Jermann 1998) and heteroskedastic AR(1) productivity, and a fully abstract general model covering a broad class of rational-expectations equilibrium systems.&lt;/p&gt;
&lt;p&gt;The central result is Theorem 1: if (1) intertemporal and risk preferences are separated (EZ-style), (2) first and higher moment drivers of the remaining model structure are separated, and (3) constraints are approximately linear, then risk aversion gamma and higher-moment parameters theta_h are irrelevant for the elasticity of any endogenous variable — including all asset prices — with respect to first moment states and lagged endogenous variables. Formally, in the solution z_t = z + Z_z&lt;em&gt;z_{t-1} + Z_x&lt;/em&gt;x_t + Z_h*h_t, the elasticity matrices Z_z and Z_x are independent of gamma and theta_h. Risk parameters affect only model intercepts and steady states (the constant z) and the elasticity with respect to higher moment states (Z_h). Thus augmenting a stochastic growth model with shocks to volatility or risk aversion has no effect on impulse responses to productivity shocks or other first-moment disturbances.&lt;/p&gt;
&lt;p&gt;In the homoskedastic special case (constant volatility), risk aversion is irrelevant for the impulse response of every variable, including all asset prices. This clarifies the Tallarini (2000) separation: it is not a separation between macroeconomic and financial variables, but between means (average equity premium, steady-state levels) and volatilities and impulse responses. Risk aversion affects the level of the equity premium but not stock price volatility or impulse responses.&lt;/p&gt;
&lt;p&gt;Numerical verification using projection methods (Caldara et al. 2012) confirms irrelevance holds even at risk aversion of 100 and unconditional volatility of volatility of 80% of baseline. A second, richer model class — with EIS of 0.3, capital adjustment cost elasticity of 3, and left-skewed gamma-distributed productivity shocks calibrated to match Bekaert and Engstrom (2017) quarterly consumption growth moments (kurtosis 4.04, skewness -0.399, matching model kurtosis of 4 and skewness of -0.82) — produces an equity premium more than three times larger than the baseline class and a stock price elasticity with respect to productivity about three times larger, yet continues to display irrelevance: risk aversion and time-varying risk have essentially no effect on the stock price elasticity with respect to productivity.&lt;/p&gt;
&lt;p&gt;The theorem extends to smooth ambiguity preferences (Klibanoff, Marinacci, Mukerji 2005) and multiplier preferences (Hansen and Sargent 2001) as long as risk adjustments remain functions of higher-moment state variables. The paper also derives the Barro-King (1984) comovement restriction under recursive preferences (Appendix C), showing that in the neoclassical structure only productivity shocks generate positive comovement of consumption, investment, and labor. This interacts with the irrelevance theorem to explain why production-economy asset pricing models face a compounded difficulty: volatility and risk-aversion shocks cannot break irrelevance within the standard structure, and they also cannot generate the required comovement without additional mechanisms.&lt;/p&gt;
&lt;p&gt;The paper provides a unified taxonomy for generating a meaningful role for risk in production economies. One can &amp;ldquo;break&amp;rdquo; irrelevance by removing one of the three assumptions: (1) allowing risk aversion to vary with economic conditions as in Campbell-Cochrane (1999) habit formation or heterogeneous agents; (2) introducing non-separability between first and higher moments in production, as in Di Tella and Hall (2022) where entrepreneurial idiosyncratic risk makes aggregate volatility endogenous; or (3) incorporating sufficient nonlinearity via occasionally binding constraints, as in Brunnermeier-Sannikov (2014) or Gourio-Ngo (2020) near the zero lower bound. Alternatively, one can &amp;ldquo;adapt&amp;rdquo; to irrelevance by driving dynamics with higher-moment shocks — volatility shocks (Basu-Bundick 2017, combined with nominal rigidities to preserve comovement) or risk-aversion shocks (Basu et al. 2024, combined with an investment reallocation channel).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-intuition-behind-the-irrelevance-theorem"&gt;Q1. What is the core intuition behind the irrelevance theorem?&lt;/h3&gt;
&lt;p&gt;The Euler equation under EZ preferences decomposes into an Intertemporal Term (characterizing expected consumption-return tradeoffs, driven by EIS) and a Risk Term (characterizing tradeoffs across unexpected future states, driven by risk aversion). In standard models, the production technology is &amp;lsquo;a perfect foresight model with shocks tacked on&amp;rsquo;: transformation across time is separated from transformation across future states. Because constraints are approximately linear, innovations to endogenous variables with respect to first-moment shocks (productivity, capital) do not contain investment or other endogenous variables, so the Risk Term is a function only of higher-moment states. Differentiating the Euler equation with respect to a first-moment state therefore eliminates the Risk Term entirely, leaving only the Intertemporal Term and making the solution for that elasticity independent of gamma and sigma.&lt;/p&gt;
&lt;h3 id="q2-how-is-the-tallarini-2000-result-clarified-and-extended"&gt;Q2. How is the Tallarini (2000) result clarified and extended?&lt;/h3&gt;
&lt;p&gt;Tallarini (2000) shows that risk aversion is irrelevant for quantity dynamics in a homoskedastic real business cycle model. This is widely interpreted as a separation between macroeconomic (quantity) and financial (price) variables. The paper shows this interpretation is incorrect. When shocks are homoskedastic, risk aversion is irrelevant not just for quantities but for all asset price dynamics, including stock price volatility. The actual separation is between means (steady states, intercepts, average equity premium — all of which depend on risk aversion) and volatilities and impulse responses (which do not). The paper extends Tallarini&amp;rsquo;s result by showing irrelevance holds for all endogenous variables including stock prices, by showing it persists under heteroskedasticity for elasticities with respect to first-moment states specifically, and by generalizing to abstract models beyond the neoclassical RBC framework.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-three-conditions-required-for-irrelevance-and-what-is-the-role-of-each"&gt;Q3. What are the three conditions required for irrelevance and what is the role of each?&lt;/h3&gt;
&lt;p&gt;The three conditions are: (1) Separation of intertemporal and risk preferences — EZ-style preferences ensure risk aversion gamma enters only the Risk Term of the Euler equation, not the Intertemporal Term. If preferences are non-separable (e.g., power utility, habit formation), gamma enters the intertemporal tradeoff and affects first-moment elasticities. (2) Separation of first and higher moment drivers in the remaining model structure — production technology and all other constraints must not link transformation of goods across time to transformation across states. If higher-moment variables appear in the production function or resource constraint (e.g., idiosyncratic risk in entrepreneurial production as in Di Tella-Hall 2022), first-moment states appear in the Risk Term and irrelevance breaks. (3) Approximate linearity of constraints — nonlinearities create interactions between current state values and forward-looking volatility. Strong enough nonlinearities (such as those introduced by occasionally binding constraints near the zero lower bound or in financial crisis models) can cause irrelevance to fail even when conditions (1) and (2) hold.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-formal-mathematical-structure-of-the-general-model-and-theorem"&gt;Q4. What is the formal mathematical structure of the general model and theorem?&lt;/h3&gt;
&lt;p&gt;The general model consists of a system of expectational equilibrium conditions E[f(z_{t+1}, x_{t+1} | z_t, x_t, h_t, z_{t-1}; Theta)] = 0, where z_t are endogenous variables, x_t are first-moment exogenous states following a heteroskedastic AR(1) with shock distribution conditional on h_t, and h_t are higher-moment states with an independent AR(1) process. The equilibrium conditions split into constraints (f0, depending only on theta_0, not gamma or theta_h) and asset-pricing Euler equations (depending on the EZ SDF, hence on gamma). The proof uses a risk-adjusted affine approximation (Assumptions 1 and 2): constraints are approximated as conditionally affine in states; the CGF of shocks is conditionally affine in h_t. Conjecturing a linear solution z_t = z + Z_z&lt;em&gt;z_{t-1} + Z_x&lt;/em&gt;x_t + Z_h*h_t and applying the method of undetermined coefficients in separate layers shows that Z_z satisfies a quadratic matrix equation depending only on theta_0 (Proposition 2, Equation 171), and Z_x satisfies a Sylvester equation also depending only on theta_0 and Z_z (Equation 172). Since neither equation involves gamma or theta_h, those parameters are irrelevant for Z_z and Z_x. Z_h and z do depend on all parameters including gamma and theta_h.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-irrelevance-theorem-interact-with-the-barro-king-1984-comovement-constraint"&gt;Q5. How does the irrelevance theorem interact with the Barro-King (1984) comovement constraint?&lt;/h3&gt;
&lt;p&gt;Barro and King (1984) show that, in the neoclassical structure, shocks other than productivity shocks fail to generate the observed positive comovement of consumption, investment, and labor. The paper derives this result under recursive preferences in Appendix C, confirming it extends to the EZ case. The comovement constraint implies that, within the neoclassical structure, the magnitude of higher-moment shocks must be limited to preserve comovement — production-economy asset pricing models typically drive business cycles with productivity shocks rather than volatility or risk-aversion shocks. But the irrelevance theorem implies that productivity shock impulse responses are independent of risk. Together, these results explain why modeling asset prices in production economies is non-trivial: one must simultaneously address comovement (ruling out large higher-moment shocks as the primary business cycle driver) and irrelevance (meaning productivity shocks cannot be enriched with risk dynamics). A successful model must either break irrelevance or adapt to it with mechanisms that also solve the comovement problem.&lt;/p&gt;
&lt;h3 id="q6-what-does-it-mean-to-break-irrelevance-and-what-are-the-main-examples"&gt;Q6. What does it mean to &amp;lsquo;break&amp;rsquo; irrelevance and what are the main examples?&lt;/h3&gt;
&lt;p&gt;Breaking irrelevance means removing one of the three conditions so that risk aversion or risk parameters enter the elasticity with respect to first-moment states. Examples: (1) Campbell-Cochrane (1999) external habit: risk aversion varies over time as consumption approaches habit, creating time-varying links between the intertemporal and risk terms of the Euler equation. Heterogeneous households (Guvenen 2009) produce similar effects. (2) Di Tella and Hall (2022): entrepreneurs face uninsurable idiosyncratic shocks, making the aggregate production function incorporate risk. Volatility is endogenous and affects how the economy responds to first-moment shocks. Colacito et al. (2014), Decker et al. (2016), and Belo (2010) similarly incorporate production risk-return tradeoffs. (3) Brunnermeier-Sannikov (2014) financial frictions and Gourio-Ngo (2020) zero lower bound: occasionally binding constraints introduce strong enough nonlinearities to break the affine approximation and generate large endogenous volatility far from the steady state. A non-separable production example is also given: if k_{t+1} = (k+i)*1{epsilon &amp;gt;= 0}, investment appears in the consumption innovation and hence in the Risk Term, causing gamma and sigma to enter the first-moment elasticity.&lt;/p&gt;
&lt;h3 id="q7-what-does-it-mean-to-adapt-to-irrelevance-and-what-are-the-main-examples"&gt;Q7. What does it mean to &amp;lsquo;adapt&amp;rsquo; to irrelevance and what are the main examples?&lt;/h3&gt;
&lt;p&gt;Adapting to irrelevance means staying within the class of models covered by the theorem but driving business cycle dynamics with shocks to higher-moment states rather than first-moment states. In this approach, risk aversion and risk parameters remain irrelevant for how the model responds to first-moment shocks (productivity, capital), but they do affect the elasticity with respect to higher-moment shocks and thus drive important dynamics. Basu and Bundick (2017) drive cycles with shocks to the volatility of time preference and maintain positive comovement of consumption, investment, and labor by incorporating nominal rigidities (New-Keynesian frictions break the Barro-King constraint). Basu et al. (2024) drive cycles with shocks to risk aversion and recover comovement via a novel investment reallocation channel between labor and capital. Dupor and Mehkari (2014) document other mechanisms that can overcome the comovement problem, including consumption-investment complementarities and externalities in leisure preferences.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-extend-irrelevance-beyond-epstein-zin-preferences"&gt;Q8. How does the paper extend irrelevance beyond Epstein-Zin preferences?&lt;/h3&gt;
&lt;p&gt;The paper shows irrelevance holds for a broader family of preferences as long as the log SDF can be written as a base component m*&lt;em&gt;{t+1} plus additional risk adjustments m&lt;/em&gt;{i,t+1} = f_tilde_i(Lambda, theta_0) * A_i * z_{t+1}, where Lambda is a generalized risk parameter vector (encompassing ambiguity aversion and other attitudes), and the associated certainty equivalent condition E_{i,t}[A_i&lt;em&gt;z_{t+1}] = -H_{i,t}[f_hat_i * A_i&lt;/em&gt;z_{t+1}] holds. This formulation covers smooth ambiguity preferences (Klibanoff et al. 2005, illustrated via Ju-Miao 2012 generalized smooth ambiguity with ambiguity aversion parameter eta) and multiplier preferences (Hansen-Sargent 2001). The key property for irrelevance to hold is that the risk adjustments are solely functions of higher-moment state variables h_t. For smooth ambiguity, irrelevance holds if belief dynamics are exogenous, as in Ilut-Schneider (2014).&lt;/p&gt;
&lt;h3 id="q9-what-numerical-exercises-are-conducted-to-validate-the-approximate-linearity-assumption"&gt;Q9. What numerical exercises are conducted to validate the approximate linearity assumption?&lt;/h3&gt;
&lt;p&gt;Two classes of models are solved using projection methods (Caldara et al. 2012), which provide the highest accuracy among available solution methods and capture time variation in risk premiums that second-order perturbation methods cannot. Class 1 replicates Tallarini (2000): EIS = 1, elasticity of investment = 10, normally distributed shocks (gamma shape parameter = 600), calibrated to HP-filtered output volatility of about 1.5% per quarter. Class 2 introduces larger frictions: EIS = 0.3, elasticity of investment = 3, left-skewed gamma shocks with shape parameter 6 (implying kurtosis = 4, skewness = -0.82, consistent with Bekaert-Engstrom 2017 empirical moments of quarterly consumption growth: kurtosis 4.04, skewness -0.399). For both classes, risk aversion is varied up to 100 and the unconditional volatility of volatility up to 80% of the baseline volatility. In both classes, the stock price elasticity with respect to productivity shows essentially no variation with risk aversion or volatility-of-volatility (though a slight negligible median decline is noted), while the equity premium and the stock price elasticity with respect to volatility respond clearly to those risk parameters. The exercise also shows Class 2 produces an equity premium more than three times larger than Class 1 and a stock price elasticity with respect to productivity about three times larger, yet irrelevance persists.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-paper-relate-to-and-differ-from-backus-ferriere-and-zin-2015"&gt;Q10. How does the paper relate to and differ from Backus, Ferriere, and Zin (2015)?&lt;/h3&gt;
&lt;p&gt;Backus, Ferriere, and Zin (2015) is the closest predecessor, providing irrelevance results for several specific models of time-varying risk and time-varying ambiguity. However, the paper argues they share the common misinterpretation of the Tallarini property as a separation between quantities and prices. The present paper extends their results into a fully abstract, general model structure with arbitrary equilibrium conditions and arbitrary shock distributions, proving irrelevance without tying it to specific model structures. This generality allows the paper to clarify that the separation is between means and volatilities, not between macro and finance variables. The paper also provides a clearer account of how models generate meaningful risk dynamics by breaking or adapting to the three theorem conditions.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-relationship-between-the-papers-results-and-risk-adjusted-affine-approximations-in-the-prior-literature"&gt;Q11. What is the relationship between the paper&amp;rsquo;s results and risk-adjusted affine approximations in the prior literature?&lt;/h3&gt;
&lt;p&gt;The proof builds directly on the risk-adjusted affine approximation methodology of Jermann (1998), Malkhozov (2014), and Lopez, Lopez-Salido, and Vazquez-Grande (2018). These approximations preserve exact equality for the nonlinear expectation and certainty equivalent equations (not linearizing them) while linearizing other constraints. Special cases of the irrelevance result appear in the second- and third-order perturbation solutions of Schmitt-Grohe and Uribe (2004) and Van Binsbergen et al. (2012), which this paper unifies and generalizes. The use of entropy (the conditional cumulant generating function operator) to summarize higher-order terms is motivated by Backus et al. (2014), who show entropy effectively summarizes asset pricing properties of pricing kernels. The conditionally affine CGF assumption (Assumption 2) generalizes the normal-shock setting where CGFs are exactly affine in h_t.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-scope-conditions-and-limitations-of-the-theorem"&gt;Q12. What are the scope conditions and limitations of the theorem?&lt;/h3&gt;
&lt;p&gt;The theorem applies under three maintained assumptions: (1) separation of preferences (EZ-style or the broader class in Section 4.4), (2) separation of first and higher moment drivers in all model constraints including government, financial sector, labor markets, and endowment processes, and (3) approximate linearity — formally, that the affine approximation (Assumptions 1 and 2) is accurate. The theorem does NOT apply when: constraints are strongly nonlinear due to occasionally binding constraints (ZLB, financial crisis regimes); production incorporates endogenous risk-return tradeoffs; risk aversion varies endogenously with the state (habit formation, wealth distribution with heterogeneous agents); or belief dynamics are endogenous in the ambiguity case. The paper cannot provide a complete characterization of when nonlinearities are &amp;lsquo;strong enough&amp;rsquo; to break irrelevance — numerical evidence suggests simply increasing risk aversion or vol-of-vol is insufficient, but occasionally binding constraints in the literature have been shown to be sufficient. The theorem also assumes the first and higher moment state shocks are independent (Equation 54), a modeling assumption that drives the separation.&lt;/p&gt;
&lt;h3 id="q13-what-do-the-results-imply-for-how-the-field-should-model-asset-prices-in-production-economies"&gt;Q13. What do the results imply for how the field should model asset prices in production economies?&lt;/h3&gt;
&lt;p&gt;The theorem implies that meaningful risk modeling in production economies is fundamentally more demanding than in endowment economies. In endowment economies, adding EZ preferences with high risk aversion or stochastic volatility directly affects how asset prices respond to the endowment process. In production economies, these same additions have no effect on impulse responses to productivity shocks — the primary drivers of business cycles in the neoclassical structure — because productivity is a first-moment state. Successful production-economy asset pricing models must therefore either: incorporate mechanisms that connect intertemporal and risk tradeoffs in production (endogenous volatility, incomplete markets, idiosyncratic risk); introduce sufficient structural nonlinearity; or drive business cycles with higher-moment shocks combined with additional mechanisms to preserve comovement. The paper suggests that the limited success of long-run risk and disaster risk models in production economies is not a failure of calibration but a logical consequence of the theorem&amp;rsquo;s conditions being satisfied.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;First moment states&lt;/strong&gt;: Exogenous state variables that affect expected values of the model structure (e.g., productivity level, capital stock) but not the higher central moments of the shock distributions. In the general model, x_t with shock distribution having zero mean conditional on h_t but variance and higher moments controlled entirely by h_t, not x_t itself.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Higher moment states&lt;/strong&gt;: Exogenous state variables that control the conditional higher central moments (variance, skewness, kurtosis) of the shock distributions but not their means — e.g., stochastic volatility of productivity h_t. Risk aversion and parameters governing higher moments (theta_h) are irrelevant for elasticities with respect to first-moment states but are critical for elasticities with respect to higher-moment states.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Irrelevance (in this paper&amp;rsquo;s sense)&lt;/strong&gt;: The property that risk aversion gamma and higher-moment parameters theta_h do not enter the matrices Z_z and Z_x in the solution z_t = z + Z_z&lt;em&gt;z_{t-1} + Z_x&lt;/em&gt;x_t + Z_h*h_t. These parameters are irrelevant for impulse responses and dynamic elasticities with respect to first-moment states, though they do affect steady states (z), model intercepts, and elasticities with respect to higher-moment states (Z_h).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Breaking irrelevance&lt;/strong&gt;: Removing one of the three theorem conditions — separability of preferences, separability of first and higher moment drivers in constraints, or approximate linearity — so that risk aversion or risk parameters enter the first-moment elasticities. Requires economically substantive modifications such as endogenous risk-return tradeoffs in production, habit formation, or occasionally binding constraints.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Adapting to irrelevance&lt;/strong&gt;: Staying within the class of models covered by the theorem — accepting that risk parameters do not affect first-moment impulse responses — but driving business cycle dynamics primarily with shocks to higher-moment states (volatility, risk aversion). Requires additional mechanisms (nominal rigidities, reallocation channels) to maintain positive comovement of consumption, investment, and labor, which higher-moment shocks cannot generate in the neoclassical structure alone.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Risk-adjusted affine approximation&lt;/strong&gt;: A solution method that preserves the nonlinear expectation and certainty equivalent equations exactly (not linearizing them, thereby retaining all risk effects) while log-linearizing the remaining constraints. The resulting solution is affine in the state variables, with the CGF of shocks assumed to be conditionally affine in the higher-moment states h_t. This approach captures higher-order risk terms while maintaining analytical tractability.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Entropy operator&lt;/strong&gt;: The conditional matrix operator H_t[u] = log E_t[exp(u - E_t[u])], equivalent to the vectorized conditional cumulant generating function (CGF) evaluated at 1. Used to represent all higher-order terms in the equilibrium conditions compactly; the key technical tool enabling the proof to separate expectational terms (independent of risk parameters) from entropy terms (functions of higher-moment states).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Means-volatilities separation&lt;/strong&gt;: The corrected characterization of Tallarini (2000)&amp;rsquo;s result: risk aversion affects model means (intercepts, steady states, average equity premium) but not volatilities or impulse responses of any variable — including asset prices — when shocks are homoskedastic. This reinterpretation replaces the widely held but incorrect view that Tallarini establishes a separation between macroeconomic and financial variables.&lt;/p&gt;</description></item><item><title>Balancing Work and Care: How Workplace Factors Can Mitigate the Gendered Impacts of Caregiving</title><link>https://macropaperwarehouse.com/papers/balancing-work-and-care-how-workplace-factors-can-mitigate-the-gendered-impacts-of-caregiving/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/balancing-work-and-care-how-workplace-factors-can-mitigate-the-gendered-impacts-of-caregiving/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper examines how workplace environments shape the economic consequences that fall on mothers — but not fathers — when a child is diagnosed with cancer. The motivation is a gap in the caregiving-and-labor-markets literature: while the earnings penalties from childbirth are well-documented, less is known about caregiving shocks that arrive later in childhood, or about whether and how the firm, occupation, or industry a parent works in moderates those penalties.&lt;/p&gt;
&lt;p&gt;The empirical setting is Australia. The authors use the ABS Person Level Integrated Data Asset (PLIDA), a longitudinal administrative database linking tax records (ATO, 2005–2022), Medicare health records, and 2011 Census occupation and hours data. A distinctive feature is matched employer-employee identifiers, enabling construction of workplace characteristics at the firm, occupation, and industry levels. The sample comprises 3,258 families in which a child (age 4–18, average age 12.98) began chemotherapy between 2012 and 2023 and both parents were employed two years before treatment. Pre-diagnosis average earnings are $37,639 for mothers and $79,702 for fathers (CPI-adjusted to 2012).&lt;/p&gt;
&lt;p&gt;The identification strategy is a dynamic difference-in-differences (DiD) model following Fadlon and Nielsen (2019, 2021). The treatment group consists of parents whose children started chemotherapy between 2012 and 2017; the control group consists of parents whose children will receive the same diagnosis later, between 2018 and 2023, with placebo treatment assigned six years before actual treatment. Individual fixed effects absorb time-invariant heterogeneity; year fixed effects absorb common trends. Childhood cancer — specifically chemotherapy-requiring cancer — is treated as a largely random shock with no pre-trend in earnings or employment between treated and control families before diagnosis.&lt;/p&gt;
&lt;p&gt;Main findings on the average effects: Maternal earnings fall by $5,608 in the year chemotherapy begins (14.9% of baseline earnings). The earnings decline persists for at least three years even as measured caregiving intensity (child healthcare service use) returns to baseline by year 3, leaving earnings approximately 9.7% below baseline in year 3 (−$3,645). The primary mechanism is a reduction in hours worked rather than outright job exit: employment falls by 4.9 percentage points in year 0, peaking at a decline of 5.6 percentage points two years post-treatment, a modest reduction relative to the earnings loss. Job-to-job transitions are not significantly elevated. Mental health service use (therapy, antidepressants, anxiolytics) shows no significant change for either parent, ruling out a mental health channel and reinforcing that caregiver time demands drive the result. Fathers experience no statistically significant change in earnings, employment, or job transitions across all specifications.&lt;/p&gt;
&lt;p&gt;Subgroup heterogeneity: The earnings penalty is substantially larger for mothers of younger children (under 12): −$9,443 in year 0, equivalent to 25.8% of that subgroup&amp;rsquo;s baseline earnings. For children with above-median healthcare utilization, the year-0 penalty is −$7,826 (21.6%).&lt;/p&gt;
&lt;p&gt;Workplace moderation — three dimensions are examined at the firm, occupation, and industry levels:&lt;/p&gt;
&lt;p&gt;(1) Gender pay gap: Mothers in occupations with below-average gender pay gaps face lower earnings losses ($5,782 vs $8,409; 16.5% vs 18.1%). The effect is significant at the occupation level but not at the firm or industry level.&lt;/p&gt;
&lt;p&gt;(2) Work hour intensity: Mothers in firms with below-median weekly hours face a year-0 earnings loss of $3,240 (9.9%) versus $7,159 (15.6%) in high-hours firms — a difference of $3,919, significant at the firm level. A parallel gap holds at the occupation level. When both firm and occupation are low-hours, the combined loss equals $2,519; when both are high-hours, it reaches $9,357 — a fourfold difference.&lt;/p&gt;
&lt;p&gt;(3) Female representation in the top 20% of earners: Mothers at firms where women are the majority of top-20%-earners suffer a penalty of $3,856 (8.3%) versus $7,799 (23.4%) elsewhere — a $3,943 mitigation at the firm level. At the occupation level the corresponding figures are $4,240 (9.2%) versus $8,356 (25.0%). Female representation in middle or bottom earnings tiers carries no significant moderating effect.&lt;/p&gt;
&lt;p&gt;In the combined specification (all firm- and occupation-level variables simultaneously), female representation in the top 20% and work hour intensity remain jointly significant; the gender pay gap loses significance, consistent with these variables being correlated. In the polar comparison between fully supportive jobs (low hours, high female senior representation, low occupation gender pay gap) and fully unsupportive jobs (opposite), the difference is dramatic: mothers in supportive jobs suffer a −$6,280 year-0 earnings hit that recovers fully by year 1, while mothers in unsupportive jobs face −$10,416 in year 0 widening to −$13,882 in year 3 before partially recovering in year 4.&lt;/p&gt;
&lt;p&gt;Policy implications (with scope conditions): The results support policies that reduce greedy-work norms and increase female representation in senior roles as instruments for attenuating the gendered economic cost of caregiving shocks. The study does not isolate specific workplace policies (e.g., formal paid leave) but identifies observable correlates of supportive environments. Effects are identified among working parents of children requiring chemotherapy; they do not generalize to cancer not requiring chemotherapy or other types of caregiving shocks without further evidence. Notably, fathers&amp;rsquo; outcomes are unresponsive to workplace factors, suggesting that social norms or intra-household bargaining — not workplace barriers per se — are the primary constraints on paternal caregiving adjustment.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The authors use a later-treated dynamic DiD, comparing parents whose children began chemotherapy 2012–2017 (treated) to parents whose children will begin the same treatment 2018–2023 (control), with the control group&amp;rsquo;s placebo treatment assigned six years before their actual treatment. Individual fixed effects absorb time-invariant heterogeneity; year fixed effects absorb macro shocks. The parallel trends assumption is validated by showing: (1) no statistically significant differences in pre-cancer demographic, socioeconomic, or workplace characteristics between treated and control groups (Figure 1); and (2) no pre-trend in earnings or employment in years -4 and -3 relative to baseline (Table A3, estimates small and insignificant). The main threats acknowledged are (a) non-random selection into workplace types — mothers who anticipate greater caregiving loads may sort into more family-friendly jobs — and (b) differences in baseline wage levels across job types. On (a), the authors argue the direction of selection bias goes the wrong way: if selection were driving results, mothers in supportive workplaces (who selected there due to caregiving preferences) would have weaker labor market attachment and larger post-shock earnings declines; instead the opposite is found. On (b), the authors show that absolute dollar declines in less-supportive workplaces also correspond to larger percentage declines relative to baseline, so the pattern is not an artifact of higher baseline wages in high-hour jobs (though Appendix Table A2 confirms mothers in high-hour and high-senior-female firms do have higher baseline earnings of around $46,000–$50,000 vs $32,000–$33,000).&lt;/p&gt;
&lt;h3 id="q2-how-is-the-caregiving-shock-defined-and-what-does-this-imply-for-external-validity"&gt;Q2. How is the caregiving shock defined and what does this imply for external validity?&lt;/h3&gt;
&lt;p&gt;The shock is defined as initiation of chemotherapy by the child, identified from Medicare prescription records using ATC codes beginning with L01 (excluding methotrexate L01BA01) and adding immunomodulators with chemotherapy-like effects. Chemotherapy initiation is treated as a reliable, time-consistent marker because it typically follows immediately from diagnosis of cancers such as acute lymphoid leukemia, astrocytoma, and neuroblastoma. The authors note explicitly that estimates do not represent the effects of childhood cancer not requiring chemotherapy (e.g., early-stage cancers treated with surgery, radiation, or immunotherapy alone). This restriction to chemotherapy-requiring cancers likely selects a sample with above-average caregiving intensity.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-main-mechanism-through-which-the-earnings-decline-operates"&gt;Q3. What is the main mechanism through which the earnings decline operates?&lt;/h3&gt;
&lt;p&gt;The primary mechanism is a reduction in hours worked rather than outright job exit. The employment decline (approximately 4.5–5.0 percentage points in years 0–2 per Table A3) is modest relative to the earnings loss of $5,608. A back-of-envelope calculation in footnote 6 shows that if 5% of mothers left the labor market at average earnings, the implied earnings drop would be only $1,882, far below the observed $5,608. Job-to-job transitions (probability of switching employer) are not significantly elevated. Mental health service use (psychological therapy, antidepressant/anxiolytic/antipsychotic prescriptions) shows no significant change for either parent (Appendix Figure A4), ruling out mental health deterioration as a channel. The persistence of earnings losses beyond the period of peak healthcare service use (which returns to baseline by year 3, per Appendix Figure A2) is consistent with stalled career trajectories — foregone promotions or skill development — or with continued but less-measured caregiving demands.&lt;/p&gt;
&lt;h3 id="q4-at-which-organizational-level-firm-occupation-or-industry-do-workplace-moderators-operate-most-strongly"&gt;Q4. At which organizational level (firm, occupation, or industry) do workplace moderators operate most strongly?&lt;/h3&gt;
&lt;p&gt;Firm and occupation levels are the dominant levels; industry-level measures are consistently insignificant for all three moderating variables. The authors interpret this as follows: industry-level measures are too broad to capture the specific work arrangements and norms that affect caregiving balance. At the occupation level, structural characteristics — profession-wide agreements, flexibility of task-based roles, part-time feasibility — directly govern how feasible it is to reduce hours without exiting employment. At the firm level, immediate workplace culture and specific HR policies apply. The relative contribution of firm vs occupation varies by the moderator: work hour intensity effects are significant at both firm and occupation levels, female senior representation is significant at both, while the gender pay gap effect is significant only at the occupation level.&lt;/p&gt;
&lt;h3 id="q5-why-does-female-representation-in-senior-roles-top-20-of-earners-mitigate-the-earnings-penalty-while-middle-and-bottom-tier-representation-does-not"&gt;Q5. Why does female representation in senior roles (top 20% of earners) mitigate the earnings penalty while middle and bottom tier representation does not?&lt;/h3&gt;
&lt;p&gt;The authors argue that women in the top-20% of earners — effectively leadership positions — are better positioned to advocate for and implement caregiving-supportive policies (paid leave, flexible scheduling). Representation in lower tiers may be indicative of a caregiving-friendly workforce composition but lacks the organizational power to shape policies. This is supported empirically: the moderating interaction is significant and economically large for top-20% female representation at both the firm (mitigating the penalty by $3,943) and occupation levels (mitigating by $4,116), while interactions for the middle 50–80% and bottom 50% earnings tiers are not statistically significant in most specifications.&lt;/p&gt;
&lt;h3 id="q6-why-does-the-occupational-gender-pay-gap-matter-for-the-earnings-penalty-but-not-the-firm-level-or-industry-level-gap"&gt;Q6. Why does the occupational gender pay gap matter for the earnings penalty but not the firm-level or industry-level gap?&lt;/h3&gt;
&lt;p&gt;The authors offer two explanations. First, occupations define the day-to-day nature of work — task structure, required hours, flexibility — in ways that make caregiving more or less compatible. Occupations that accommodate part-time and flexible scheduling tend to attract more women and develop norms that support caregiving, which in turn narrows occupational gender pay gaps. At the firm level, the same firm often contains diverse occupations with heterogeneous norms, so firm-level gender pay gap is a noisier signal. At the industry level, the measure is too aggregated. Second, narrow occupational gender pay gaps may reflect the collective bargaining power of women in female-dominated occupations (e.g., nursing), which translates into formal caregiving protections. A firm or industry may exhibit a wide gender pay gap due to male dominance in senior or high-earning roles even when specific female-dominated occupations within that firm/industry have caregiving-friendly norms. However, in the combined specification including all workplace factors simultaneously, the gender pay gap variable loses statistical significance, suggesting its initial effect was partly mediated by correlated factors (hours intensity and female senior representation).&lt;/p&gt;
&lt;h3 id="q7-how-does-the-combined-supportive-vs-unsupportive-comparison-work-and-what-does-it-show"&gt;Q7. How does the combined &amp;lsquo;supportive vs unsupportive&amp;rsquo; comparison work and what does it show?&lt;/h3&gt;
&lt;p&gt;Supportive jobs are defined as those satisfying all three criteria: low work hour intensity at both firm and occupation levels, high female representation in the top 20% of earners at both firm and occupation levels, and low gender pay gap at the occupation level (N = 2,708 mother-years). Unsupportive jobs are the opposite on all criteria (N = 2,339). Event study estimates (Table A9, Figure 3) show stark divergence. In supportive jobs, the year-0 penalty is −$6,280, and earnings recover quickly to statistically insignificant levels by years 1–4. In unsupportive jobs, the year-0 penalty is −$10,416, it widens to −$10,658 in year 2 and −$13,882 in year 3, before partially recovering in year 4. Pre-treatment estimates are not significantly different from zero in both subsamples, supporting parallel trends within each group.&lt;/p&gt;
&lt;h3 id="q8-what-heterogeneity-is-documented-by-child-and-family-characteristics"&gt;Q8. What heterogeneity is documented by child and family characteristics?&lt;/h3&gt;
&lt;p&gt;Appendix Figure A3 presents two subgroup analyses. Mothers of children under age 12 at diagnosis experience a year-0 earnings loss of −$9,443 (25.8% of baseline earnings of $36,567), substantially larger than the average. Mothers of children with above-median healthcare utilization (measured by number of medical appointments in the year following treatment initiation) experience a year-0 loss of −$7,826 (21.6% of baseline earnings of $36,278). These patterns are consistent with the interpretation that caregiving intensity — driven by child age and treatment severity — scales the maternal earnings penalty.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-conducted"&gt;Q9. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s main robustness arguments are: (1) pre-trend validation (Figures 1 and 2, Table A3) confirming no anticipatory effects and balanced pre-characteristics; (2) the selection-direction argument for workplace heterogeneity — the selection story would predict larger penalties in supportive workplaces but the opposite is found; (3) showing that absolute earnings declines in less-supportive workplaces also represent larger proportional declines relative to baseline, ruling out a level-effect interpretation; (4) the mental health non-result (Appendix Figure A4) confirming earnings effects are not confounded by parental mental health deterioration; (5) separate combined specification (Table A8) testing all workplace moderators simultaneously to address multicollinearity. The paper does not report explicit placebo tests using alternative shocks or falsification samples, nor does it report results restricted to narrow geographic areas or specific cancer types.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-prior-literature-on-caregiving-shocks"&gt;Q10. How does this paper relate to prior literature on caregiving shocks?&lt;/h3&gt;
&lt;p&gt;The paper builds most directly on three prior studies using Nordic or European administrative data: Eriksen et al. (2021, Journal of Health Economics) on childhood health shocks and parental labor supply; Breivik and Costa-Ramon (2024, Review of Economics and Statistics) on children&amp;rsquo;s health shocks and parental earnings and mental health; and Vaalavuo et al. (2023, Demography) on gender inequality from child health shocks on parental trajectories. All three find significant maternal earnings or employment losses and no or small paternal effects. The present paper&amp;rsquo;s contribution relative to these is the explicit examination of how firm-, occupation-, and industry-level workplace characteristics moderate the maternal penalty — a dimension the prior literature has not addressed. It also connects to Fadlon and Nielsen (2019, 2021) on the methodology and to the broader child-penalty literature reviewed by Cortes and Pan (2023, Journal of Economic Literature). On workplace mechanisms it connects to Goldin (2014) on &amp;lsquo;greedy jobs&amp;rsquo; and Goldin and Katz (2016) on pharmacy as a family-friendly profession.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The findings suggest that maternal earnings losses from caregiving shocks can be substantially mitigated by workplace environments characterized by lower work hour intensity and higher female representation in senior earnings tiers. This points to policies promoting: (1) reduced greedy-work norms — discouraging long-hours cultures and enabling part-time flexibility without disproportionate wage penalties; (2) greater female representation in leadership and high-earning positions, which appears to create cultural and policy environments more accommodating of caregiving. Scope conditions: the results apply to working mothers (and fathers) of children requiring chemotherapy in Australia, where Medicare provides universal healthcare coverage and existing social insurance exists. The paper explicitly does not identify specific causal mechanisms (e.g., it cannot isolate the effect of formal paid leave from culture). On fathers, the implication is that workplace factors alone are unlikely to induce fathers to increase caregiving, pointing instead to the need to shift social norms around paternal caregiving and intra-household bargaining.&lt;/p&gt;
&lt;h3 id="q12-how-do-the-australian-institutional-context-and-data-compare-to-european-studies"&gt;Q12. How do the Australian institutional context and data compare to European studies?&lt;/h3&gt;
&lt;p&gt;Australia&amp;rsquo;s PLIDA dataset is exceptional in combining population-level coverage, employer-employee identifiers (enabling firm-level workplace measures), and Medicare healthcare records (enabling both shock identification via chemotherapy and caregiving-intensity proxying via healthcare utilization). The employer identifiers are critical for this paper&amp;rsquo;s contribution — most comparable European studies cannot construct firm-level workplace characteristics. The Australian context differs from Nordic studies in terms of family policy generosity (less universal paid parental leave), but Medicare provides universal healthcare access. Pre-diagnosis earnings ($37,639 for mothers vs $79,702 for fathers) indicate a large pre-existing earnings gap, consistent with a majority-male breadwinner household structure in the sample.&lt;/p&gt;
&lt;h3 id="q13-do-fathers-outcomes-respond-to-any-workplace-factor"&gt;Q13. Do fathers&amp;rsquo; outcomes respond to any workplace factor?&lt;/h3&gt;
&lt;p&gt;In almost all specifications, fathers&amp;rsquo; earnings, employment, and job changes show no statistically significant effects of the caregiving shock and no significant interactions with workplace characteristics (Appendix Tables A4 and A6). One exception: in Table A4, the interaction between the cancer shock and working at a firm with above-median work hours is negative and significant at the 5% level for fathers, suggesting that fathers who work in high-hours firms do experience some earnings reduction — consistent with them reducing hours in an environment that penalizes deviations from long hours. However, the authors note the effect is substantially smaller relative to baseline earnings than the corresponding maternal effect. The broader pattern implies that workplace flexibility does not appear to be the binding constraint preventing fathers from taking on more caregiving; social norms and intra-household bargaining are posited as more important.&lt;/p&gt;
&lt;h3 id="q14-what-are-the-data-limitations-and-caveats"&gt;Q14. What are the data limitations and caveats?&lt;/h3&gt;
&lt;p&gt;First, work hours at the firm and occupation levels are constructed from the 2011 Census, which is a single cross-section; work hour norms may have shifted between 2011 and the 2012–2023 sample period. Occupation and industry codes also come from the 2011 Census, so parents who changed occupation between 2011 and their baseline year may be misclassified. Second, employment status is inferred from positive ATO earnings in a financial year, a coarser measure than actual employment spells. Third, the sample is restricted to firms with at least 10 employees, which excludes small-firm workers. Fourth, the analysis uses dollar earnings levels, not log earnings, which means baseline wage differences across workplace types can affect the interpretation of absolute dollar results (though the authors show percentage effects are also larger in less-supportive workplaces). Fifth, the study identifies workplace correlates of smaller penalties but does not isolate the causal effect of any specific policy. Sixth, the paper covers only cancer requiring chemotherapy — typically more intensive cancers — so results may overstate average caregiving-shock effects.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Caregiving shock&lt;/strong&gt;: In this paper, a sudden, largely unanticipated increase in caregiving demands on parents triggered by a child&amp;rsquo;s initiation of chemotherapy. Distinguished from the chronic caregiving burden of childbirth; specifically refers to health events that arrive later in childhood and impose large, time-intensive care requirements.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Later-treated dynamic DiD&lt;/strong&gt;: The paper&amp;rsquo;s identification design, following Fadlon and Nielsen (2019, 2021), in which the control group consists of parents who will receive the same treatment (child&amp;rsquo;s cancer diagnosis) at a later date. The control group&amp;rsquo;s placebo treatment year is set six years before their actual treatment, enabling estimation of time-path effects relative to diagnosis while accounting for pre-existing differences via individual fixed effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Work hour intensity&lt;/strong&gt;: Median weekly hours worked by employees at a given firm or in a given occupation (from the 2011 Census), used as a proxy for &amp;lsquo;greedy job&amp;rsquo; characteristics — workplaces that reward continuous long-hours presence and penalize deviations. High work hour intensity captures both above-full-time norms and the likely presence of evening and weekend work requirements.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Female representation in the top 20% of earners&lt;/strong&gt;: A binary indicator equal to one when women are the majority (above 50%) of workers in the top quintile of earnings at a given firm or occupation. The paper distinguishes this from female representation in middle and lower earnings tiers to isolate the effect of women&amp;rsquo;s presence in positions with organizational power to influence workplace policies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Supportive job&lt;/strong&gt;: As defined operationally in this paper: a job in which the worker&amp;rsquo;s firm and occupation both have below-median work hour intensity, both have majority female representation in the top 20% of earners, and the occupation has a below-average gender pay gap. Mothers in supportive jobs suffer smaller and shorter-lived earnings penalties following a caregiving shock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Greedy occupation&lt;/strong&gt;: Borrowed from Goldin (2014), and used in this paper to describe occupations that disproportionately reward workers who supply long, often inflexible, hours. In the paper&amp;rsquo;s empirical framework, these are occupations with above-median work hour intensity, which are shown to amplify maternal earnings losses after a caregiving shock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Caregiving intensity&lt;/strong&gt;: The time-varying burden of care associated with a child&amp;rsquo;s illness, proxied in this paper by the volume of child healthcare service utilization (Medicare items: GP visits, specialist consultations, diagnostic imaging, prescriptions). Caregiving intensity peaks at year 0 (treatment initiation), declines significantly by year 2, and returns to baseline by year 3 — yet maternal earnings penalties persist beyond this return to baseline.&lt;/p&gt;
&lt;!-- flags: Employment figures cited in the text (4.9 pp in year 0; peak of 5.6 pp in year 2) differ slightly from Table A3 values (-0.045 = 4.5 pp in year 0; -0.050 = 5.0 pp in year 2). This is a within-paper discrepancy in the IZA working paper version. Layer 1 reports the text-stated figures as authored. --&gt;</description></item><item><title>Bargaining with renegotiation in models with on-the-job search</title><link>https://macropaperwarehouse.com/papers/bargaining-with-renegotiation-in-models-with-on-the-job-search/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/bargaining-with-renegotiation-in-models-with-on-the-job-search/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper resolves a long-standing theoretical impasse in labor search models: how to model wage bargaining when workers search on the job (OJS) and the quit rate depends on the wage. Shimer (2006) showed that this wage-dependent turnover creates a potentially non-convex bargaining set, causing the Nash bargaining solution to break down and generating equilibrium multiplicity. Gottfries introduces renegotiation — wages are fixed under a contract that expires at a Poisson rate γ, after which a new wage is bargained — as the device that simultaneously restores uniqueness and nests the earlier models of Pissarides (1994), Mortensen (2003), and Shimer (2006) as limit cases.&lt;/p&gt;
&lt;p&gt;The model is a continuous-time, frictional labor market with risk-neutral firms and workers. Unemployed workers receive job offers at rate λu; employed workers receive outside offers at rate λe; matches dissolve exogenously at rate δ. Wages are determined through non-cooperative alternating-offers bargaining in the spirit of Rubinstein (1982) and Binmore et al. (1986), with worker bargaining power β. The key innovation is that contracted wages last until renegotiation, which arrives at a Poisson rate γ(F), where F indexes match quality (and hence wage expectations about future renegotiations). As γ → ∞ (continuous renegotiation), the model converges to Pissarides (1994): values solve the Nash bargaining solution with perfectly transferable values, and worker turnover is independent of the current contracted wage. As γ → 0 (no renegotiation), the model converges to the unique equilibrium from Shimer (2006) and Mortensen (2003), with wages playing a strong role in retaining workers. Equilibrium uniqueness follows because renegotiation makes match types payoff-relevant — wage expectations about future negotiations differ across types, so the Nash product cannot be constant on the support, pinning down the initial condition for the wage differential equation.&lt;/p&gt;
&lt;p&gt;The main mechanism is a turnover-retention channel that amplifies worker bargaining power. Because a higher wage reduces the quit rate, and marginal quits are bilaterally inefficient (the firm loses its profits when the worker leaves), agreeing on a higher wage partially recoup losses through longer match duration. This acts as an additional source of worker surplus share on top of the primitive bargaining power β. The strength of this channel is governed by θ — the expected fraction of the discounted match duration covered by a given contracted wage. Higher θ (less frequent renegotiation) means wages matter more for turnover and workers extract more surplus. Lower θ (more frequent renegotiation) attenuates the channel.&lt;/p&gt;
&lt;p&gt;Calibrated to US labor market data — a 45% monthly job-finding rate (Shimer 2012), a 3.2% monthly job-to-job transition rate (Moscarini and Thomsson 2007), a 5% unemployment rate, a 5% annual discount rate, and targeting a labor share of 2/3 and a lognormal wage-offer distribution with scale parameter σ = 0.16 (Gottfries and Teulings 2017) and a mean-to-minimum wage ratio of 1.7 (Hornstein et al. 2007) — the model implies sharply different primitive bargaining powers depending on the assumed renegotiation frequency. Under continuous renegotiation (γ = ∞), the calibrated bargaining power of workers is β = 0.46. Under never-renegotiated wages (γ = 0), β = 0.02. The implication is that the correct inference about worker bargaining power from observed wage distributions is very sensitive to the assumed renegotiation regime.&lt;/p&gt;
&lt;p&gt;For minimum wages, the paper proves that, holding firm entry and the reservation wage constant, any minimum wage increase raises the entire wage distribution in the sense of first-order stochastic dominance (Proposition 2). However, the extent of spillovers above the minimum wage depends critically on renegotiation frequency. In a high-commitment economy (low γ) versus a low-commitment economy (high γ) with identical pre-policy wage distributions, the high-commitment economy exhibits strictly larger wage spillovers throughout the support above the minimum (Proposition 3). The intuition is that a spike in the mass of workers at the minimum wage creates a strong incentive for firms to offer higher wages to reduce costly turnover — but this incentive only materializes when wages are sticky enough that turnover responds appreciably to them. With continuous renegotiation, the spillover vanishes entirely and only a mass point at the minimum wage remains. In the limit of no renegotiation, the model resembles the wage-posting model, which produces especially large spillovers by construction.&lt;/p&gt;
&lt;p&gt;An extension endogenizes the contract length. Firms optimally choose the renegotiation frequency after observing the match type. Two regimes emerge: when worker bargaining power is sufficiently high or productivity rises quickly relative to profits, firms prefer continuous renegotiation; otherwise, an interior contract length strictly above zero is optimal, and firms with all the bargaining power prefer no renegotiation. This implies that the polar assumptions of full commitment or no commitment standard in the literature arise only as boundary cases.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-theoretical-problem-this-paper-addresses"&gt;Q1. What is the core theoretical problem this paper addresses?&lt;/h3&gt;
&lt;p&gt;Shimer (2006) demonstrated that when a worker&amp;rsquo;s quit rate depends on the contracted wage, the bargaining set can become non-convex, violating a key condition for the Nash bargaining solution. He proposed a non-cooperative alternating-offers bargaining game but showed that it produces a continuum of equilibria. The existing literature responded either by removing bargaining (wage posting, all bargaining power to firms) or by making turnover independent of the wage (counteroffers by the incumbent firm). Gottfries provides a solution that preserves both bargaining and wage-dependent turnover by introducing renegotiation.&lt;/p&gt;
&lt;h3 id="q2-how-does-renegotiation-restore-equilibrium-uniqueness"&gt;Q2. How does renegotiation restore equilibrium uniqueness?&lt;/h3&gt;
&lt;p&gt;Without renegotiation and with homogeneous productivities (as in Shimer 2006), the match type F is not payoff-relevant: only the current contracted wage matters, so the Nash product is constant on the wage support and any wage in that support is a potential equilibrium outcome. With renegotiation, each type F is associated with a distinct expected future wage (wage expectation), which is payoff-relevant because it governs future turnover. Different types therefore face different Nash products, and the product cannot be constant across types. This forces the Nash product to be increasing to the left of the bargaining outcome and decreasing to the right for each type, providing a unique interior maximum and a unique initial condition w(0) = max{βx(0) + (1−β)wr, wmin}. The paper also shows that alternative refinements — large-friction limits or the case where λe = 0 — yield the same unique equilibrium.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-model-nest-pissarides-1994-and-mortensen-2003"&gt;Q3. How does the model nest Pissarides (1994) and Mortensen (2003)?&lt;/h3&gt;
&lt;p&gt;As γ → ∞ (continuous renegotiation, θ → 0), the contracted wage becomes irrelevant because future wages are renegotiated almost immediately. The worker&amp;rsquo;s quit decision is then independent of the current wage, so values solve the standard Nash bargaining solution with perfectly transferable values, exactly as in Pissarides (1994). As γ → 0 (no renegotiation, θ → 1), the wage lasts the full duration of the match, turnover responds maximally to wages, and the equilibrium values correspond to Mortensen (2003, Section 4.3.4) with a unique initial condition (rather than the multiplicity in Shimer 2006). Intermediate values of γ correspond to no prior model.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-mechanism-by-which-workers-receive-a-share-of-surplus-exceeding-their-bargaining-power-β"&gt;Q4. What is the mechanism by which workers receive a share of surplus exceeding their bargaining power β?&lt;/h3&gt;
&lt;p&gt;When a worker bargains for a higher wage, she reduces her quit probability. Marginal quits are bilaterally inefficient because the firm loses its profits when the worker leaves to a marginally better job (even though the transition is socially efficient once the new employer&amp;rsquo;s value is counted). The reduction in inefficient separations increases the joint match surplus. Formally, the extra surplus share comes from the term λe · [w&amp;rsquo;(F)/(δ+ρ+λe(1−F))] · [(δ+ρ+λe(1−F))/(δ+ρ+γ(F)+λe(1−F))] · Π(F,w(F)), which is the density of incoming offers per unit wage increase multiplied by the fraction of the match duration covered by the contracted wage, multiplied by the profit level lost at each marginal quit. This term is zero when γ → ∞ (continuous renegotiation) and is largest when γ = 0 (no renegotiation).&lt;/p&gt;
&lt;h3 id="q5-what-is-θ-and-what-role-does-it-play"&gt;Q5. What is θ and what role does it play?&lt;/h3&gt;
&lt;p&gt;θ is defined for the homogeneous-productivity case as the expected fraction of the expected discounted match duration that an agreed wage remains in force. It captures the marginal relative importance of the current contracted wage versus the wage expectation (which governs future renegotiated wages). θ = 1 corresponds to no renegotiation (the contracted wage lasts the whole match), θ → 0 corresponds to continuous renegotiation. The renegotiation rate is γ(F) = [(1−θ)/θ] · (δ+ρ+λe(1−F)). A small increase in the wage by w&amp;rsquo;(F)dF decreases turnover by θ dF in the homogeneous case, so θ directly scales the turnover-retention channel and hence workers&amp;rsquo; effective surplus share.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-calibration-reveal-about-the-relationship-between-renegotiation-assumptions-and-inferred-bargaining-power"&gt;Q6. What does the calibration reveal about the relationship between renegotiation assumptions and inferred bargaining power?&lt;/h3&gt;
&lt;p&gt;Holding transition rates fixed (λu = 0.45, λe = 0.181, δ = 0.024 per month) and targeting a 2/3 labor share and a lognormal wage-offer distribution (σ = 0.16, mean-min ratio 1.7), the calibrated worker bargaining power β is 0.46 under continuous renegotiation (γ = ∞) and only 0.02 under no renegotiation (γ = 0). The calibrated productivity distribution also differs markedly: no-renegotiation requires a much fatter right tail in firm productivities to match the same wage distribution because the labor share falls sharply in the upper tail when bargaining power is low and wages are infrequent renegotiated.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-paper-prove-about-minimum-wage-spillovers"&gt;Q7. What does the paper prove about minimum wage spillovers?&lt;/h3&gt;
&lt;p&gt;Proposition 2 proves that, holding firm entry constant and adjusting unemployment benefits to keep the reservation wage constant, a minimum wage increase raises the equilibrium wage distribution in the sense of first-order stochastic dominance. Proposition 3 proves that, comparing a high-commitment economy H (lower γH) and a low-commitment economy L (higher γL) that have identical pre-policy wage distributions (and therefore βH &amp;lt; βL), the high-commitment economy H exhibits strictly higher wages at every rank F after a small minimum wage increase. The mechanism is that a mass of workers at the minimum wage creates a dense region of outside options, making it worthwhile for firms to accept higher wages to reduce turnover — but only when committed wages are sticky enough to affect actual turnover.&lt;/p&gt;
&lt;h3 id="q8-what-happens-to-the-wage-distribution-spike-at-the-minimum-wage-when-renegotiation-is-frequent"&gt;Q8. What happens to the wage distribution spike at the minimum wage when renegotiation is frequent?&lt;/h3&gt;
&lt;p&gt;Under the baseline assumption that workers move when indifferent (no mass points), the equilibrium has no spike; the mass at the minimum wage spreads continuously upward. When this assumption is relaxed and workers may stay when indifferent (following Shimer 2006), an equilibrium with a mass point at the minimum wage exists. Equation (19)/(20) show the equilibrium mass point at the minimum wage is increasing in the renegotiation rate γ (higher γ → larger spike). This occurs because with frequent renegotiation, spillovers above the minimum wage are small, so the density just above the minimum is high, which in turn supports a large mass at the minimum. The paper parameterizes this with φ = 0.04 (ratio of mass at minimum wage to density just above) and illustrates with θ = 0.02 (long contracts) and θ = 0.5 (short contracts).&lt;/p&gt;
&lt;h3 id="q9-how-does-endogenizing-the-contract-length-change-the-predictions"&gt;Q9. How does endogenizing the contract length change the predictions?&lt;/h3&gt;
&lt;p&gt;When firms choose the renegotiation frequency after observing the match type, two regimes emerge. In the first, the firm would not benefit from raising the wage above the continuous-renegotiation Nash-bargaining level: this happens when worker bargaining power is sufficiently high or productivity increments are large relative to profits. Firms then choose continuous renegotiation (γ = ∞) for that match type. In the second regime, lower turnover makes it profitable to commit to a higher wage via a longer contract; firms pick an interior γ satisfying the envelope condition. With all bargaining power to the firm (β = 0), the optimum is no renegotiation (infinite contract length). The equilibrium in the endogenous-contract model satisfies a differential equation that coincides with the wage-posting model differential equation in the interior region, providing a microfoundation for wage-posting results even when workers have some bargaining power. The model also provides a uniqueness justification for equilibria in Coles (2001) and Coles and Mortensen (2016).&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-brügemann-gautier-and-menzio-2015"&gt;Q10. How does this paper relate to Brügemann, Gautier, and Menzio (2015)?&lt;/h3&gt;
&lt;p&gt;Brügemann, Gautier, and Menzio (2015) identify a similar surplus-retention mechanism in a model where a single firm bargains successively with many workers: agreeing on a high wage with one worker is &amp;lsquo;cheap&amp;rsquo; because the firm can recoup part of the cost through lower wages agreed with subsequent workers. Gottfries&amp;rsquo; mechanism is the bilateral analogue: within a single match, a higher wage is cheap because it reduces wasteful turnover and extends the profitable match duration. Both models generate workers capturing a surplus share above their primitive bargaining power, but through distinct channels.&lt;/p&gt;
&lt;h3 id="q11-what-assumptions-are-needed-for-uniqueness-and-what-relaxing-them-implies"&gt;Q11. What assumptions are needed for uniqueness and what relaxing them implies?&lt;/h3&gt;
&lt;p&gt;Two key restrictions are imposed. First, Markov strategies are required and wage functions must be weakly increasing in match type F; without this, equilibria exist in which workers accept lower-productivity jobs for a higher current wage, creating decreasing wage functions. Second, workers must move with positive probability when indifferent between offers, which eliminates mass points on the support. Shimer (2006) showed that when indifferent workers never move, multiple equilibria with mass points exist. Relaxing the second restriction opens the door to a spike at the minimum wage in the minimum wage application. Alternative refinements — large-friction limits, the limiting case as λe → 0, or as β → 0 — all single out the same unique equilibrium.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The main policy implication is that the spillover effects of minimum wage increases depend critically on the degree of wage commitment in the labor market. In economies where wages are rarely renegotiated (higher θ), minimum wage increases spread substantially up the wage distribution; in economies with continuous renegotiation, only a spike at the minimum results with little or no spillover. This has direct implications for empirical studies of minimum wages: the observed pattern of spillovers is informative about the prevailing renegotiation regime. The scope conditions are: (i) partial equilibrium (firm entry and reservation wage are held fixed); (ii) all matches remain profitable at the minimum wage (wmin &amp;lt; x(0)); (iii) random rather than directed search. The paper does not provide an empirical test or identification strategy for the renegotiation frequency itself.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-limits-and-caveats"&gt;Q13. What are the limits and caveats?&lt;/h3&gt;
&lt;p&gt;The model treats the renegotiation frequency as an exogenous parameter (except in Section 6). The calibration does not structurally identify the renegotiation frequency from data; it instead illustrates sensitivity. The analysis of minimum wages is partial equilibrium — firm entry and reservation wages are held fixed — and the paper notes that general equilibrium effects (entry, reservation wages) are ambiguous in sign and difficult to identify empirically. The model has no on-the-job search effort endogeneity or worker heterogeneity (workers are homogeneous ex ante). The wage-posting and counteroffers models studied in the literature require strong commitment assumptions that this model relaxes but does not fully endogenize in a dynamic contracting sense.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Renegotiation (frequency parameter γ)&lt;/strong&gt;: The Poisson rate at which a contracted wage expires and a new wage is bargained. In the paper&amp;rsquo;s own sense, γ indexes the degree of wage commitment: γ = 0 means the contracted wage lasts the entire match (perfect commitment, no renegotiation); γ → ∞ means the wage is continuously reset (no commitment). The frequency γ governs how much the contracted wage — versus future renegotiated wages — matters for the worker&amp;rsquo;s turnover decision, and hence how much of the match surplus the worker captures.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bilateral inefficiency of transitions&lt;/strong&gt;: The paper defines a job-to-job transition as bilaterally inefficient when the value to the worker at the new job is less than the total surplus of the existing match. Since the firm loses its profits when the worker quits, the pair jointly would prefer the worker to stay — yet the worker moves whenever her individual value is higher elsewhere. The gap between individual and joint incentives is the source of bilateral inefficiency; it is what makes turnover-reduction through higher wages mutually beneficial and gives workers extra bargaining power beyond β.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Match type (F) and wage expectation&lt;/strong&gt;: In the model, F is a match quality drawn from the uniform distribution on [0,1] upon meeting. F determines both the productivity x(F) and the wage expectation — the anticipated outcome of future renegotiations. Critically, the wage expectation is the payoff-relevant state variable that differs across types and thereby distinguishes matches, restoring equilibrium uniqueness. Higher F is associated with higher wage expectations, lower turnover, and greater match surplus.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Commitment parameter (θ)&lt;/strong&gt;: Defined for the homogeneous-productivity case as the expected fraction of the expected discounted match duration for which the currently agreed wage remains in force. θ = 1 corresponds to no renegotiation; θ → 0 to continuous renegotiation. A one-unit wage increase reduces turnover by θ in equilibrium, so θ directly scales the turnover-retention channel and the extra surplus share flowing to workers beyond their primitive bargaining power β.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Minimum wage spillover&lt;/strong&gt;: The paper uses &amp;lsquo;spillover&amp;rsquo; to mean the upward shift in wages paid by firms above the minimum wage that results from a minimum wage increase. Mechanically, a minimum wage creates a mass of workers at the floor; if turnover responds to wages (i.e., commitment is high), firms above the minimum prefer to raise wages to avoid losing workers to the mass point competitors, spreading the effect. The paper proves (Proposition 3) that spillovers are strictly larger in higher-commitment (lower γ) economies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Markov-perfect equilibrium (MPE) of the bargaining game&lt;/strong&gt;: The equilibrium concept applied to the alternating-offers bargaining game. In an MPE, offer and acceptance rules depend only on the current match type F, not on prior bargaining history. This restriction, combined with the renegotiation structure, is what allows the paper to derive a unique differential equation for the wage function w(F) and a unique initial condition, yielding the unique equilibrium wage distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Turnover-retention channel&lt;/strong&gt;: The mechanism by which a higher contracted wage reduces the worker&amp;rsquo;s quit probability and thereby increases the joint match surplus. Because marginal quits are bilaterally inefficient, a small wage increase generates a surplus gain proportional to the density of arriving outside offers times the expected fraction of the match covered by the contracted wage times firm profits — exactly the extra term that elevates the worker&amp;rsquo;s effective surplus share above β. This channel is the paper&amp;rsquo;s central contribution to understanding why workers capture more than their bargaining power suggests.&lt;/p&gt;</description></item><item><title>Codification, Technology Absorption, and the Globalization of the Industrial Revolution</title><link>https://macropaperwarehouse.com/papers/codification-technology-absorption-and-the-globalization-of-the-industrial-revolution/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/codification-technology-absorption-and-the-globalization-of-the-industrial-revolution/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Why did the First Industrial Revolution (IR) spread to Meiji Japan—and to essentially no other non-Western country—during the first wave of globalization? The paper tests Mokyr&amp;rsquo;s hypothesis that &amp;ldquo;technical literacy,&amp;rdquo; i.e., the codification of engineering, commercial, and industrial knowledge in the local vernacular, was a necessary condition for absorbing IR technologies. The motivating puzzle: after opening to trade (1858) and the Meiji Restoration (1868), 80% of Japanese exports were still primary products as late as ~1883 and real per capita GDP growth was only 0.6%/yr (1870-1883/85); then in a brief 13-year window (1883-1896) the manufacturing export share tripled and stabilized at around 60% of exports until WWII.&lt;/p&gt;
&lt;p&gt;Data and setup: The authors build several novel datasets. (1) A cross-language measure of codification: scraping national/major libraries and WorldCat for technical books (agriculture, applied sciences, commerce, industry, technology) in 33 languages, 1500-1930. (2) &amp;ldquo;British Patent Relevance&amp;rdquo; (BPR): the cosine similarity (TF-IDF, unigrams+bigrams) between the digitized synopses of all British patents 1780-1852 (from Woodcroft 1857) and a hand-curated corpus of 460 English-language 19th-century technical manuals matched to SITC industries. BPR measures the world supply of codifiable IR knowledge by industry and is deliberately not based on what Japan translated (to avoid endogeneity). (3) The first harmonized, bilateral, industry-level trade dataset for the 19th century: 37 regions, 93 industries, quinquennial 1880-1910, built from reporting countries Japan, US, Belgium, Italy. Outcomes are annualized industry export growth ({1880,1885} to {1905,1910}) and, in robustness, productivity/comparative-advantage growth following Costinot et al. (2012) and Amiti-Weinstein (2018).&lt;/p&gt;
&lt;p&gt;Main findings (with magnitudes): A Japanese industry with a one-standard-deviation higher BPR experienced annual export growth ~12 percentage points faster and annual productivity (comparative-advantage) growth ~1.2 percentage points faster (coefficients 0.121*** and 0.012***). Cross-sectionally, the BPR-growth relationship is positive and significant only for Japan and other codifying countries: for non-Japan regions the BPR coefficient is negative (-0.030***), while English-, French-, and the &amp;ldquo;top-4 codified&amp;rdquo; (English/French/German/Italian) regions show positive coefficients (0.042**, 0.032**, 0.078***), smaller than Japan&amp;rsquo;s. Low-income and Asian regions tend negative (divergence), not always significant. Time-series: regressing Japanese export growth from 1875 to varying end-years, the BPR coefficient is negative/significant in the 1875-1880 placebo window (Japan resembled the periphery), flips around 1890, and is positive and significant at 1% by 1895—coinciding with Japan&amp;rsquo;s catch-up in codification.&lt;/p&gt;
&lt;p&gt;Mechanism and the Meiji &amp;ldquo;natural experiment&amp;rdquo;: In 1870, 84% of all technical books were in four languages (English, French, German, Italian); an Arabic-only reader had access to just 71 technical books. Japan started ordinary but codified explosively: technical-book growth jumped from 1.6%/yr (1600-1860) to 8.8%/yr (1870-1900); translated technical books rose from 8 (1500-1860) to 608 by 1900; Japanese technical books in the NDL grew from 706 (1880) to 2,823 (1890). State provision solved a public-goods/coordination problem: the government built English-Japanese dictionaries (ETSJ 1862/1866, FSEJ 1871) creating standardized Japanese jargon from Chinese glyphs, and 74% of identified technical-book translators (1870-1885) were government employees. Implication: low-cost vernacular access to technical knowledge was a necessary (not sufficient) condition for IR diffusion; where regions were linguistically/geographically distant from Western Europe, codification required state provision (a Gerschenkronian role for the state).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-the-main-threats-to-it"&gt;Q1. What is the identification strategy and the main threats to it?&lt;/h3&gt;
&lt;p&gt;Two-pronged. (1) Cross-sectional: regress region-industry export growth on BPR interacted with region-group dummies, with exporter fixed effects, exploiting that BPR is global (not Japan-specific) and that Japan was uniquely a codifier in the periphery. If codification is the mechanism, only codifying regions should show a positive BPR-growth link. (2) Time-series: exploit the sharp timing of Japanese codification (two well-demarcated periods—pre vs. post technical literacy in the 1880s) by estimating the BPR coefficient on Japanese export growth from 1875 to rolling end-years. The 1875-1880 window serves as a placebo (Japan not yet literate). Main threat is omitted-variable bias: that BPR is correlated with distance to the technology frontier, fundamental comparative advantage, Meiji institutional reforms, or industry steam-intensity. The cross-section addresses the &amp;lsquo;BPR matters everywhere&amp;rsquo; and income/geography confounds; the timing addresses slow-moving confounds (literacy, Tokugawa culture, gradual reforms) since reforms like tax/banking/railroads were mostly in place by 1875, 15-37 years before the BPR effect appears.&lt;/p&gt;
&lt;h3 id="q2-how-are-the-cross-section-and-time-series-results-distinguished-from-confounders-empirically"&gt;Q2. How are the cross-section and time-series results distinguished from confounders empirically?&lt;/h3&gt;
&lt;p&gt;In the cross-section, income terciles (High/Medium/Low) and an Asia dummy are added: no region group replicates Japan&amp;rsquo;s positive pattern; the poorest and Asian regions show negative (divergence) coefficients. The placebo (1875-1880) yields a negative significant BPR coefficient for Japan itself—identical in sign to non-codifiers—then flips positive/significant by 1895, which conventional &amp;lsquo;opening to trade&amp;rsquo; (1858) or &amp;lsquo;Meiji Restoration&amp;rsquo; (1868) stories cannot explain because the effect appears 37 and 27 years later, respectively.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Japan&amp;rsquo;s BPR coefficient is larger (though not always significantly) than that of European codifiers, consistent with Japan having more to learn from British patents as a late industrializer. Among non-codifiers, low-income and Asian regions show negative BPR-growth relationships (divergence). Within codifiers, English- and French-speaking regions individually have positive but smaller and less precisely estimated coefficients; pooling the top-4 codified languages sharpens significance (0.078***). The time-series point estimates for Japan slowly decline after 1900 (not significantly), consistent with Japan shifting to Second Industrial Revolution technologies and becoming less reliant on older IR ones.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;(1) Alternative patent corpora: results are nearly identical using British patents 1853-1879 (full text and AI-summarized) and US patents 1836-1860 and 1861-1879 (coefficients 0.121, 0.116, 0.111, 0.115), though later/US patents lower the R-squared, suggesting the 1780-1852 IR patents best explain Japanese export growth. (2) Productivity instead of exports (Costinot et al. 2012 comparative-advantage growth): qualitatively the same, 1.2 pp/yr for a 1-SD BPR increase, with deterioration in non-codifiers. (3) Confounders: controlling for British-colony status (insignificant) and industry steam-power intensity (French 1860s data) does not affect results. (4) Sample selection: dropping non-manufacturing sectors, excluding Asian destination markets, and dropping major export products (textiles, iron/metal) all leave the results intact, indicating broad-based change.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q5. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;It builds on Mokyr (2011) on &amp;rsquo;technical knowledge&amp;rsquo;/&amp;lsquo;access costs&amp;rsquo; for European industrialization, extending it outside Europe with a Gerschenkronian twist (state as provider of the codification public good). It contributes to the technology-adoption-lags literature (Comin and Hobijn 2010; ~45-year average lags) by offering a friction explanation. It departs from prior Meiji studies (Sussman-Yafeh 2000; Tang; Morck-Nakamura; Bernhofen-Brown) that found banking, railroads, constitutional/monetary reforms had little measurable growth impact—offering codification as the resolution to &amp;lsquo;what drove the Meiji Miracle,&amp;rsquo; consistent with Broadberry et al. (2025) dating Japan&amp;rsquo;s convergence to ~1890 driven by manufacturing productivity. It also extends the knowledge-codification literature (Dittmar 2011; Brown 2024; Abramitzky-Sin 2014) by linking codified vernacular knowledge directly to industry growth rather than indirect outcomes like city growth.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Public provision of technical knowledge in the vernacular can relax a critical bottleneck to industrialization, especially for regions linguistically/geographically distant from the technology frontier where the market undersupplies this public good. Scope conditions: codification is necessary but NOT sufficient. The Meiji model required complementary investments—language/jargon standardization, mass education for absorptive capacity (literacy &amp;gt;90% for army conscripts by 1909; ~40% of elementary class time on science), tacit-knowledge acquisition (2,400 hired foreigners providing 9,506 person-years of training; study-abroad missions), and tax capacity (1873 Land Tax Reform). China&amp;rsquo;s post-1949 codification under Zhou did not yield sustained growth until Maoist policies (Great Leap, Cultural Revolution) ended—&amp;rsquo;the exception that proves the rule.&amp;rsquo;&lt;/p&gt;
&lt;h3 id="q7-what-external-validity-evidence-is-offered-beyond-japan"&gt;Q7. What external-validity evidence is offered beyond Japan?&lt;/h3&gt;
&lt;p&gt;The Meiji codification model was studied and transplanted by Park Chung Hee in South Korea (took power 1961; KIST; researcher counts rose sharply) and Zhou Enlai in China (premier 1949; Russian-language translation drive with USSR as the &amp;lsquo;Britain&amp;rsquo;). In 1950, Japan had ~70,000 technical books, China ~1,000, Korea &amp;lt;100; China surpassed 30,000 by the early 1960s. Korea&amp;rsquo;s per capita income clearly rises after Park; China&amp;rsquo;s codification did not translate into growth until after 1976. These are explicitly presented as suggestive/non-causal, plus appendix discussions of British India and Late Imperial Russia.&lt;/p&gt;
&lt;h3 id="q8-what-are-notable-caveats-and-measurement-choices"&gt;Q8. What are notable caveats and measurement choices?&lt;/h3&gt;
&lt;p&gt;BPR uses British 1780-1852 patent synopses and English manuals deliberately (Britain as IR leader; Japan hired British instructors and used British textbooks; avoids endogeneity from Japanese translation choices). It excludes tacit knowledge and secrecy-protected innovation by design. English codification is likely underestimated (British Library was un-scrapable after a 2023 cyberattack; Library of Congress used instead). German patents/trade data were excluded for coverage/reliability reasons. Linguistic-distance evidence on 1870/1913 GDP is explicitly not interpreted causally. The aggregate growth correlations for Japan, Korea, and China are described as suggestive, not causal.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Codification (of technical knowledge)&lt;/strong&gt;: The creation of a means of transmitting engineering, commercial, and industrial knowledge—via language creation and written messages (manuals, textbooks, dictionaries)—that does not require direct contact between the knowledge originator and the recipient (Cowan and Foray 1997). In the paper&amp;rsquo;s sense it is a non-rival public good that the market undersupplies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Technical literacy / technical knowledge&lt;/strong&gt;: Following Stevens (1995) and Mokyr, the codified engineering, commercial, and industrial practices a practitioner needs to set up and run modern factory-based manufacturing; the paper measures it as the stock of vernacular technical books (agriculture, applied sciences, commerce, industry, technology), excluding theoretical/hard-science and non-firm subjects like medicine.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;British Patent Relevance (BPR)&lt;/strong&gt;: An industry-level measure equal to the cosine similarity (TF-IDF weighted) between the vectorized text of British patent synopses (1780-1852) and the vectorized text of English technical manuals for that industry; it proxies how much codifiable IR knowledge a given industry stood to gain, and is independent of what was actually translated into Japanese.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Access costs&lt;/strong&gt;: Mokyr&amp;rsquo;s (2011) term for the cost of obtaining usable technical knowledge; the paper argues vernacular codification (dictionaries, translations) lowered these costs, and that linguistic distance from English/Latin-Greek roots and physical distance from Europe raised them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Technology absorption / absorptive capacity&lt;/strong&gt;: The complementary conditions needed to use codified knowledge—prior language/jargon development, literacy and scientific training, and tacit knowledge—all of which the Meiji state invested in (dictionaries, compulsory education, &amp;rsquo;live machines&amp;rsquo;/foreign instructors, study-abroad missions).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Defensive modernization (Gerschenkronian state role)&lt;/strong&gt;: The paper&amp;rsquo;s reading that an existential external threat aligned the Japanese elite behind aggressive state-led adoption of Western science, casting the state as the critical agent supplying the codification public good in late industrialization—a Gerschenkronian extension of Mokyr applied outside Europe.&lt;/p&gt;</description></item><item><title>Corrigendum to "Job Ladders by Firm Wage and Productivity" [Review of Economic Dynamics 58C (2025) 101307]</title><link>https://macropaperwarehouse.com/papers/corrigendum-to-job-ladders-by-firm-wage-and-productivity-review-of-economic-dynamics-58c-2025-101307/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/corrigendum-to-job-ladders-by-firm-wage-and-productivity-review-of-economic-dynamics-58c-2025-101307/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation. On-the-job search models typically organize firms along a &amp;ldquo;job ladder&amp;rdquo; — a common ranking by workers of available jobs — but they disagree on whether the rung is best captured by a firm&amp;rsquo;s average wage or its productivity, and empirical guidance has been scarce. Bertheau and Vejlin ask: (i) Is average wage or productivity the better empirical measure of a firm&amp;rsquo;s location on the job ladder? (ii) How does job creation across these ladders vary in the cross-section and over the business cycle? (iii) Do recessions slow reallocation into better firms (a &amp;ldquo;sullying&amp;rdquo; effect) or speed it up (a &amp;ldquo;cleansing&amp;rdquo; effect)? This matters for models of aggregate labor-market fluctuations and any imperfect-labor-market model that assumes some jobs are more desirable than others.&lt;/p&gt;
&lt;p&gt;Data and strategy. The authors build matched employer-employee data from Danish administrative registers covering all employment relationships at DAILY frequency from 1992 to 2013, merged with firm financial-accounting data (sales, value added, capital stock, FTE employment, workforce composition). The sample is restricted to manufacturing, services, and trade (industries present from 1992); aggregate unemployment ranges from 3% to 10% over the period, spanning several recessions. Daily timing removes the time-aggregation bias of quarterly data (Bertheau and Vejlin 2022 show quarterly data overstate the EE transition rate by ~30%). Firms are ranked within industry-year cells by (a) residualized average hourly wage and (b) total factor productivity (TFP) estimated via the Olley-Pakes (1996) control-function approach (investment data available from 1999). Following Haltiwanger et al. (2018b), &amp;ldquo;low&amp;rdquo; firms are the bottom employment-weighted quintile and &amp;ldquo;high&amp;rdquo; firms the top two quintiles. Net employment change is decomposed into a net poaching (employer-to-employer/EE) channel and a net nonemployment channel; EE transitions are direct moves with under seven days of nonemployment. Taber and Vejlin (2020) find 80% of EE transitions are voluntary, so poaching flows reveal worker preferences. Cyclical indicators are the change in the unemployment rate (first difference) and the level (HP-filtered deviation from trend).&lt;/p&gt;
&lt;p&gt;Main findings (magnitudes). (1) Productivity is the better job-ladder measure. Residualized wage and TFP are only weakly correlated (Spearman 0.32). Cross-sectionally, the high-vs-low gap in net job creation is far larger for TFP (0.52% vs -0.39%) than for wages (0.26% vs 0.22%), and the net-poaching differential is larger for productivity (0.75%) than wages (0.61%), since workers move up the productivity ladder faster than the wage ladder. (2) Cyclicality differs by ladder. A one-percentage-point rise in the CHANGE in unemployment raises the high-low differential job-creation rate by 0.30 pp for TFP — about 32% of the average TFP differential — driven entirely by the nonemployment channel (0.38 pp), while the poaching channel pulls the opposite way (-0.08 pp). This is a cleansing effect: low-productivity firms both fire more workers to nonemployment AND stop hiring from nonemployment in recessions. For the WAGE ladder the total differential instead contracts by 0.08 pp, because high-wage firms stop poaching (-0.21 pp) — the wage ladder breaks down (a sullying effect). (3) Measurement matters. Using sales per worker instead of TFP yields 0.12 pp on the change-in-unemployment indicator (~40% smaller than TFP&amp;rsquo;s 0.30), and with the LEVEL of unemployment the sign flips: TFP gives +0.11 pp but sales per worker gives -0.08 pp — matching Haltiwanger et al. (2021) on US LEHD data, implying their result reflects sales-per-worker proxying, not a US-Denmark difference.&lt;/p&gt;
&lt;p&gt;Implications. Productivity (not the spot wage) is what workers climb toward, consistent with sequential-auction/outside-option models (Postel-Vinay and Robin 2002). Business-cycle labor models need endogenous hiring rates, since firms shut down hiring rather than only firing in recessions.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-empirical-strategy-for-ranking-firms-and-what-are-the-main-threats-to-it"&gt;Q1. What is the empirical strategy for ranking firms, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Firms are ranked within 2-digit NACE industry-year cells (68 industries) on two dimensions: (a) residualized average hourly wage (regressing firm average wage on workforce tenure, education, age, gender, plus year FE) and (b) TFP from an Olley-Pakes (1996) control-function production function using value added, capital stock, FTE employment, and workforce composition, estimated separately by industry. Quintiles are employment-weighted, so results are interpreted as effects on the average worker. To avoid reclassification bias, firms are ranked on year t-1 measures for flows in year t. Threats: (i) Olley-Pakes uses investment as the productivity proxy but investment data exist only from 1999, so coefficients are estimated post-1999 and back-applied, assuming production technology did not change materially over 1992-2013 — an explicit assumption. (ii) They cannot use Ackerberg-Caves-Frazer (2015) or Levinsohn-Petrin (2003) because detailed intermediate-input data are missing for most firms/years. (iii) AKM firm fixed effects are avoided because the large share of small firms induces limited-mobility bias; residualized average wages are used instead (Haltiwanger et al. 2021 find no difference between AKM FE and average wages). Results are robust to an unresidualized wage measure.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-channels-and-how-are-they-distinguished-empirically"&gt;Q2. What are the two channels and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;Net job creation is decomposed as Net Job Creation = Net Poaching (EE hires minus EE separations) + Net Nonemployment (hires from minus separations to nonemployment). EE/poaching transitions are direct employer changes with fewer than seven days of nonemployment between jobs (threshold varied, results similar). Poaching flows are treated as primarily voluntary (80% per Taber and Vejlin 2020), so they reveal the job ladder; nonemployment flows capture involuntary separations and hiring from the jobless pool. The daily data are essential to cleanly separate EE moves from moves through a nonemployment spell.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-across-firm-types-and-channels-is-documented"&gt;Q3. What heterogeneity across firm types and channels is documented?&lt;/h3&gt;
&lt;p&gt;Cross-section: high-wage firms grow mainly via net poaching (0.21%) plus a little net nonemployment (0.06%); low-wage firms LOSE workers to poaching (-0.40%) but GAIN strongly via nonemployment (0.62%), so they still grow (0.22%). Low-productivity firms also lose via poaching (-0.47%) but, unlike low-wage firms, grow only marginally via nonemployment (0.08%), so they shrink overall (-0.39%). Low-type firms (both rankings) have more churn (higher hires and separations) than high-type firms. Over the cycle (Table 3, change in unemployment): when unemployment rises, low-productivity firms contract more (-1.02 pp) than high (-0.71 pp), driven by the nonemployment margin (-1.05 vs -0.67 pp) and by hiring from nonemployment rather than separations (hiring is more cyclically sensitive, consistent with Shimer 2012). High-wage firms contract more than low-wage firms; for high-wage firms separations to nonemployment rise sharply (0.26 pp vs 0.04 pp for low-wage), consistent with Mueller (2017) and Zullig (2022) that high residual-wage workers are more cyclically sensitive. Low-wage firms net-gain through poaching in recessions (0.08 pp) because poaching separations fall more than poaching hires.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-cyclicality-regression-estimates-in-detail"&gt;Q4. What are the cyclicality regression estimates in detail?&lt;/h3&gt;
&lt;p&gt;Regressions of differential (high-minus-low) flow rates on a cyclical indicator (times 100), with seasonal dummies and a time trend, 82 quarterly observations. Change-in-unemployment, TFP: Total 0.30 pp (SE 0.10, ***), Poaching -0.08 (0.04, *), Nonemployment 0.38 (0.09, ***). Level-of-unemployment, TFP: Total 0.11 (0.05, **), Nonemployment 0.13 (0.04, ***), Poaching -0.02 (ns). Change-in-unemployment, Wage: Total -0.08 (0.06, ns), Poaching -0.21 (0.08, ***), Nonemployment 0.13 (0.06, **). Level-of-unemployment, Wage: Total -0.17 (0.03, ***), Poaching -0.15 (0.03, **&lt;em&gt;), Nonemployment -0.02 (ns). The authors note that a 2-pp rise in unemployment (typical in a recession) raises the TFP differential job-creation rate by ~66% (2&lt;/em&gt;0.30/0.91) of its mean.&lt;/p&gt;
&lt;h3 id="q5-how-robust-are-the-results-to-alternative-measures-and-classifications"&gt;Q5. How robust are the results to alternative measures and classifications?&lt;/h3&gt;
&lt;p&gt;Cross-sectional results are similar across TFP, value added per worker, and sales per worker, and across three high/low cutoffs (baseline top-2/bottom-1 quintiles; Haltiwanger 2021 top-2/bottom-3; Haltiwanger 2015 top-1/bottom-1). TFP consistently yields the largest net-poaching differential, so it is argued superior, though cross-sectional differences are minor. The key DIVERGENCE is in business-cycle estimates: sales per worker underestimates cyclicality (0.12 vs 0.30 pp on change-in-unemployment) and FLIPS sign on the level indicator (-0.08 vs +0.11 pp), a pattern confirmed across all three classifications. Value added per worker and an alternative OLS-based TFP measure both track baseline TFP closely and, crucially, do NOT produce the sign switch on the level indicator — isolating sales per worker as the outlier. Ranking on profits or employment growth (unreported) gives qualitatively similar results to TFP.&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-relate-to-and-differ-from-the-closest-prior-work"&gt;Q6. How does this paper relate to and differ from the closest prior work?&lt;/h3&gt;
&lt;p&gt;Closest empirical work is Haltiwanger et al. (2018a, 2021) on US LEHD data: 2018a concludes firm wage beats firm size as a job-ladder proxy and that high-wage firms are more cyclically sensitive; 2021 finds whether recessions cleanse depends on the cyclical indicator, using sales per worker as a productivity proxy. This paper adds direct TFP (LEHD lacks it), uses daily rather than quarterly data (removing time-aggregation bias, ~30% on EE rates), and shows the wage-ranking results replicate Haltiwanger qualitatively while TFP gives different and stronger conclusions. The wage-vs-sales sign discrepancy is shown to be a measurement artifact, not a US-Denmark institutional difference. Theoretically it is closest to Audoly (2020) and Moscarini and Postel-Vinay (2013), in which better (high-type) firms are more cyclically sensitive because they poach more in expansions when the unemployed pool is small; the paper finds support for this poaching margin using TFP but, being empirical, focuses on which firm characteristic best measures the ladder. It differs from Sorkin (2018), which identifies good firms via revealed preference but does not link them to productivity, and complements Lochner and Schulz (forthcoming) on sorting.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-theoreticalpolicy-implications-and-their-scope-conditions"&gt;Q7. What are the theoretical/policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Recessions speed productivity-enhancing reallocation (cleansing via the nonemployment channel) but impede progression up the wage ladder (sullying via the poaching channel). A central modeling implication: the cleansing effect is driven only PARTLY by the classical Mortensen-Pissarides (1994) channel of firing unproductive workers; equally important, low-productivity firms STOP HIRING from nonemployment in recessions. Models with exogenous arrival rates cannot fit this (more jobs should be created from nonemployment when unemployment is high); endogenous hiring decisions are needed (e.g., Lise and Robin 2017, where low aggregate states shift the vacancy distribution toward high types). Scope conditions: estimates come from Denmark&amp;rsquo;s flexicurity labor market (low firing/hiring regulation, decentralized firm-level wage bargaining, mobility closer to the US than to France/Italy — a Dane is ~2x more likely than a French/Italian worker to make a voluntary EE move, a US worker 2.5x), 1992-2013, manufacturing/services/trade only; means-tested social assistance prevents separating active from inactive nonemployment. Magnitudes are conditional on the chosen productivity measure — using sales per worker would understate or reverse the cleansing finding.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-nature-of-this-record-corrigendum"&gt;Q8. What is the nature of this record (corrigendum)?&lt;/h3&gt;
&lt;p&gt;The DOI 10.1016/j.red.2025.101320 is a corrigendum to the original RED article 101307 (2025). The full-text file provided is the underlying working paper (IZA Discussion Paper No. 15872, January 2023), itself a heavily revised version of an earlier IZA paper, &amp;lsquo;Employment Reallocation over the Business Cycle: Evidence from Danish Data,&amp;rsquo; a chapter of Bertheau&amp;rsquo;s PhD dissertation. The summary reflects the substantive paper content; the corrigendum itself (corrections to the published version) is not detailed in the provided text.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Job ladder&lt;/strong&gt;: A common ranking by workers of available jobs from less to more desirable; the paper tests whether the rung is best indexed by a firm&amp;rsquo;s average wage or its TFP, treating the measure that best predicts voluntary (poaching) moves up as the true ladder.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Net poaching channel&lt;/strong&gt;: Net employer-to-employer (EE) flows — hires poached from other firms minus separations to other firms (direct moves with under seven days of nonemployment). Treated as primarily voluntary (80% per Taber and Vejlin 2020) and thus revealing of the job ladder.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Net nonemployment channel&lt;/strong&gt;: Net flows between a firm and the nonemployment pool — hires from nonemployment minus separations to nonemployment; not distinguished by type of nonemployment because Danish means-tested assistance prevents separating active from inactive jobseekers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cleansing effect&lt;/strong&gt;: In this paper&amp;rsquo;s sense, recessions direct/retain employment in more productive firms: the high-low productivity gap in job creation WIDENS in recessions, as low-productivity firms both separate more workers to nonemployment and stop hiring from it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sullying effect&lt;/strong&gt;: Workers are matched to better firms at a lower rate in bad times: the differential net POACHING rate between high and low firms shrinks in recessions, so the (especially wage) job ladder breaks down and workers get stuck in low-rung firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;TFP (Olley-Pakes control function)&lt;/strong&gt;: Revenue-based total factor productivity estimated via the Olley-Pakes (1996) two-step method, using firm investment as a proxy for unobserved productivity; preferred over labor productivity/sales per worker because it nets out capital intensity and better predicts employment growth and net poaching.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time-aggregation bias&lt;/strong&gt;: The distortion in measured EE transitions when employment is observed only at low (e.g., quarterly) frequency, which conflates EE moves with moves through short nonemployment spells; daily Danish data avoid it (quarterly data overstate EE rates by ~30%, Bertheau and Vejlin 2022).&lt;/p&gt;</description></item><item><title>Diet, Economic Development and Climate Change</title><link>https://macropaperwarehouse.com/papers/diet-economic-development-and-climate-change/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/diet-economic-development-and-climate-change/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Food production accounts for roughly one-third of global greenhouse gas (GHG) emissions, and richer nations contribute disproportionately through meat-intensive diets and input-intensive farming. This paper asks how much of that disparity will be exported to the developing world as it grows, and which policies can most cost-effectively reduce agricultural emissions during that transition. The answer requires separately identifying two distinct channels—demand-side dietary change and supply-side technological change—and tracing their general equilibrium consequences through global food markets.&lt;/p&gt;
&lt;p&gt;The authors build a quantitative multi-country general equilibrium model calibrated to 90 countries (plus a rest-of-world aggregate) and 47 food products for 2010. The demand side features nested non-homothetic CES preferences, which allow income elasticities to differ across food products—the core mechanism of the nutrition transition. The supply side, built on Farrokhi and Pellegrina (2023), operates at a granular grid-cell level covering the Earth&amp;rsquo;s surface, with producers on each plot choosing both which crop to grow and whether to use a modern, input-intensive (higher-GHG) technology or a traditional, labor-intensive one—the core mechanism of agricultural modernization. GHG emissions are tracked from both production and transportation. Data on calorie intake come from FAO Food Balance Sheets; emissions from Poore and Nemecek (2018) and EDGAR-FOOD; yields from FAO-GAEZ (approximately 1.1 million fields).&lt;/p&gt;
&lt;p&gt;A key methodological contribution is an identification result for income elasticities that requires no price data. In open-economy models, trade shares provide a sufficient statistic for consumer prices, so the model&amp;rsquo;s implicit Marshallian demand equations can be estimated using only expenditure shares and bilateral trade flows—a cleaner identification than prior closed-economy approaches. Structural elasticity estimates are validated against reduced-form regressions that regress product-level log absorption on log GDP per capita interacted with the product&amp;rsquo;s GHG intensity; the cross-method correlation has a slope of 0.64–0.77 and R² of 0.93–0.95.&lt;/p&gt;
&lt;p&gt;Four empirical patterns motivate the model. First, diet composition alone drives large variation in emissions: if the whole world adopted the US diet (holding total calories fixed), the food share of global GHG emissions would rise from 30% to 42%; adopting the Argentinian diet would raise it to 74%; adopting the Ethiopian diet would lower it to 12%. Second, GHG emissions per capita from food rise strongly with GDP per capita (elasticity 0.39 in the cross-section); about one-third of this is a pure scale effect (more calories) and two-thirds is a compositional shift toward higher-emission foods (elasticity of emissions per calorie with respect to GDP per capita is 0.23–0.28). Third, products with higher GHG emissions per calorie have higher income elasticities; a 1% rise in a product&amp;rsquo;s GHG intensity is associated with a 0.17–0.21% higher income elasticity, robust to excluding all meat products. Fourth, emissions from fertilizers and energy use as a share of total agricultural emissions rise with GDP per capita (slope 0.82), indicating that agricultural modernization independently amplifies GHG emissions within each crop.&lt;/p&gt;
&lt;p&gt;Model decompositions reveal that about two-thirds of the cross-sectional correlation between food emissions per capita and GDP per capita is attributable to intrinsic dietary preferences (culture, religion, demographics) rather than to income itself, and about one-half of the correlation for emissions per calorie. This implies that the causal effect of economic growth on emissions is substantially smaller than raw correlations suggest.&lt;/p&gt;
&lt;p&gt;Policy counterfactuals (Table 4) are the paper&amp;rsquo;s centerpiece. A uniform 10% TFP shock across all modern agricultural, non-agricultural, and input producers raises global welfare by 14.9% and increases global agricultural GHG emissions by 5.0% (approximately 0.6 Gt CO₂ from production, 0.004 Gt from transport). Shutting down the nutrition transition channel reduces this emission increase by 28%; shutting down agricultural modernization reduces it by a further 16%; shutting both down reduces it by 42%—so the two mechanisms together account for more than one-third of the growth-induced emission increase. Crucially, ignoring general equilibrium supply responses would overstate the emission impact of economic growth by 100%: higher food demand raises production prices, which dampens both consumption growth and further technology adoption.&lt;/p&gt;
&lt;p&gt;For dietary restrictions: a global no-beef mandate would reduce agricultural GHG emissions by 20%, at a global welfare cost of 0.6%, with large concentrated losses in major beef-producing and consuming countries (Argentina −3–5%; Uruguay −4%). A global vegetarian mandate would reduce emissions by 30% (approximately the same 20% figure is given in the abstract with apparent inconsistency but Table 4 column 3 shows −20% for no-beef and −30% for vegetarian), at a welfare cost of 2.8% globally and with greater inequality impacts for developing countries. Back-of-the-envelope calculations that ignore general equilibrium overstate the emission reductions from dietary restrictions by roughly one-third.&lt;/p&gt;
&lt;p&gt;For food trade policy: raising trade costs enough to cut transportation emissions by 75% reduces total agricultural GHG emissions by 11.9%, but at a global welfare cost of 17.8%—a ratio far worse than dietary policies. The welfare loss is highly unequal: countries in the bottom quartile of the GDP per capita distribution face welfare losses of up to 41% (the abstract states this figure; Table 4 col. 2 shows the Q4/Q1 inequality worsening by 4.9 percentage points in the eat-local scenario). The conclusion is that dietary policies dominate food trade policies on both effectiveness and equity grounds.&lt;/p&gt;
&lt;p&gt;Transportation emissions account for only about 5% of agricultural GHG (0.7 Gt CO₂ vs. 16.5 Gt from production), so policies targeting transport emissions alone have limited aggregate impact.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-for-income-elasticities-and-why-is-it-novel"&gt;Q1. What is the core identification strategy for income elasticities, and why is it novel?&lt;/h3&gt;
&lt;p&gt;Standard non-homothetic CES estimation requires price data because the demand equation depends on price indices. In a closed economy this problem is severe. The authors show that in an open economy, bilateral trade shares provide a sufficient statistic for variety price indices: averaging trade shares across a country&amp;rsquo;s import partners yields a geometric mean of production prices that can be differenced out using fixed effects. The key estimating equation (40) regresses an adjusted expenditure share on log income per capita, with fixed effects absorbing production-price variation through the set of import partners. No price data is needed. This is exact—not an approximation—unlike the approximate methods in Comin et al. (2021) or Caron and Fally (2022), which either impose additional assumptions about price variation across consumer groups or require proxies for crop-specific trade costs such as gravity variables.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-threats-to-identification-and-how-are-they-addressed"&gt;Q2. What are the main threats to identification and how are they addressed?&lt;/h3&gt;
&lt;p&gt;The key concern is that income is correlated with prices and preference shifters that also affect food expenditure shares. In the reduced-form regressions (equation 1), country-year and product-year fixed effects control for country-specific factors (including regional technology change) and global product-specific factors (including product-specific technological progress). In the structural estimation (equation 40), the model&amp;rsquo;s functional form is used to control fully for endogeneity arising through prices, since trade shares substitute out unobservable price indices exactly. The close agreement between reduced-form and structural income elasticity estimates (slope 0.64–0.77, R² 0.93–0.95 in cross-validation) is reassuring that the two quite different identifying assumptions yield similar results. One remaining concern is unobservable preference shifters (ai,k and ã_i,s), which appear as residuals; identification requires income variation orthogonal to these shifters, and the authors follow the precedent of assuming fixed effects are sufficient. Household-level data from Brazil&amp;rsquo;s Consumer Expenditure Survey (POF) bolster the reduced-form patterns using within-country income variation.&lt;/p&gt;
&lt;h3 id="q3-how-are-the-nutrition-transition-and-agricultural-modernization-distinguished-empirically-and-in-the-model"&gt;Q3. How are the nutrition transition and agricultural modernization distinguished empirically and in the model?&lt;/h3&gt;
&lt;p&gt;These are fundamentally different economic mechanisms. The nutrition transition operates through demand: as incomes rise, consumers shift toward food products that, for reasons of taste or nutrition, happen to have higher GHG emissions per calorie. It is a between-product phenomenon captured by non-homothetic income elasticities. Agricultural modernization operates through supply: as wages rise, producers substitute away from labor-intensive traditional technologies toward input-intensive modern technologies (fertilizers, machinery) that emit more GHG per calorie of output, for any given crop. It is a within-product phenomenon captured by the endogenous technology-choice margin in the agricultural production model. In the counterfactual decompositions, the authors shut down each channel independently: the nutrition transition is shut down by setting all within-sector income elasticity parameters (ε_k) equal; agricultural modernization is shut down by fixing the land share in each technology exogenously. Doing so reveals that the nutrition transition accounts for 28% and modernization for 16% of the emission increase from a 10% TFP shock (jointly 42%), with the remainder attributable to scale effects and general equilibrium price responses.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-role-of-general-equilibrium-supply-responses-and-why-do-they-matter-so-much"&gt;Q4. What is the role of general equilibrium supply responses and why do they matter so much?&lt;/h3&gt;
&lt;p&gt;A central finding is that ignoring supply-side equilibrium price responses would overstate the emission impact of economic growth by 100%. The mechanism is straightforward: economic growth raises income and thus food demand, which pushes up production prices (because agricultural supply is upward-sloping due to limited land and heterogeneous productivity across grid cells). Higher prices dampen consumption, which partially offsets the demand-driven emission increase. For dietary restriction policies, back-of-the-envelope calculations that simply remove the GHG attributable to banned food products overstate the emission reduction by roughly one-third, because consumers substitute toward other food products and global agricultural production reorganizes. The model&amp;rsquo;s general equilibrium structure is therefore essential for obtaining credible policy counterfactuals, and a main conclusion of the paper is that the literature&amp;rsquo;s existing back-of-the-envelope calculations in environmental science substantially overstate both the emission risks from growth and the emission benefits from dietary policies.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented-across-countries-and-products"&gt;Q5. What heterogeneity is documented across countries and products?&lt;/h3&gt;
&lt;p&gt;Across countries: diet composition varies enormously. Counterfactual calculations show that if all countries adopted the Argentinian diet (holding total calories fixed), the global food share of total emissions would rise to 74%; adopting the Ethiopian diet would lower it to 12%, compared to the factual 30%. The income elasticity of the agricultural sector as a whole is 0.39, close to Comin et al. (2021)&amp;rsquo;s 0.37. Rich countries have a higher share of modern technology in production, higher fertilizer and energy use per unit of land, higher food GHG per capita, and higher food GHG per calorie. About two-thirds of the cross-sectional gradient in food GHG per capita is attributable to intrinsic preferences rather than income per se. Religion is documented as one driver: Islamic-majority countries show lower preference for pork; Hindu-majority countries show higher preference for lamb, mutton, and poultry relative to other meats. Across products: GHG emissions per 1,000 kcal range from above 35 kg CO₂ for beef and coffee to below 5 kg CO₂ for wheat and rye. Income elasticity parameters (ε_k) range from lowest for staples (yams, sweet potatoes, millet, sorghum, rice) to highest for luxury fruits and vegetables (berries, asparagus, cucumbers, watermelon). Notably, the income-GHG gradient persists after excluding all meat products: vegetables and fruits have higher GHG per calorie than staples, so the nutrition transition is broader than a simple meat-consumption story.&lt;/p&gt;
&lt;h3 id="q6-how-do-the-diet-restriction-and-food-trade-policy-counterfactuals-compare-on-welfare-and-effectiveness"&gt;Q6. How do the diet restriction and food trade policy counterfactuals compare on welfare and effectiveness?&lt;/h3&gt;
&lt;p&gt;Diet restriction (no-beef): global GHG emissions fall 20%, global welfare falls 0.6%. The welfare effect is highly concentrated—Argentina experiences −3–5% welfare loss, Uruguay approximately −4% in the no-beef scenario, because they are large meat producers and exporters. Inequality between rich (Q4) and poor (Q1) countries worsens by 1.0 percentage point. Diet restriction (vegetarian): global GHG emissions fall 30%, global welfare falls 2.8%. Inequality worsens by 6.0 percentage points, indicating developing countries bear more of the cost because a larger share of their income goes to food, and their income sources (agriculture) are more directly affected. Food trade policy (&amp;rsquo;eat local&amp;rsquo;, raising trade costs to cut transportation emissions by 75%): global GHG emissions fall 11.9%, but global welfare falls 17.8%—roughly 25–30 times the welfare cost per percentage point of emission reduction compared to dietary policies. Inequality worsens substantially more: Q4/Q1 ratio worsens by 4.9 percentage points. Countries in the bottom GDP quartile face welfare losses up to 41%. The paper concludes that dietary restrictions are both substantially more effective in reducing GHG emissions and far more equitable in their welfare consequences than food trade policies.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-share-of-agricultural-ghg-from-transportation-versus-production-and-what-are-the-implications"&gt;Q7. What is the share of agricultural GHG from transportation versus production, and what are the implications?&lt;/h3&gt;
&lt;p&gt;In the 2010 data, GHG emissions from food transportation account for approximately 5% of total agricultural GHG (0.7 Gt CO₂ out of approximately 17.2 Gt total). Production accounts for 95% (16.5 Gt CO₂). This has two implications. First, in the economic growth counterfactual, transportation emissions increase by 2.2%, but because transportation is only 5% of total, its contribution to total emission growth (0.004 Gt) is negligible. Second, it implies that policies targeting food &amp;lsquo;food miles&amp;rsquo; or local eating are poorly targeted: even a dramatic 75% reduction in transportation emissions only mechanically eliminates 4.6% of total agricultural GHG, and the actual general equilibrium reduction (11.9%) comes mostly from production effects (agricultural trade restrictions reduce global production and consumption), accompanied by very large welfare costs.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-and-validation-exercises-are-conducted"&gt;Q8. What robustness checks and validation exercises are conducted?&lt;/h3&gt;
&lt;p&gt;The paper provides several validation exercises. (1) The reduced-form income elasticity regressions are run both with all crops and excluding all meat products (beef, lamb and mutton, pig meat, poultry), yielding nearly identical coefficients of 0.176 and 0.175 (columns 1 and 2 of Table 1), and with country-year and product-year fixed effects (columns 3–4), showing similar results across specifications. (2) The structural income elasticities are compared to the reduced-form estimates, with a cross-method slope of 0.64–0.77 and R² of 0.93–0.95, reassuring given the two methods make different identifying assumptions. (3) Model fit is checked against six untargeted empirical regularities (Figure 6): declining agricultural employment share, rising input cost share, rising modern technology land share, rising food GHG per capita, rising calories per capita, and rising food GHG per calorie—all with GDP per capita. The model matches the sign and approximate magnitude of each relationship. (4) Household-level estimates using Brazil&amp;rsquo;s POF survey replicate the cross-country finding that higher-GHG products have higher income elasticities, controlling for fixed effects, food price proxies, and excluding meat. (5) The decomposition of the cross-sectional income-emissions gradient shows that equalizing comparative advantage (column 3) or trade costs (column 4) across countries leaves the gradient approximately unchanged, supporting the focus on preferences and technology.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-prior-work-and-where-does-it-depart-from-it"&gt;Q9. How does this paper relate to prior work and where does it depart from it?&lt;/h3&gt;
&lt;p&gt;The paper sits at the intersection of several literatures. It builds on Farrokhi and Pellegrina (2023) for the granular grid-cell production model with technology choice; on Costinot, Donaldson, and Smith (2016) for the agricultural field structure; and on Comin, Lashkari, and Mestieri (2021) for non-homothetic CES preferences and the identification of income elasticities. Key departures: (a) Relative to Comin et al. (2021), the authors extend identification to nested CES preferences and to an open-economy without requiring price data—their method is exact rather than approximate. (b) Relative to the environmental science literature (e.g., Hoolohan et al., 2013; Perignon et al., 2017; Tilman et al., 2011), the paper endogenizes general equilibrium supply responses, which the authors show dramatically attenuate the effect of both income growth and dietary policies on emissions. (c) Relative to prior quantitative spatial models of climate change (e.g., Shapiro 2016 on trade costs and CO₂), this paper focuses on agricultural emissions specifically and introduces nutrition transition and technology choice. (d) The authors claim to be the first to analyze both dietary restrictions and food trade policies on agricultural emissions within quantitative trade models. (e) Relative to Chen et al. (2022), who use a computable general equilibrium model with general equilibrium supply adjustments, this paper includes far more food products (47 vs. their smaller set) and endogenizes technology choice, both of which are quantitatively important for capturing the nutrition transition.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-papers-mechanism-for-why-vegetable-and-fruit-consumption-also-raises-ghg-emissions-as-income-rises-even-without-meat"&gt;Q10. What is the paper&amp;rsquo;s mechanism for why vegetable and fruit consumption also raises GHG emissions as income rises, even without meat?&lt;/h3&gt;
&lt;p&gt;The paper notes in footnote 1 that the positive correlation between income elasticities and GHG emissions per calorie persists even when meat products are excluded from the sample (Table 1, columns 3–4). The reason is that vegetables and fruits—which become more preferred as countries grow richer—emit more GHG per calorie than staple foods such as yams and potatoes. Staples require little processing or refrigeration and are typically produced with traditional, low-input technologies. By contrast, fresh fruits and vegetables (especially high-value items such as berries, asparagus, grapes, and coffee) require more energy-intensive transportation, storage, and sometimes greenhouse production. This means that the nutrition transition generates rising emissions not merely through the beef channel emphasized in much of the public debate, but through a broader shift away from calorie-dense staples toward diverse, lower-calorie-density products that happen to have higher GHG footprints per calorie.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-model-imply-about-the-environmental-kuznets-curve-for-food-emissions"&gt;Q11. What does the model imply about the Environmental Kuznets Curve for food emissions?&lt;/h3&gt;
&lt;p&gt;The paper explicitly tests for and finds no evidence of an Environmental Kuznets Curve (EKC) in food emissions—that is, no inverse-U shape in which emissions per capita eventually decline as countries become very rich, as might be expected if wealthy nations adopt more sustainable diets or stricter environmental regulations. The income-emission relationship is found to be approximately log-linear across all levels of development (footnote 8). This is consistent with the broader empirical literature on the EKC (cited survey by Dinda, 2004). The implication is that there is no automatic &amp;lsquo;greening&amp;rsquo; of diets as countries develop; active policy intervention would be needed.&lt;/p&gt;
&lt;h3 id="q12-how-is-economic-development-modeled-in-the-policy-counterfactuals-and-what-are-the-scope-conditions"&gt;Q12. How is economic development modeled in the policy counterfactuals, and what are the scope conditions?&lt;/h3&gt;
&lt;p&gt;Economic development is modeled as a uniform 10% increase in TFP for three types of agents: (i) modern agricultural producers, (ii) non-agricultural producers, and (iii) agricultural input producers (fertilizers, machinery, pesticides). Traditional agricultural technology is not subject to productivity growth, following Gollin, Parente, and Rogerson (2007). This creates both income effects (via higher wages) and substitution effects (via changes in relative input prices that favor modern, input-intensive technology). The scope conditions are important: the results apply specifically to a uniform global TFP shock, not to individual-country development. For individual-country TFP shocks, the analytical decomposition (equation 34) shows that general equilibrium income spillovers to foreign countries can attenuate the nutrition transition if foreign incomes fall (e.g., due to terms-of-trade effects). The model does not incorporate dynamics (it is a static model calibrated to 2010), so it cannot directly speak to transition paths or time horizons for emission convergence.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-welfare-implications-for-developing-countries-under-different-policies-and-why-do-dietary-policies-dominate"&gt;Q13. What are the welfare implications for developing countries under different policies, and why do dietary policies dominate?&lt;/h3&gt;
&lt;p&gt;Under economic growth (10% TFP shock), global welfare rises 14.9% with a modest increase in Q4/Q1 inequality of 0.4 percentage points, indicating relatively even welfare gains. Under no-beef, global welfare falls 0.6% but inequality worsens by 1.0 pp; under vegetarian, welfare falls 2.8% and inequality worsens by 6.0 pp—developing countries lose more because more of their income is spent on food and the agricultural sector is a larger share of their economy. Under eat-local (food trade restrictions), welfare falls 17.8% and the Q4/Q1 ratio worsens by 4.9 pp, with countries in the bottom GDP quartile facing losses up to 41%. The stark dominance of dietary policies over trade policies reflects two structural features: (a) food trade restrictions reduce the gains from comparative advantage in food production, which are particularly large for food-exporting developing countries; and (b) the welfare cost per unit of GHG reduction is far higher for trade policies because they distort production allocation without addressing the underlying demand-side emissions driver.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Nutrition Transition&lt;/strong&gt;: As defined and used in this paper: the demand-side process by which rising income causes consumers to shift their caloric intake away from staple foods (yams, potatoes, rice, millet) toward food products with higher GHG emissions per calorie (meats, fruits, vegetables, coffee). The transition is captured in the model by non-homothetic income elasticity parameters ε_k that are higher for more emissions-intensive products and is operative even after excluding all meat products.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agricultural Modernization&lt;/strong&gt;: As defined and used in this paper: the supply-side process by which rising wages induce producers to substitute from traditional, labor-intensive agricultural technology (τ=0, no purchased intermediate inputs) toward modern, input-intensive technology (τ=1, fertilizers, machinery, pesticides), which emits more GHG per calorie of output. This operates within each crop and is captured in the model by endogenous technology choice at the plot level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-Homothetic CES Preferences (Nested)&lt;/strong&gt;: A three-tier preference structure in which the expenditure share of a food product k depends on income through a product-specific parameter ε_k that governs how fast the product&amp;rsquo;s preference weight grows with utility. Products with higher ε_k have higher income elasticities; the overall income elasticity of the agricultural sector (0.39 in this paper&amp;rsquo;s calibration) is an expenditure-weighted average of the ε_k values. The nested structure allows the agricultural sector&amp;rsquo;s income elasticity relative to non-agriculture to be determined separately from the income elasticities of individual food products within agriculture.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implicit Marshallian Demand&lt;/strong&gt;: The demand equation derived from non-homothetic CES preferences by substituting out unobservable price indices using a base good, yielding a demand specification that depends on observable expenditure shares and income rather than on prices directly. In this paper&amp;rsquo;s open-economy extension, trade shares further substitute out unobservable variety price indices, making the estimation equation fully price-data-free.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GHG Emission Intensity (per calorie)&lt;/strong&gt;: In this paper: the parameter φ_k (crop-specific) and φ_τ (technology-specific), where φ_kτ = φ_k × φ_τ is the kg CO₂-equivalent emitted per 1,000 kcal of crop k produced under technology τ. This is the key cross-product heterogeneity that, combined with income elasticity heterogeneity, drives the environmental consequences of the nutrition transition. In the data: ranges from below 5 kg CO₂ per 1,000 kcal for wheat and rye to above 35 kg for beef and coffee.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Grid-Cell Production Model&lt;/strong&gt;: A representation of the agricultural supply side in which the Earth&amp;rsquo;s land surface is divided into approximately 1.1 million fields (FAO-GAEZ), each with agro-climatically determined potential yields by crop and technology that are independent of market conditions. Within each field, a continuum of plots is allocated to crops and technologies via Fréchet productivity draws, yielding smooth aggregate supply functions and allowing for realistic specialization patterns and technology gradients across geography.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Back-of-the-Envelope (Demand Mechanism) Benchmark&lt;/strong&gt;: In this paper: a partial-equilibrium counterfactual calculation that takes observed or baseline food demand quantities and simply attributes changes to them from a policy without allowing supply prices, production, or trade flows to adjust. The paper systematically compares model general equilibrium results against this benchmark (column 9 of Table 4) to quantify how much supply-side adjustments matter, finding that the back-of-the-envelope approach overstates the emission impact of economic growth by approximately three times, and overstates the emission reduction from dietary policies by roughly one-third.&lt;/p&gt;</description></item><item><title>Dispersion Over the Business Cycle: Passthrough, Productivity, and Demand</title><link>https://macropaperwarehouse.com/papers/dispersion-over-the-business-cycle-passthrough-productivity-and-demand/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/dispersion-over-the-business-cycle-passthrough-productivity-and-demand/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Carlsson, Clymo, and Joslin use Swedish manufacturing firm-level microdata for 1998–2013 to separately identify and characterize the cyclical behavior of physical productivity (TFPQ) shocks and demand shocks at the firm level, two forces that are observationally equivalent under the standard CES-demand benchmark. The paper&amp;rsquo;s central contribution is threefold: it documents new empirical facts about dispersion cyclicality, estimates a non-constant-elasticity (non-CES) demand curve directly from firm-level price and quantity data, and embeds those estimates into a quantitative heterogeneous-firm model to study the aggregate consequences of each type of dispersion shock.&lt;/p&gt;
&lt;p&gt;The data combine four Swedish register sources: the Företagens Ekonomi (FEK) survey for bookkeeping variables; the Industrins Varuproduktion (IVP) survey for 8-digit product-level price and quantity data used to construct firm-level price indices; the Konjunkturstatistik för Industrin (KFI) survey for quarterly capacity-utilization data; and additional investment deflators. The unbalanced panel contains 3,181 unique manufacturing firms and 15,044 firm-year observations. TFPQ is measured using a Cobb-Douglas value-added production function with factor utilization adjustment; factor elasticities are estimated via cost shares at the 2-digit sector level, yielding an average labor share of 0.735.&lt;/p&gt;
&lt;p&gt;Demand is estimated using the Gopinath-Itskhoki-Rigobon (GIR) flexible demand curve, which nests CES as the limiting case. TFPQ innovations instrument for price in a second-order approximation, following Foster, Haltiwanger, and Syverson (2008). The main-sample estimates yield theta = 2.94 (average elasticity) and eta = 4.27 (super-elasticity), both significant at the 1% level. The second-order price term is statistically significant at the 5% level in all three samples, decisively rejecting CES. These estimates imply that a 5% price increase raises the demand elasticity from 2.94 to 3.74, while a 5% price reduction reduces it to 2.42, creating a &amp;ldquo;real rigidity&amp;rdquo; in the sense of Ball and Romer (1990): raising price loses many customers while lowering it gains few.&lt;/p&gt;
&lt;p&gt;Incomplete passthrough of TFPQ shocks is a central empirical finding. OLS estimates yield beta_z = -0.124; first-difference estimates yield -0.097. Even in the subsample of firms that adjusted all product-level prices in a given year, TFPQ passthrough remains near -0.10, ruling out Calvo or menu-cost price stickiness as the sole driver. Longer-horizon (two- and three-year) first-difference regressions produce similar estimates, ruling out Rotemberg gradual adjustment as well. The non-CES demand curve alone implies a static-optimal passthrough of theta/(theta + eta) = 3/(3 + 4.3) = 41%, so real rigidity explains most of the incompleteness even before accounting for adjustment costs. Demand shocks pass through to prices at a rate of 0.209-0.235, a non-zero result rationalized in the quantitative model by input adjustment costs.&lt;/p&gt;
&lt;p&gt;On cyclicality of dispersion, both TFPQ and demand shock dispersion are countercyclical, but demand dispersion rises by more and is more robust across recession episodes. In 2009 (the Great Recession), the IQR of demand shock growth was 56% above its non-recession average, while the IQR of TFPQ shock growth rose 36%. Sales dispersion rose 58% (IQR) in 2009. A semi-structural variance decomposition shows that demand shocks account for 63% of average sales growth dispersion and approximately 80% of its increase in 2009; TFPQ dispersion contributes only marginally to sales dispersion because the TFPQ variance is shrunk by a factor of roughly 25 on its way to sales growth through the chain of low passthrough and demand elasticity. Demand accounts for about 50% of average price growth dispersion and 40% of its cyclical increase in 2009; TFPQ accounts for about 10% of price dispersion on average.&lt;/p&gt;
&lt;p&gt;The quantitative heterogeneous-firm model extends Bloom (2009) and Bloom et al. (2018) to continuous time with both TFPQ and demand shocks, non-CES demand (theta = 3, eta = 4.3 from the estimates), and non-convex input adjustment costs on a composite scale factor covering both capital and labor. The resale loss kappa = 0.3565 is taken from Bloom et al. (2018). The model is calibrated to match IQRs of 0.2 for TFPQ and demand shock log-changes in the low-uncertainty state, consistent with pre-crisis Swedish data. For the high-uncertainty state, the calibration targets the Great Recession peaks: a 30% rise in TFPQ dispersion (sigma_z(2) = 1.38 sigma_z(1)) and a 60% rise in demand dispersion (sigma_epsilon(2) = 1.90 sigma_epsilon(1)), reflecting the empirical finding that demand dispersion increases more.&lt;/p&gt;
&lt;p&gt;A simulated transition to the high-uncertainty state causes aggregate output to fall by 3.5%. Decomposing into the Bloom (2009) &amp;ldquo;volatility effect&amp;rdquo; (realized shocks drawn from the high-dispersion distribution, firms believe low) and &amp;ldquo;uncertainty effect&amp;rdquo; (firms believe high, shocks drawn from low distribution), the paper finds both effects are negative in the non-CES model, in sharp contrast to Bloom (2009) where the volatility effect is positive (the Oi-Hartman-Abel effect). Non-CES demand amplifies the total output decline by approximately 40% relative to the CES model (peak fall 2.5% vs. 1.75%), primarily by reversing the sign of the volatility effect. Increased demand dispersion drives almost all of the first-year output decline and the majority of the uncertainty effect; TFPQ dispersion is the main driver of the negative volatility effect via markup dispersion. The inaction rate among firms jumps from 50% to 95% on impact of the uncertainty shock, then recovers within one year. TFPQ uncertainty induces little wait-and-see behavior because firms optimally adjust inputs by only 23% of the TFPQ shock size (versus 200% under CES), so uncertainty about TFPQ translates mainly into markup uncertainty. Demand uncertainty triggers strong wait-and-see behavior because demand directly maps one-for-one into desired input use.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-core-identification-strategy-for-separating-tfpq-and-demand-shocks-and-what-are-the-main-threats"&gt;Q1. What is the paper&amp;rsquo;s core identification strategy for separating TFPQ and demand shocks, and what are the main threats?&lt;/h3&gt;
&lt;p&gt;The authors identify TFPQ from a utilization-adjusted Cobb-Douglas value-added production function, then estimate demand using TFPQ innovations as instruments for price. TFPQ innovations are valid instruments because they shift marginal cost without directly shifting demand, tracing out the demand curve. The utilization adjustment (from the KFI managerial survey) is critical: without it, demand shocks that reduce utilization would appear as negative TFPQ shocks, biasing demand elasticity estimates upward and breaking instrument validity. The paper validates the adjustment by showing that firms reporting &amp;lsquo;insufficient demand&amp;rsquo; exhibit 15% lower utilization on average, and 23% lower during the Great Recession. A second threat is quality change in firm-level prices; the authors address this with (a) robustness using the Eslava et al. (2023) CUPI quality-adjusted price index and (b) a single-product-firm subsample. Demand and passthrough results are similar across all three price index approaches. The within-firm focus (demeaning by firm and sector-year fixed effects throughout) mitigates cross-sectional comparability issues but limits misallocation-level analyses analogous to Hsieh and Klenow (2009).&lt;/p&gt;
&lt;h3 id="q2-how-is-the-non-ces-demand-curve-identified-and-what-exactly-does-the-super-elasticity-parameter-eta-measure"&gt;Q2. How is the non-CES demand curve identified, and what exactly does the super-elasticity parameter eta measure?&lt;/h3&gt;
&lt;p&gt;The GIR demand curve is q = (1 - eta * log p)^(theta/eta). A second-order approximation around the firm&amp;rsquo;s average price yields log q = -theta * p_hat - (eta&lt;em&gt;theta/2) * p_hat^2 + fixed effects + epsilon, where p_hat is the firm&amp;rsquo;s demeaned log relative price. Regressing real sales on p_hat and p_hat^2, instrumented by demeaned TFPQ and its square, recovers theta = -b1 and eta = 2&lt;/em&gt;b2/b1. Because p_hat is demeaned at the firm level, the estimates capture within-firm nonlinearity in the price-sales relationship, not cross-sectional heterogeneity in elasticity levels. The parameter eta is the &amp;lsquo;super-elasticity&amp;rsquo;: it measures how much the demand elasticity itself changes with the price. When eta &amp;gt; 0, a firm that raises its price faces an increasingly elastic demand curve (loses customers rapidly), and one that lowers its price faces a less elastic curve (gains customers slowly). The estimated eta = 4.27 in the main sample is roughly half the value of 10 studied (but not estimated) in Klenow and Willis (2016) and larger than the approximately 2 used in Berger and Vavra (2019).&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-distinguish-the-volatility-effect-from-the-uncertainty-effect-in-the-quantitative-model"&gt;Q3. How does the paper distinguish the &amp;lsquo;volatility effect&amp;rsquo; from the &amp;lsquo;uncertainty effect&amp;rsquo; in the quantitative model?&lt;/h3&gt;
&lt;p&gt;Following Bloom (2009), the paper simulates two counterfactuals. The uncertainty effect holds shocks drawn from the low-dispersion distribution (s=1) but lets firms believe that the high-uncertainty state (s=2) has arrived; this isolates the precautionary wait-and-see channel. The volatility effect draws shocks from the high-dispersion distribution (s=2) but lets firms believe they are in the low-uncertainty state; this isolates the direct effect of realizing more extreme shocks on aggregate output. In the non-CES model, both effects are negative. The uncertainty effect is dominated by demand uncertainty because demand shocks directly affect desired input use one-for-one, so uncertainty about future demand creates strong incentives to pause investment. TFPQ uncertainty induces little wait-and-see behavior because the optimal scale adjustment to a TFPQ shock is only 23% of the shock magnitude (vs. 200% under CES). The volatility effect is dominated by TFPQ dispersion because realized TFPQ shocks generate markup dispersion via incomplete passthrough, creating misallocation. Under CES, the volatility effect from TFPQ is positive (OHA effect: convex output-productivity relationship); non-CES demand makes the output-productivity relationship concave for eta large enough, flipping the sign.&lt;/p&gt;
&lt;h3 id="q4-what-mechanism-makes-tfpq-passthrough-so-low-in-both-the-data-and-the-model"&gt;Q4. What mechanism makes TFPQ passthrough so low in both the data and the model?&lt;/h3&gt;
&lt;p&gt;Two mechanisms operate. First, non-CES demand itself: when eta &amp;gt; 0, raising price increases the demand elasticity, and lowering price decreases it. This means the benefit to revenue from a price cut (following a productivity gain that reduces costs) is muted because the firm gains fewer customers than under CES. The static optimal passthrough is theta/(theta + eta) = 3/(7.3) = 41%. Second, non-convex input adjustment costs further reduce passthrough by making firms reluctant to change their scale in response to TFPQ shocks. In the model, the investment threshold is nearly flat across a wide range of TFPQ values (shown in Figure 6, left panel), reflecting that optimal scale barely responds to productivity. Together these mechanisms reproduce TFPQ passthrough of 20-30% in model-simulated data vs. 10-24% in the actual data, both far below the CES benchmark of 100%. The paper also verifies that low passthrough persists in the subsample of flexible-price firm-years, ruling out sticky prices as the primary driver.&lt;/p&gt;
&lt;h3 id="q5-why-does-demand-shock-dispersion-rather-than-tfpq-dispersion-dominate-the-variance-decompositions-of-sales-and-price-growth"&gt;Q5. Why does demand shock dispersion, rather than TFPQ dispersion, dominate the variance decompositions of sales and price growth?&lt;/h3&gt;
&lt;p&gt;The contribution of TFPQ dispersion to sales dispersion is (1-theta)^2 * beta_z^2 * Var(z). With beta_z = -0.097 and theta = 2.99, the TFPQ variance is shrunk by approximately (1-2.99)^2 * (0.097)^2 = 4 * 0.0094 ≈ 0.04, so only about 4% of TFPQ variance propagates to sales variance. This extremely small multiplier reflects two successive attenuation steps: low TFPQ passthrough to prices (beta_z^2 ≈ 0.01) and a small price-to-sales elasticity. Demand shocks, by contrast, affect sales directly through the demand curve without a price intermediary: the contribution is ((1-theta)*beta_epsilon + 1)^2 * Var(epsilon). With beta_epsilon = 0.209 and theta = 2.99, the multiplier is ((1-2.99)*0.209 + 1)^2 = (1 - 0.416)^2 = 0.34, about eight times larger than for TFPQ even though both shocks have similar variance. The cyclical increase is even more skewed toward demand because demand dispersion rises by 56% vs. 36% for TFPQ in 2009.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-relate-to-tfpr-dispersion-and-what-does-it-say-about-using-tfpr-as-a-sufficient-statistic"&gt;Q6. How does the paper relate to TFPR dispersion, and what does it say about using TFPR as a sufficient statistic?&lt;/h3&gt;
&lt;p&gt;TFPR = p * z. For arbitrary passthrough, TFPR growth = beta_epsilon * delta_epsilon + (beta_z + 1) * delta_z. Because passthrough from both shocks is incomplete, TFPR growth reflects a mixture of both underlying shocks. The paper shows via a variance decomposition of TFPR that TFPQ is the main driver of TFPR growth dispersion—accounting for roughly 60% on average—because low passthrough means prices move little, leaving TFPQ changes to dominate TFPR. However, this finding obscures the importance of demand shocks for aggregate outcomes: demand dispersion is the dominant driver of sales growth dispersion and wait-and-see behavior, yet TFPR growth dispersion mostly reflects TFPQ. A researcher relying on TFPR dispersion to infer uncertainty would correctly detect productivity uncertainty but would miss the more cyclically important demand uncertainty channel.&lt;/p&gt;
&lt;h3 id="q7-how-do-the-oi-hartman-abel-oha-and-wait-and-see-mechanisms-work-differently-under-non-ces-vs-ces-demand"&gt;Q7. How do the Oi-Hartman-Abel (OHA) and wait-and-see mechanisms work differently under non-CES vs. CES demand?&lt;/h3&gt;
&lt;p&gt;Under CES demand, sales of each firm are s = z^(theta-1) * exp(epsilon), and aggregate output is E[z^(theta-1)] which is convex in z, so a mean-preserving spread in TFPQ raises aggregate output (OHA effect). Under the estimated non-CES parameters (theta=3, eta=4.3), the approximate relationship yields output proportional to z^0.82, which is concave, so a mean-preserving spread in TFPQ reduces aggregate output. The mechanism is that under non-CES demand, TFPQ shocks pass through incompletely to prices and thus create markup dispersion: high-productivity firms have high markups, low-productivity firms have low markups, and the resulting misallocation reduces total output even relative to a social planner who would set p=mc. For wait-and-see: under CES, optimal input adjustment to a TFPQ shock equals (theta-1) times the shock, which is 200% for theta=3; under non-CES with eta=4.3, it is only (theta^2/(theta+eta) - 1) * shock = 0.233 * shock = 23%. This means firms adjust scale very little in response to TFPQ uncertainty, dampening the wait-and-see channel for TFPQ. TFPQ uncertainty then causes uncertainty about markups, which is costly but does not trigger large investment adjustments.&lt;/p&gt;
&lt;h3 id="q8-what-role-do-adjustment-costs-play-and-how-robust-are-the-results-to-the-structure-of-those-costs"&gt;Q8. What role do adjustment costs play, and how robust are the results to the structure of those costs?&lt;/h3&gt;
&lt;p&gt;Non-convex adjustment costs on a composite firm-scale factor x = k^alpha * l^(1-alpha) create an inaction region: firms neither invest nor disinvest until shocks are sufficiently large. In the low-uncertainty state, the model generates a yearly inaction rate of 25.4% (consistent with pre-crisis Swedish data showing roughly 15%). When uncertainty rises, the inaction region widens, the inaction rate jumps to 95% on impact, and firms let their scale shrink via depreciation. The baseline calibration uses the resale loss kappa = 0.3565 from Bloom et al. (2018). The paper also calibrates kappa to the Swedish inaction rate (kappa = 0.1165), which delivers qualitatively identical dynamics but a smaller amplitude recession (1.7pp vs. 3.5pp output fall). The paper also solves a version with adjustment costs only on capital (as in Bachmann and Bayer, 2013): the wait-and-see effect is dampened but the qualitative results hold—demand uncertainty still dominates TFPQ uncertainty in driving wait-and-see, and non-CES demand still reverses the sign of the OHA effect.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-role-of-the-price-wedge-and-time-varying-passthrough"&gt;Q9. What is the role of the price wedge and time-varying passthrough?&lt;/h3&gt;
&lt;p&gt;The passthrough equation residual (price wedge, tau) captures price changes unexplained by TFPQ and demand shocks. It could reflect un-modeled shocks (e.g., financial constraints, as Gilchrist et al. (2017) document for Sweden), markup decisions, or measurement error. The price wedge makes a meaningful contribution to both average sales/price dispersion and to the rise in 2009. Time-varying passthrough is also documented: TFPQ passthrough is countercyclical (more negative in recessions), while demand passthrough is procyclical (falls in recessions when firms receive more extreme idiosyncratic demand shocks). Redoing the variance decomposition with year-by-year passthrough estimates makes demand&amp;rsquo;s contribution to sales dispersion in 2009 even larger, because firms adjust prices less to demand shocks during the recession, leaving more of the demand shock impact in sales.&lt;/p&gt;
&lt;h3 id="q10-what-heterogeneity-is-documented-across-industries-and-firm-types"&gt;Q10. What heterogeneity is documented across industries and firm types?&lt;/h3&gt;
&lt;p&gt;Sectoral demand elasticity estimates from the pooled 22-sector sample yield an average theta of 3.89 and median of 2.73 for the linear CES model; for the non-linear model, average theta is 3.26 and average eta is 7.42, with substantial positive skew. The median non-linear eta of 5.37 is larger than the pooled estimate of 4.27, indicating the pooled estimate is pulled down by some sectors with smaller deviations from CES. Key empirical results (greater cyclicality of demand dispersion, incomplete TFPQ passthrough) hold within each major sector and across balanced panels, the single-product subsample, and the CUPI price-index sample. Time-varying passthrough is also found to be systematically higher by about 25% in the post-2008 period compared to the pre-2008 period, suggesting a structural shift in how demand shocks transmit to prices, though the paper does not investigate the source of this change.&lt;/p&gt;
&lt;h3 id="q11-what-robustness-checks-are-run-on-the-demand-and-passthrough-estimates"&gt;Q11. What robustness checks are run on the demand and passthrough estimates?&lt;/h3&gt;
&lt;p&gt;Demand estimation robustness: (1) piece-wise linear specification (elasticity of 2 below average price, 4 above average price, significant at 0.1% level); (2) balanced panel; (3) excluding the Great Recession; (4) using Statistics Sweden firm identifiers instead of authors&amp;rsquo; own; (5) CUPI price index; (6) single-product firms; (7) sector-by-sector estimation; (8) including firm and sector-year fixed effects directly in the nonlinear regression (rather than pre-demeaning). All exercises confirm statistically significant eta and broadly similar theta. Passthrough robustness: (1) OLS vs. IV (lagged shocks) vs. first-differences; (2) balanced panel; (3) single-product subsample; (4) two-period lagged instruments (beta_z = -0.294, beta_epsilon = 0.249); (5) flexible-price subsample; (6) longer-horizon (two- and three-year) first differences for TFPQ. Corroboration: TFPQ innovations are positively associated with reported process innovations in Eurostat CIS data (7% greater TFPQ growth for process innovators); negative demand shocks are correlated with managers reporting &amp;lsquo;insufficient demand&amp;rsquo; in KFI data (8% lower demand growth).&lt;/p&gt;
&lt;h3 id="q12-how-does-this-paper-differ-from-and-relate-to-bloom-2009-and-bloom-et-al-2018"&gt;Q12. How does this paper differ from and relate to Bloom (2009) and Bloom et al. (2018)?&lt;/h3&gt;
&lt;p&gt;Bloom (2009) and Bloom et al. (2018) model a single composite firm-level shock (implicitly TFPR) in a CES-demand economy, finding that uncertainty shocks reduce output through wait-and-see behavior but generate a positive volatility effect (OHA) that partly offsets the uncertainty effect. The present paper adds two departures: (1) it separates TFPQ and demand shocks and shows they have distinct empirical and aggregate implications; (2) it replaces CES demand with an estimated non-CES demand curve. Departure (2) reverses the OHA effect, amplifying the total output decline by around 40% relative to the CES model. Departure (1) shows that the uncertainty channel operates primarily through demand, while TFPQ operates primarily through the volatility channel. The quantitative model uses the same non-convex adjustment cost structure and calibration approach as Bloom et al. (2018) to ensure comparability. The paper also relates to Bachmann and Bayer (2013) and Mongey and Williams (2017), who find smaller aggregate effects with adjustment costs only on capital; the present paper notes that adjustment costs on both capital and labor are needed for large wait-and-see effects, but qualitative conclusions are unchanged with capital-only costs.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-policy-and-theoretical-implications-of-the-findings"&gt;Q13. What are the policy and theoretical implications of the findings?&lt;/h3&gt;
&lt;p&gt;First, policies aimed at reducing firm-level demand uncertainty (e.g., demand stabilization, aggregate demand management) have larger aggregate output effects than policies addressing productivity uncertainty, because demand uncertainty triggers wait-and-see investment behavior while TFPQ uncertainty is largely absorbed in markups without changing investment much. Second, TFPQ dispersion is still harmful but through misallocation: policies that reduce markup dispersion induced by productivity differentials can raise aggregate output without requiring reduced dispersion per se. Third, the finding that TFPR dispersion is a poor proxy for demand shock dispersion has implications for how researchers use TFPR as a measure of misallocation or uncertainty: it conflates two distinct forces with different aggregate implications. Fourth, the estimated super-elasticity provides a data-disciplined input for calibrating models with real rigidities, directly relevant for the Ball-Romer nominal non-neutrality question—higher real rigidities amplify the output effects of monetary policy shocks. The authors flag this as a natural extension. The scope conditions are: Swedish manufacturing, annual data 1998-2013, partial equilibrium model (aggregate price level exogenous), firms with matching price and utilization data (large-firm bias).&lt;/p&gt;
&lt;h3 id="q14-what-additional-findings-are-documented-regarding-the-cyclicality-of-other-firm-level-variables"&gt;Q14. What additional findings are documented regarding the cyclicality of other firm-level variables?&lt;/h3&gt;
&lt;p&gt;Beyond TFPQ and demand dispersion, the paper documents that dispersion of sales growth, price growth, labor, intermediate goods, and capacity utilization are all countercyclical. The IQR of sales growth was 58% above the non-recession average in 2009 and 9% above in 2001; the IQR of price growth was 83% above in 2009 and 5% above in 2001. The one notable exception is investment, which displays procyclical dispersion (less dispersed during the Great Recession). The paper also documents that roughly 30% of firms report insufficient demand at all their plants in the survey data; average capacity utilization is 88% with median 91% and standard deviation of 14.1%; and about 25% of firm-year observations involve utilization at or above 100%.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Physical total factor productivity (TFPQ)&lt;/strong&gt;: Firm-level quantity productivity: output per unit of inputs, measured from a utilization-adjusted Cobb-Douglas value-added production function. Distinct from revenue TFP (TFPR = p*z) because it abstracts from demand conditions and price-setting. In this paper, TFPQ is estimated within firm over time using the cost-share approach and a capacity-utilization correction from managerial survey data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Demand shock (epsilon)&lt;/strong&gt;: The idiosyncratic component of a firm&amp;rsquo;s demand curve that captures its ability to sell more (or fewer) units at a given price in a given year, reflecting changes in customer base size or customers&amp;rsquo; willingness to pay. Estimated as the residual from the GIR demand curve after controlling for firm fixed effects, sector-time fixed effects, and the firm&amp;rsquo;s own price.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-CES demand curve / super-elasticity (eta)&lt;/strong&gt;: A demand specification adapted from Gopinath, Itskhoki, and Rigobon (2010) in which the demand elasticity is not constant but rises with the firm&amp;rsquo;s price. The parameter eta (estimated at 4.27 in the main sample) governs how fast the elasticity rises with the price: when eta &amp;gt; 0, firms gain few customers by cutting price (elasticity falls as price falls) and lose many customers by raising price (elasticity rises as price rises). This is the source of &amp;lsquo;real rigidity&amp;rsquo; that makes incomplete TFPQ passthrough optimal.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Incomplete TFPQ passthrough&lt;/strong&gt;: The empirical finding that firms reduce their prices by far less than one-for-one in response to a productivity gain (estimated beta_z = -0.097 to -0.124, far from the CES benchmark of -1). The paper attributes this primarily to non-CES demand real rigidity (which implies an optimal static passthrough of only 41% given the estimated parameters) and secondarily to adjustment costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Oi-Hartman-Abel (OHA) effect&lt;/strong&gt;: The positive &amp;lsquo;volatility effect&amp;rsquo; in standard CES-demand uncertainty models: because output is a convex function of TFPQ under CES, a mean-preserving spread in productivity raises aggregate output (lucky firms expand more than unlucky firms contract). The paper overturns this result by showing that with non-CES demand (eta sufficiently large), the output-productivity relationship becomes concave, so TFPQ dispersion reduces aggregate output via markup misallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wait-and-see channel&lt;/strong&gt;: The mechanism by which uncertainty about future shocks causes firms with non-convex input adjustment costs to pause investment: firms prefer to remain inactive and let inputs depreciate rather than invest or disinvest, at the risk of having to pay an irreversibility cost if the shock turns out to have been in the opposite direction. In this paper, this channel is driven primarily by demand uncertainty because demand shocks determine how many units a firm can sell and hence its desired input level; TFPQ uncertainty does not trigger strong wait-and-see behavior because the optimal scale response to TFPQ shocks is small under non-CES demand.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Markup dispersion / misallocation&lt;/strong&gt;: Dispersion across firms in the ratio of price to marginal cost, arising in this paper from incomplete TFPQ passthrough: firms with high productivity set high markups rather than passing through productivity gains as price cuts. The resulting wedge between prices and marginal costs means that resources are misallocated (too little output at high-productivity firms relative to the social optimum), reducing aggregate output. This is the channel through which TFPQ dispersion harms the aggregate economy in the model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Price wedge (tau)&lt;/strong&gt;: The residual from the passthrough regression: the component of firm price changes unexplained by the estimated TFPQ and demand shocks. Interpreted as capturing un-modeled shocks (financial constraints, markup adjustments) and potentially measurement error. The price wedge makes a meaningful contribution to both average sales/price dispersion and to the Great Recession increase in dispersion.&lt;/p&gt;</description></item><item><title>Distortions, Producer Dynamics, and Aggregate Productivity: A General Equilibrium Analysis</title><link>https://macropaperwarehouse.com/papers/distortions-producer-dynamics-and-aggregate-productivity-a-general-equilibrium-analysis/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/distortions-producer-dynamics-and-aggregate-productivity-a-general-equilibrium-analysis/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks how institutional distortions to factor markets affect not only the static allocation of inputs across farms but also the dynamic choices — crop selection and productivity-enhancing investment — that determine the long-run distribution of farm productivities and hence aggregate agricultural TFP. The question matters because prior work on misallocation has largely treated the productivity distribution as exogenous; this paper endogenizes it, showing that the dynamic channels can be quantitatively larger than the static factor-misallocation channel.&lt;/p&gt;
&lt;p&gt;The empirical foundation is the Vietnam Access to Resources Household Survey (VARHS), a balanced panel of 2,118 farm households surveyed biennially from 2006 to 2016 across twelve provinces in north and south Vietnam. Vietnam provides a natural laboratory: post-1986 reforms decollectivized agriculture nationally, but deeply divergent pre-reform institutions (collective agriculture in the north for more than three decades; private household farming in the south throughout) produced durable differences in land-market functioning, crop-choice restrictions, and property-rights security. Measured TFP is more than 2.5 times higher in the south than the north (the observed log TFP ratio implies roughly a 2.5-fold level difference). The elasticity of land use with respect to farm TFP is 0.554 in the south versus 0.152 in the north, and the elasticity of labor use is 0.382 versus 0.122 — three to four times larger in the south — indicating far more efficient resource allocation in the south. The share of perennial-crop farmers (high-value cash crops, especially coffee) is 33% in the south and roughly 5% in the north. Average biennial TFP growth is 6.2% in the south versus 2.6% in the north.&lt;/p&gt;
&lt;p&gt;The authors build a dynamic general equilibrium model of heterogeneous farm managers (following Lucas 1978) in which farm productivity has four components: a permanent farmer-specific component, a random transitory component, an endogenous managerial ability component accumulated through investment, and a crop-specific component tied to endogenous crop choice. Institutional distortions are modeled as idiosyncratic revenue taxes correlated with farm productivity (following Restuccia and Rogerson 2008), with the key parameter being the elasticity of distortions with respect to farm productivity (rho). A higher rho means more-productive farms face proportionately larger distortions, which (i) compresses the gap between large and small farms in equilibrium factor use, and (ii) reduces the private return to investing in ability. The model also incorporates government-imposed crop restrictions that force a fraction of farms to grow rice regardless of profitability. The model is calibrated to south Vietnam moments: average TFP growth, dispersion in TFP and growth, the land-size distribution, the measured elasticity of distortions, and crop shares. Measurement error in output and inputs is explicitly modeled following Bils, Klenow, and Ruane (2021); the estimated BKR statistic is 0.906 for the south and 0.987 for the north, indicating relatively limited measurement error by manufacturing-sector standards.&lt;/p&gt;
&lt;p&gt;The main counterfactual imposes north Vietnam distortion parameters on the south-calibrated benchmark economy. Three distortion parameters differ: (1) the distortion elasticity rho rises from 0.79 (south) to 0.91 (north); (2) crop-specific distortions flip sign — in the south perennials face lower effective taxes than rice (phi_perennial = 1.61 &amp;gt; 1), while in the north perennials face higher effective taxes than rice (phi_perennial = 0.68 &amp;lt; 1); (3) the share of farms subject to government-imposed crop restrictions rises from 23% to 43%.&lt;/p&gt;
&lt;p&gt;The counterfactual experiment produces four main quantitative results. First, aggregate TFP falls by 41% relative to the benchmark, accounting for 61% of the observed productivity gap between north and south Vietnam (the observed ratio is 0.42; the counterfactual ratio is 0.59). Second, the average biennial farm TFP growth rate falls by 1.6 percentage points (from 6.23% to 4.60%), accounting for just under half of the observed 3.6 percentage-point north-south gap. Third, TFP dispersion (standard deviation of log TFP) falls by 8 percentage points, more than half of the 14-percentage-point lower dispersion observed in the north. Fourth, the share of perennial farmers collapses from 33% to 9%, closely matching the observed 5% in the north.&lt;/p&gt;
&lt;p&gt;Channel decomposition reveals that static factor misallocation alone reduces output by 19.4% (one-third of the total 40.8% gap, proportionately allocated), while the crop-choice channel reduces output by 8.0% and the farm-ability channel (endogenous investment) reduces output by 31.6%. Together, the dynamic channels (crop choice plus farm ability) account for approximately two-thirds of the total productivity loss, more than doubling the contribution of static misallocation. Among individual distortions, the distortion elasticity rho alone accounts for a 38.3% output reduction, crop-specific distortions account for 7.5%, and government crop restrictions account for only 1.4%. The key mechanism is that a small increase in rho (from 0.79 to 0.91) has large productivity consequences because the productivity cost is convex in rho and accelerates as rho approaches one — at rho = 1, distortions fully absorb all incremental profits from higher ability, eliminating investment incentives entirely.&lt;/p&gt;
&lt;p&gt;The paper shows that measurement error has limited impact on the north-south comparison (since the main experiment is a within-survey, within-country comparison), but substantially inflates the level gains from removing all distortions: removing measurement error from the model more than doubles the estimated gains from moving to a first-best economy, underscoring that measurement error matters most in cross-economy level comparisons.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-the-main-threats-to-validity"&gt;Q1. What is the identification strategy and the main threats to validity?&lt;/h3&gt;
&lt;p&gt;The identification exploits within-country, within-survey variation between north and south Vietnam, which share a common currency, survey instrument, and price measurement methodology. The main threat is that technology and geography differ across regions beyond institutions. The paper addresses this in two ways. First, it restricts comparisons to the two rice-growing delta regions — the Red River Delta (north) and Mekong Delta (south) — where technology and geographic differences are minimal, and shows the same patterns hold: measured distortion elasticity in the Mekong Delta is 0.79 versus 0.94 in the Red River Delta, and growth is higher and productivity more dispersed in the south. Second, the paper uses FAO Global Agro-Ecological Zones data to show land quality differences are negligible between north and south and, if anything, slightly favor the north; when scaled through the production function (land share times span-of-control = 0.35), land quality cannot account for the observed TFP gap. A second threat is measurement error inflating wedge dispersion and the estimated distortion elasticity. The paper addresses this by embedding explicit measurement error in the calibration and by using the Bils-Klenow-Ruane (2021) methodology, finding BKR statistics of 0.91 (south) and 0.99 (north), suggesting measurement error is modest in agriculture relative to manufacturing. The calibrated true distortion elasticity for the south is rho = 0.79, versus a measured elasticity of 0.86, a bias of around 0.06 — consistent with BKR estimates.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-productivity-channels-and-how-is-each-measured"&gt;Q2. What are the three productivity channels and how is each measured?&lt;/h3&gt;
&lt;p&gt;The three channels are (1) static factor misallocation, (2) crop distribution, and (3) farm ability. Each is isolated by a sequential decomposition: for factor misallocation, counterfactual distortions rho and phi are imposed while holding the crop and ability distributions fixed at benchmark-economy values, yielding an output loss of 19.4%. For crop distribution, the crop shares are adjusted to the counterfactual economy while holding within-crop ability distributions fixed at benchmark values; output falls by 8.0%. For farm ability, the ability distribution conditional on crop type is adjusted to the counterfactual while holding crop shares fixed; output falls by 31.6%. The sum (59.0%) exceeds the total gap (40.8%) because of negative interactions among channels — factor misallocation has a smaller bite when the productivity distribution is more compressed, as in the counterfactual.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-role-of-the-distortion-elasticity-parameter-rho-and-why-does-it-generate-outsized-productivity-losses-from-a-small-change"&gt;Q3. What is the role of the distortion elasticity parameter rho and why does it generate outsized productivity losses from a small change?&lt;/h3&gt;
&lt;p&gt;The parameter rho governs the extent to which more productive farms face proportionately larger distortions. At rho = 0, distortions are orthogonal to productivity; at rho = 1, distortions grow one-for-one with productivity, fully taxing away any incremental profit from increasing ability. The investment return to moving up the ability ladder is proportional to the incremental profit gained, which equals (1 - rho) times the increment in revenue. As rho rises toward 1, this return collapses toward zero. Because the South&amp;rsquo;s calibrated rho is already 0.79 — close to 1 on the relevant scale — a further increase to 0.91 is disproportionately large in terms of investment disincentives. The paper demonstrates this asymmetry explicitly in Appendix C.6: a symmetric increase and decrease of rho by 0.1 (set to the observed North-South difference in measured elasticity) reduces output by 42% when rho rises but only 39% when rho falls, driven primarily by the farm-ability channel (27 log points versus 21 log points difference in log output).&lt;/p&gt;
&lt;h3 id="q4-how-do-crop-specific-distortions-and-government-crop-restrictions-work-and-what-is-their-quantitative-contribution"&gt;Q4. How do crop-specific distortions and government crop restrictions work and what is their quantitative contribution?&lt;/h3&gt;
&lt;p&gt;Crop-specific distortions phi_i create wedges that differ across crop types. In the south, phi_perennial = 1.61 (perennial growers face lower effective taxes than rice farmers), while in the north phi_perennial = 0.68 (perennial growers face higher effective taxes). This reversal in relative distortions discourages perennial farming in the north both directly (lower profits) and dynamically (perennial farmers, who tend to be higher-ability, invest less). Unilaterally imposing north crop-specific distortions on the south benchmark reduces output by 7.5%. Government-imposed crop restrictions force a fraction omega of farms to grow rice regardless of profitability, with omega rising from 23% to 43% north-south. This channel has the smallest impact (1.4% output loss) because: (a) a large fraction of restricted farmers would have chosen rice anyway, and (b) back-of-envelope calculation shows the loss amounts to reducing productivity of only about 7% of farmers (the 20 percentage-point change in omega times the 33% perennial share) by about 20% (measured perennial-rice TFP gap).&lt;/p&gt;
&lt;h3 id="q5-what-empirical-evidence-motivates-the-endogenous-investment-mechanism"&gt;Q5. What empirical evidence motivates the endogenous investment mechanism?&lt;/h3&gt;
&lt;p&gt;Table 3 shows that in both north and south Vietnam, farm investment (cash or labor investment in irrigation or soil/water conservation) and extension-service participation are positively correlated with farm TFP and negatively correlated with farm-level distortion wedges, indicating that more distorted farms invest less. In the south, both investment and extension services are significantly positively associated with subsequent TFP growth. In the north, only extension-service participation is positively associated with future growth, while physical investment is not — suggesting the return to investment is suppressed in the north. The data also document a life-cycle profile (Figure 3) in which farm TFP rises steeply for young farms and then levels off, much more sharply in the south than in the north, consistent with faster ability accumulation in the less-distorted south.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-is-documented-across-crop-types-within-each-region"&gt;Q6. What heterogeneity is documented across crop types within each region?&lt;/h3&gt;
&lt;p&gt;In the south, perennial farmers have higher average output (log output 10.6 vs. 9.9 for rice), more land (3.9 acres vs. 2.4), more labor, higher TFP (above mean relative to rice), and far higher biennial TFP growth (10.9% vs. 4.9%). In the north, the pattern reverses: perennial farmers underperform relative to rice farmers in output (-0.583 log points, significant), land, labor, and TFP (-0.413 log points). This reversal occurs because crop-specific distortions disproportionately penalize perennial farming in the north. Despite the average gaps, there is substantial productivity overlap across crop types within both regions (Figure A.1), with many unproductive perennial farmers and productive rice farmers coexisting. This overlap motivates the paper&amp;rsquo;s modeling of crop selection as a utility-cost decision with idiosyncratic taste heterogeneity (Frechet distribution), rather than a pure productivity-cutoff rule.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-conducted-and-what-do-they-show"&gt;Q7. What robustness checks are conducted and what do they show?&lt;/h3&gt;
&lt;p&gt;Four robustness exercises are conducted. First, re-calibrating with fixed quadratic investment-cost curvature (zeta = 2) instead of the estimated 1.74 yields a counterfactual output ratio of 58.8%, similar to the baseline 59.2%. Second, lowering the targeted average growth rate by 2 percentage points (addressing the concern that aggregate TFP growth partly reflects economy-wide technology rather than ability investment) produces a counterfactual output ratio of 58.6% — essentially unchanged. Third, lowering the targeted growth rate by 4 percentage points produces 62.4%, still economically large. Fourth, two model extensions are explored: (a) incorporating a hump-shaped life-cycle profile with a young-to-old transition produces a 43% productivity loss, similar to the 41% baseline; (b) allowing entrants to draw ability from a distribution dependent on the exiting predecessor&amp;rsquo;s ability produces a 57% productivity loss — larger than baseline because investment creates positive spillovers to future entrants.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-the-prior-misallocation-literature"&gt;Q8. How does this paper relate to and differ from the prior misallocation literature?&lt;/h3&gt;
&lt;p&gt;The paper builds on Restuccia and Rogerson (2008) and Hsieh and Klenow (2009), who model static misallocation via idiosyncratic wedges. It contributes three extensions. First, it endogenizes the farm productivity distribution by adding investment and crop choice, so that the same wedges that generate static misallocation also distort dynamics — this doubles the productivity cost. Second, the experiment is a within-country comparison between two regions rather than a comparison against a hypothetical undistorted economy, avoiding the criticism that the undistorted benchmark is unrealistic. The re-calibrated north model accounts for 100% of the observed north-south TFP ratio (40.7% model vs. 42% data). Third, the dynamic model generates falsifiable predictions about farm TFP growth rates, TFP dispersion, and crop distributions — all of which move in the right directions — providing a richer validation test than static models allow. The paper also relates to Hsieh and Klenow (2014), who document faster life-cycle productivity growth in less distorted economies (India and Mexico vs. US), and to Adamopoulos and Restuccia (2020), who study land reform in Vietnam but with exogenous productivity distributions; the current paper finds that endogenizing productivity distributions significantly amplifies the costs of distortions. The measurement-error treatment follows Bils, Klenow, and Ruane (2021) and Adamopoulos et al. (2022).&lt;/p&gt;
&lt;h3 id="q9-what-does-the-model-imply-about-a-hypothetical-undistorted-economy"&gt;Q9. What does the model imply about a hypothetical undistorted economy?&lt;/h3&gt;
&lt;p&gt;Removing all distortions (rho = 0, phi_i = 1 for all crops, omega = 0, sigma_epsilon = 0) increases TFP by a factor of 3.37 relative to the south benchmark (Appendix C.5, Table C.11), meaning the first-best economy is more than three times as productive. Static reallocation gains alone (holding the productivity distribution fixed) account for roughly 70% of this gap. The remaining gains come from the endogenous shift in the ability distribution — in the undistorted economy, lower ability farmers invest less (because higher general equilibrium wages lower profits) but higher ability farmers invest more (because distortions no longer claw back incremental profits). The net result is a more polarized ability distribution with a heavier right tail, consistent with the concentrated structure of agriculture in advanced economies. Importantly, the paper cautions that abstracting from measurement error inflates the estimated undistorted-economy gains by more than a factor of two: a model without measurement error yields gains more than twice as large as the calibrated model that accounts for it.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central policy implication is that institutions distorting factor markets — particularly those that generate a positive correlation between farm productivity and the effective tax rate (captured by rho) — reduce agricultural TFP through three compounding channels, with two-thirds of the loss arising from dynamic distortions (investment suppression and crop selection) rather than static factor reallocation. This means that standard static calculations of misallocation costs substantially understate the true costs. Land accumulation restrictions that prevent productive farms from expanding (the historical legacy in north Vietnam, where 82.8% of Red River Delta agricultural land was state-allocated) are particularly costly because they are the empirical analog of high rho. The scope conditions are: (1) the analysis applies to the Vietnamese agricultural context in 2006-2016, a period well after initial reform but still characterized by persistent institutional differences; (2) the model abstracts from occupational choice and structural transformation, which other work has shown amplify distortion costs further; (3) the main results are robust to the north-south within-country design but level estimates (gains from the first-best) are sensitive to measurement error treatment. The paper suggests that reducing the productivity-distortion correlation — e.g., through secure land titles and functioning land rental markets — would unlock gains exceeding what static misallocation calculations imply.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Distortion elasticity (rho)&lt;/strong&gt;: The parameter governing how strongly institutional distortions — modeled as idiosyncratic revenue taxes — are correlated with farm-level productivity. A higher rho means more productive farms face proportionately larger distortions, compressing both static factor allocation and the dynamic return to investing in ability. In the paper&amp;rsquo;s calibration, rho = 0.79 for south Vietnam and 0.91 for north Vietnam; the difference accounts for the majority of the measured North-South productivity gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Managerial ability ladder&lt;/strong&gt;: The endogenous component of farm productivity that farmers accumulate through investment. A farmer at ability node h has productivity phi^h; investing e units of output raises ability to the next node with probability x = (e/a)^(1/zeta). The investment return depends on the incremental profit gain from higher ability, which is suppressed when the distortion elasticity rho is large, creating a tight link between static institutional distortions and dynamic farm growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crop-specific distortion (phi_i)&lt;/strong&gt;: A factor in the distortion specification that captures institutional barriers differentially affecting specific crops. In south Vietnam, phi_perennial = 1.61, meaning perennial-crop growers face lower effective taxes than rice farmers; in north Vietnam, phi_perennial = 0.68, reversing the ranking. This parameter embeds market-access barriers, infrastructure gaps, and regulatory disadvantages specific to particular crops.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Government-imposed crop restriction (omega)&lt;/strong&gt;: The share of farms legally required to grow rice regardless of relative profitability or household preferences, reflecting Vietnamese national food-security policies. The restriction is more prevalent in the north (43% of farms) than the south (23%). Unlike idiosyncratic distortions, crop restrictions enter the model as a direct constraint on the discrete crop-choice decision rather than as a tax on revenue, and the paper finds their productivity cost is relatively small (1.4% output loss) because many restricted farmers would have chosen rice anyway.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic misallocation&lt;/strong&gt;: The productivity losses arising from distortions&amp;rsquo; effects on farms&amp;rsquo; forward-looking decisions — specifically the choice of crop (crop selection) and investment in managerial ability — as opposed to the static misallocation of given factor inputs across farms with fixed productivities. In the paper, dynamic misallocation accounts for two-thirds of the total productivity gap, more than doubling the contribution of static factor misallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;BKR measurement-error statistic&lt;/strong&gt;: A diagnostic from Bils, Klenow, and Ruane (2021) that estimates the ratio of true wedge dispersion to observed wedge dispersion using the cross-term in a regression of log output changes on log wedge, log input, and their interaction. Values near one indicate little measurement error; values near zero indicate the observed wedge is mostly noise. The paper finds BKR = 0.906 for south Vietnam and 0.987 for north Vietnam, indicating measurement error is modest and is unlikely to confound the north-south comparison.&lt;/p&gt;</description></item><item><title>Environmental Subsidies to Mitigate Net-Zero Transition Costs</title><link>https://macropaperwarehouse.com/papers/environmental-subsidies-to-mitigate-net-zero-transition-costs/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/environmental-subsidies-to-mitigate-net-zero-transition-costs/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether public subsidies to green-technology producers, financed by a carbon tax, can materially reduce the macroeconomic cost of reaching net-zero CO2 emissions by 2060. The motivation is a market-structure failure that standard environmental models ignore: the abatement goods sector is initially immature and highly concentrated, with 10 percent of firms capturing roughly 80 percent of operating revenue (Eurostat/Ecorys data). Under such conditions a carbon tax alone raises the cost of abatement inputs, depresses competition, and generates a deep and prolonged GDP recession — even if it achieves the emissions target. The paper shows that redirecting carbon tax revenues toward subsidizing this sector can substantially offset the recession.&lt;/p&gt;
&lt;p&gt;The analytical vehicle is an environmental dynamic stochastic general equilibrium (E-DSGE) model for the world economy, built by merging three bodies of work: the DICE climate block (Nordhaus 1992, 2018), a real-business-cycle production structure in the spirit of Smets and Wouters (2007), and an endogenous market-structure framework for the abatement goods sector following Bilbiie, Ghironi, and Melitz (2012). Firm entry into the abatement sector responds to expected future profits, which depend on sunk costs. Two margins of adjustment are distinguished: the intensive margin (existing firms expanding production) and the extensive margin (startups creating new varieties). Competition in the abatement sector is a central object of analysis: higher firm numbers reduce the abatement price, which in turn lowers the carbon tax burden on final-goods producers.&lt;/p&gt;
&lt;p&gt;The model is estimated using Bayesian methods on five annual world time series from 1961 to 2019: real GDP growth, real consumption growth, CO2 emissions growth, the change in surface temperature anomaly, and the growth rate of environment-related patents (OECD). Because the model has stochastic growth trends, the authors use the extended-path solution method (Fair and Taylor 1983) rather than standard linearization, and an inversion filter to form the likelihood function. Posterior draws from 320,000 MCMC iterations (8 parallel chains, ~30 percent acceptance) pin down five structural parameters and ten shock parameters. Estimated initial output growth is approximately 4.99 percent per year and the initial emissions-to-output decoupling rate is 1.13 percent per year, both consistent with Nordhaus (1992) benchmarks. The temperature elasticity to radiative forcing (ξ_T) is estimated at 0.084, the abatement-sector exit rate at 0.06, and the entry congestion cost at 5.63.&lt;/p&gt;
&lt;p&gt;The paper implements projections from 2019 to 2100 under three IPCC-aligned scenarios (SSP1–1.9, SSP2–4.5, SSP3–7.0), focusing on the Paris Agreement target of limiting warming to below 2 degrees Celsius. In the laissez-faire (no-policy) scenario, emissions peak near 57 Gt CO2 in 2060 and 70 Gt in 2100, producing roughly 4 degrees Celsius of warming by 2100, with damages reaching 4 percent of GDP per year. In the below-2-degree scenario with a carbon tax only, the carbon tax must rise to approximately $480 per ton by 2080, abatement cost reaches 3.4 percent of GDP in 2060, and cumulative GDP loss from 2019 to 2060 totals $258 trillion (averaging $6.3 trillion per year, or 4.9 percent of 2019 world GDP). This is the baseline against which subsidies are evaluated.&lt;/p&gt;
&lt;p&gt;Two subsidy experiments are run, both fully financed by carbon tax revenue (budget neutral by construction). First, a subsidy targeted only at incumbent abatement firms (intensive margin): this immediately compresses the abatement price from 2.5 times to 1.5 times the price of the final good, reduces aggregate abatement cost from 2 percent to 0.8 percent of GDP in 2040, and brings the carbon tax needed to hit the emissions target down from $300 to $160 per ton in 2040. However, by lowering incumbents&amp;rsquo; labor costs and raising the equilibrium wage, the intensive-margin subsidy raises the cost of startup entry and reduces the number of abatement firms over time, deteriorating long-run competition.&lt;/p&gt;
&lt;p&gt;Second, an optimal subsidy that allocates carbon revenues between incumbents and startups. The optimal split is determined by maximizing social welfare (the infinite discounted sum of household utility) over a grid of subsidy shares. The welfare function is concave in the startup share, with a maximum at 60 percent of revenues to startups and 40 percent to incumbents. Under this optimal policy, the number of firms in the abatement sector nearly doubles relative to the baseline by 2050, the abatement price falls sharply, and the carbon tax needed to achieve the same emissions path drops to $125 per ton in 2040 versus $300 in the no-subsidy baseline. Cumulative GDP loss from 2019 to 2060 falls to $141 trillion ($138 trillion in one presentation, $141 trillion in another), saving approximately $120 to $123 trillion relative to the carbon-tax-only scenario, equivalent to roughly $2.9 trillion per year. The abatement price is reduced by more than a factor of 2.5 under the optimal subsidy regime.&lt;/p&gt;
&lt;p&gt;Present-value GDP subsidy multipliers (the ratio of discounted GDP gain to discounted subsidy expenditure) exceed 2.0 through 2035 and remain above 1.78 through 2060, with consumption multipliers ranging from 1.42 to 1.90 over the same horizon. These large multipliers reflect the competition-enhancing effect of startup subsidies: by accelerating firm entry, the policy lowers abatement prices for all final-goods producers, amplifying the direct subsidy impact. The largest GDP gains are concentrated in the first decade (2019–2030), when subsidies rapidly reduce the abatement price and induce firm entry. The scope condition for these results is the below-2-degree (SSP1–1.9) scenario with a simultaneous carbon-tax-and-subsidy announcement in 2019, a world-representative aggregate model, and the assumption that carbon tax revenues are fully recycled into the abatement sector rather than used for general government expenditure.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-central-market-failure-the-paper-addresses-and-why-does-it-make-a-carbon-tax-alone-insufficient"&gt;Q1. What is the central market failure the paper addresses, and why does it make a carbon tax alone insufficient?&lt;/h3&gt;
&lt;p&gt;The abatement goods sector is initially immature and highly concentrated (10 percent of firms account for roughly 80 percent of operating revenue). In the decentralized equilibrium, each final-goods firm is atomistic with respect to climate damage and so does not voluntarily abate. The carbon tax corrects this free-rider problem, but because the abatement market is imperfectly competitive, abatement goods are priced at a monopolistic markup (the abatement price begins at 2.5 times the price of the final good). The high abatement price raises the cost of reducing emissions, depresses the optimal abatement effort, and magnifies the GDP recession. A carbon tax alone thus generates a $258 trillion cumulative GDP loss by 2060. The paper&amp;rsquo;s main point is that subsidizing entry into the abatement sector introduces competition that compresses the markup, lowering both the abatement price and the required carbon tax rate.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-models-identification-strategy-and-what-are-the-main-econometric-challenges"&gt;Q2. What is the model&amp;rsquo;s identification strategy and what are the main econometric challenges?&lt;/h3&gt;
&lt;p&gt;The model is identified through full-information Bayesian maximum likelihood on five world aggregate series, 1961–2019. Climate block parameters are largely taken from DICE (Nordhaus 1992, 2018), narrowing the estimation to five structural parameters: initial output growth rate, initial emissions-to-output decoupling rate, temperature elasticity to radiative forcing (ξ_T), abatement-sector exit rate (δ_A), and entry congestion cost (χ). The main econometric challenges are (i) stochastic growth trends, which make standard linearization around a fixed point invalid — addressed with the extended-path solution method — and (ii) forming the likelihood for a nonlinear model, addressed with an inversion filter (Fair and Taylor 1983; Guerrieri and Iacoviello 2017) rather than computationally expensive particle filters. A drawback acknowledged by the authors is that Jensen&amp;rsquo;s inequality collapses to equality in the extended-path approach, so nonlinear uncertainty from future shocks is not captured — the same limitation that applies to standard linearized DSGE models.&lt;/p&gt;
&lt;h3 id="q3-how-are-the-intensive-and-extensive-margins-of-adjustment-to-the-carbon-tax-distinguished-in-the-model-and-why-does-this-distinction-matter-for-policy"&gt;Q3. How are the intensive and extensive margins of adjustment to the carbon tax distinguished in the model, and why does this distinction matter for policy?&lt;/h3&gt;
&lt;p&gt;The intensive margin refers to incumbent abatement firms increasing the quantity produced of existing varieties. The extensive margin refers to households creating new startups that introduce additional varieties of abatement goods. The distinction matters because (i) more varieties increase competition and compress the abatement price (via a price-index formula: aggregate abatement price falls with firm numbers), and (ii) the two margins respond differently to subsidy design. A subsidy only to incumbents immediately lowers production costs and the abatement price but raises the equilibrium wage, which increases the sunk cost for prospective entrants and crowds out startup entry over time, ultimately harming competition. A subsidy to startups has a delayed effect — startups take one period to begin producing — but generates a sustained competitive effect that eventually exceeds the immediate gain from the incumbent-only policy. The welfare-maximizing policy therefore combines both, weighting startups at 60 percent.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-optimal-subsidy-split-and-how-is-it-determined"&gt;Q4. What is the optimal subsidy split and how is it determined?&lt;/h3&gt;
&lt;p&gt;The optimal split allocates 60 percent of carbon tax revenues to subsidizing startups&amp;rsquo; sunk entry costs and 40 percent to reducing incumbents&amp;rsquo; production costs (labor input subsidies). This is determined by computing the present value of household welfare (infinite discounted sum of utilities evaluated at 2019 when the policy is announced) for each value of the subsidy share on a fine grid. The welfare function is strictly concave in the startup share, rising until the startup share reaches 0.6 and declining thereafter. The intuition for concavity is that subsidizing startups has a long-horizon payoff (gradual entry and competition), while subsidizing incumbents has an immediate payoff (price reduction) but a long-run cost (reduced entry incentive). The optimum balances these dynamics.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-quantitative-effects-of-the-optimal-subsidy-on-the-carbon-tax-path-abatement-prices-and-firm-numbers"&gt;Q5. What are the quantitative effects of the optimal subsidy on the carbon tax path, abatement prices, and firm numbers?&lt;/h3&gt;
&lt;p&gt;Relative to the no-subsidy carbon-tax-only baseline: (1) The carbon tax needed to hit net-zero by 2060 falls from approximately $300 per ton in 2040 to $125 per ton under the optimal subsidy, and from approximately $390–$480 per ton in later years to correspondingly lower values. (2) The abatement price is reduced by more than a factor of 2.5 over the horizon. (3) The number of firms in the abatement goods sector nearly doubles by 2050 relative to the baseline. (4) Abatement cost as a share of output falls substantially, from the baseline peak of approximately 3.4 percent of GDP in 2060 to a lower trajectory. (5) Detrended output in 2040 improves from approximately -3 percent (baseline) to -1 percent under the optimal subsidy, and from -3.2 percent to -2 percent in 2050. These numbers are conditional on the below-2-degree warming scenario and the announced policy starting in 2019.&lt;/p&gt;
&lt;h3 id="q6-how-large-are-the-subsidy-fiscal-multipliers-and-what-drives-them"&gt;Q6. How large are the subsidy fiscal multipliers and what drives them?&lt;/h3&gt;
&lt;p&gt;GDP subsidy multipliers (present value of GDP gain per unit of present value of subsidy expenditure) are approximately 2.27 at the 2030 horizon, 2.03 at 2035, 1.89 at 2040, 1.81 at 2045, 1.78 at 2050, 1.80 at 2055, and 1.85 at 2060. Consumption multipliers are uniformly lower but remain above 1.4 throughout. The high multipliers are driven by the competition channel: each dollar of subsidy to startups reduces the abatement price for all final-goods producers economy-wide, amplifying the direct expenditure effect many times over. Multipliers exceed 2 in the early years when startup entry is most rapid and the abatement-price reduction is sharpest. The slight uptick in multipliers at the 2060 horizon reflects the long-run dynamics of the abatement sector reaching a more competitive equilibrium.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-role-of-the-dice-climate-block-and-what-simplifications-are-made-relative-to-state-of-the-art-climate-science"&gt;Q7. What is the role of the DICE climate block and what simplifications are made relative to state-of-the-art climate science?&lt;/h3&gt;
&lt;p&gt;The climate block is taken directly from DICE-1992 and DICE-2016R2 (Nordhaus 1992, 2018). It models atmospheric CO2 accumulation, radiative forcing from CO2 and non-CO2 sources, and two-box (surface and deep-ocean) temperature dynamics. Key DICE parameters (φ_11, φ_12, φ_21, φ_22, ξ_M, M_1750, damage cost a) are calibrated to match DICE values. The temperature sensitivity parameter ξ_T is estimated from the data rather than calibrated, yielding 0.084, slightly below DICE 2013 and 2016 values. The authors explicitly note that more advanced climate blocks are important for physical risk assessment but have &amp;rsquo;little added value&amp;rsquo; for transition risk analysis, which concerns the costs of policy, not the physical hazard. The non-CO2 radiative forcing follows a deterministic path that caps at F_max by 2100. The damage function is quadratic in surface temperature: Φ(T_t) = 1/(1+aT_t^2). In the laissez-faire scenario, this implies damages of 1.5 percent of GDP by 2050 and 4 percent by 2100.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-compare-to-the-standard-dice-model-and-what-does-the-comparison-reveal"&gt;Q8. How does the paper compare to the standard DICE model and what does the comparison reveal?&lt;/h3&gt;
&lt;p&gt;The authors estimate both the E-DSGE (with endogenous firm entry in the abatement sector) and a version equivalent to DICE (with perfect competition and no firm-entry dynamics) on the same data. Both models match the empirical second moments (standard deviations and autocorrelations of the five observables) comparably, so standard information criteria cannot discriminate between them. The key difference is that the E-DSGE model reproduces the standard deviation and autocorrelation of patent growth (the proxy for abatement-sector entry), which the DICE version cannot by construction (it has no entry shock). In DICE-like environments, the abatement sector is assumed competitive from the outset and the abatement price equals 1 (the final-goods price), so there are no dynamics in abatement pricing or firm numbers. This means DICE models understate transition costs when the abatement market is initially concentrated, and miss the welfare gain from competition-enhancing policies.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-role-of-the-endogenous-market-structure-mechanism-and-how-does-it-relate-to-solar-photovoltaic-markets"&gt;Q9. What is the role of the endogenous market structure mechanism and how does it relate to solar photovoltaic markets?&lt;/h3&gt;
&lt;p&gt;The paper argues the solar PV market provides historical validation of the model mechanism. From the late 1970s to 2019, the cumulative number of solar PV patents increased dramatically while module costs fell precipitously (the cost of solar PV modules in 2019 USD per watt fell 45 percent between 1990 and 2000, 58 percent between 2000 and 2010, and 81 percent between 2010 and 2019). The model predicts exactly this pattern: an initial carbon policy raises expected profits in the abatement sector, inducing entry, which intensifies competition and compresses prices. The initial abatement price in the model (2.5 times the final-goods price) eventually falls below 1 after 2040 under a carbon-tax-only policy. The paper notes the solar sector&amp;rsquo;s trajectory was partly driven by government subsidies in several countries, consistent with the model&amp;rsquo;s policy recommendation.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-main-shock-processes-in-the-model-and-what-do-impulse-response-functions-reveal"&gt;Q10. What are the main shock processes in the model and what do impulse response functions reveal?&lt;/h3&gt;
&lt;p&gt;Five structural shocks are estimated: TFP (productivity), government spending, CO2 emissions, firm entry (innovation), and temperature. All are AR(1) processes. Estimated AR(1) coefficients: productivity 0.949, government spending 0.867, CO2 emissions 0.940, firm entry 0.592, temperature 0.181 — so temperature shocks are nearly serially uncorrelated at annual frequency. Generalized impulse response functions (computed at 2019 state variables, averaged over 500 draws) show: (1) A positive productivity shock raises output and worsens emissions, stimulating abatement-sector entry and reducing the abatement price. (2) A positive CO2 emissions shock triggers a sharp abatement effort and firm entry, but depresses output by almost 5 percent in the short run. (3) A government spending shock (demand shock) raises final-good production, worsens emissions, but crowds out abatement — abatement effort and firm numbers fall 5 percent and 1.1 percent respectively. (4) A firm-entry shock raises firm numbers by nearly 10 percent at peak, reducing abatement prices and encouraging abatement effort without increasing emissions. (5) A temperature shock depresses output by more than 6 percent initially, reducing emissions and abatement effort, and shrinking the abatement sector while pushing abatement prices up.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-three-ipcc-aligned-scenarios-used-in-the-projections-and-how-do-they-differ"&gt;Q11. What are the three IPCC-aligned scenarios used in the projections and how do they differ?&lt;/h3&gt;
&lt;p&gt;The three scenarios correspond to SSP1–1.9, SSP2–4.5, and SSP3–7.0. (1) Below +2 degrees C (SSP1–1.9): carbon neutrality by 2060, followed by negative emissions (up to -10 Gt by 2100). Requires the carbon tax to rise to approximately $480 per ton by 2080. Abatement cost reaches 3.4 percent of GDP in 2060. This is the scenario used for the policy experiments. (2) Below +3 degrees C (SSP2–4.5): carbon neutrality delayed to shortly after 2100. Carbon tax rises gradually to $300 per ton by 2100. Abatement cost rises to 0.5 percent of GDP in 2050 and 1.2 percent by 2100. Detrended output falls to -3 percent by 2060. (3) +4 degrees C (SSP3–7.0): no policy, laissez-faire. Emissions peak at 57 Gt in 2060 and 70 Gt in 2100. Temperature rises approximately 4 degrees C by 2100. Damages reach 4 percent of GDP per year by 2100. Detrended output decreases from 3 percent to -1 percent by 2050 and -3 percent by 2100 due to climate damage alone.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-main-policy-implications-and-their-scope-conditions"&gt;Q12. What are the main policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central implication is that carbon tax revenues should not be recycled to households as lump-sum transfers (the conventional approach in environmental economics) but should instead be used to subsidize entry and operation in the abatement goods sector. The welfare-maximizing split is 60 percent to startups and 40 percent to incumbents. This reduces the cumulative GDP loss from $258 trillion to approximately $138–141 trillion by 2060, saving roughly $120–123 trillion total ($2.9 trillion per year on average). Scope conditions: (1) The result is conditional on the below-2-degree Paris scenario — less stringent emissions targets require lower carbon taxes and generate smaller transition costs, so the absolute gain from subsidies would be smaller. (2) The policy must be announced credibly in advance (2019 in the simulation) so that firms adjust expectations and entry decisions. (3) The model abstracts from capital, cross-country heterogeneity, sector-level differences, and physical risks from climate change. (4) Stochastic uncertainty about future shocks is not incorporated into the policy optimization (extended-path solution collapses uncertainty around the deterministic path). The authors suggest future work should evaluate the optimal policy accounting for stochastic climate and economic risks (following Cai and Lontzek 2019).&lt;/p&gt;
&lt;h3 id="q13-how-does-the-paper-relate-to-prior-e-dsge-and-iam-literature-and-what-is-novel"&gt;Q13. How does the paper relate to prior E-DSGE and IAM literature, and what is novel?&lt;/h3&gt;
&lt;p&gt;The paper positions itself relative to two literatures. First, integrated assessment models (IAMs) originating with DICE (Nordhaus 1992, 1994): IAMs provide long-run analysis but lack microfounded expectations and uncertainty. Second, E-DSGE models (Fischer and Springborn 2011; Heutel 2012; Angelopoulos et al. 2013; Golosov et al. 2014; Annicchiarico and Di Dio 2015, 2017; Diluiso et al. 2021): these have microfoundations and handle short-run dynamics well but typically operate in a linearized, stationary framework unsuited for long-run climate trends. Some prior E-DSGE work includes endogenous entry (Annicchiarico et al. 2018; Shapiro and Metcalf 2021) but focuses on short-run analysis or specific country (U.S.) settings. The paper&amp;rsquo;s novelties are: (1) Merging DICE with a BGM-style endogenous market structure for the abatement sector in a unified framework suitable for long-run analysis; (2) Nonlinear estimation of the E-DSGE model using the extended-path plus inversion-filter approach — the authors claim this is the first attempt to estimate a nonlinear E-DSGE with both environmental and macroeconomic trends; (3) Distinguishing intensive and extensive margins of abatement-sector adjustment and optimizing the subsidy split between them; (4) Computing present-value subsidy multipliers for climate policy.&lt;/p&gt;
&lt;h3 id="q14-what-are-the-main-limitations-and-caveats-acknowledged-by-the-authors"&gt;Q14. What are the main limitations and caveats acknowledged by the authors?&lt;/h3&gt;
&lt;p&gt;The authors acknowledge several limitations. (1) Capital is excluded from the production function to keep the model tractable given the focus on the abatement goods sector and endogenous entry. (2) The model is a world aggregate with no cross-country heterogeneity; a multicountry model would be needed to study distributional effects across nations. (3) The policy analysis is conditional on the below-2-degree scenario and does not account for uncertainty about future economic and climate conditions — the extended-path method does not incorporate stochastic uncertainty in the forward-looking path. (4) The analysis does not account for the positive benefits of avoided physical risk from climate change (reduced damages in alternative scenarios are noted but not attributed to subsidy policy per se). (5) Non-CO2 radiative forcing is modeled as a simple deterministic path, which simplifies the climate dynamics. (6) The comparison with DICE via second moments rather than formal model selection criteria (since the DICE version has one fewer observable and one fewer shock) limits the formal identification of the endogenous entry mechanism. (7) The model does not include labor market frictions, nominal rigidities, or financial frictions, all of which could affect transition dynamics.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Abatement goods sector&lt;/strong&gt;: In this paper, the sector producing intermediate inputs (abatement goods) purchased by final-goods firms to reduce their CO2 emissions. The sector is initially immature and highly concentrated, with high barriers to entry that prevent competition and keep abatement prices above the price of the final good. The paper models this sector with endogenous firm entry following Bilbiie, Ghironi, and Melitz (2012), distinguishing between incumbents (intensive margin) and startups (extensive margin).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Transition risk&lt;/strong&gt;: In this paper, the macroeconomic cost — in terms of GDP loss, employment diversion, and abatement expenditure — of implementing climate policy (specifically a carbon tax path) to achieve net-zero emissions by 2060. Transition risk is distinct from physical risk (climate damage to productivity); the paper focuses exclusively on transition risk and does not account for avoided physical risk when evaluating policy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous market structure&lt;/strong&gt;: The property that the number of firms (varieties) in the abatement goods sector is not fixed but responds endogenously to expected future profits, sunk entry costs, and exit shocks. Following Bilbiie et al. (2012), the paper models a free-entry condition where households create startups until the marginal cost of entry (sunk cost) equals the expected discounted value of future profits. This endogeneity allows the model to capture how carbon taxes and subsidies affect abatement-sector competition and prices over time.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive margin vs. extensive margin (abatement sector)&lt;/strong&gt;: The intensive margin refers to adjustment by existing (incumbent) abatement firms — increasing production of current varieties when demand rises. The extensive margin refers to the creation of new firms (startups) that introduce additional varieties. The paper shows these margins respond differently to subsidy design: incumbent subsidies have immediate price effects but crowd out entry; startup subsidies have delayed effects but generate lasting competitive pressure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extended-path solution method&lt;/strong&gt;: A numerical method (Fair and Taylor 1983; Adjemian and Juillard 2014) for solving nonlinear rational-expectations models with stochastic growth trends. In each period, agents are surprised by current shocks but expect future shocks to be zero on average (consistent with rational expectations). The method provides accurate solutions while accounting for model nonlinearities, and is combined with an inversion filter to form the likelihood function for Bayesian estimation. It is used here instead of standard log-linearization, which would be invalid under unbalanced growth dynamics.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Subsidy multiplier (present value)&lt;/strong&gt;: The ratio of the discounted cumulative GDP gain (or consumption gain) to the discounted cumulative subsidy expenditure over a given horizon, in the spirit of fiscal multipliers (Feve and Sahuc 2017; Leeper et al. 2017). In this paper, these multipliers measure the efficiency of redirecting carbon-tax revenues to abatement-sector subsidies. GDP multipliers exceed 2.0 through 2035 because the competition-enhancing effect of startup subsidies lowers abatement prices economy-wide, amplifying the direct expenditure impact.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Damage function&lt;/strong&gt;: The function Phi(T_t) = 1/(1 + aT_t^2) in the TFP equation, where T_t is the surface temperature anomaly and a is a calibrated damage parameter taken from DICE-2016R2. It captures the reduction in total factor productivity caused by climate change. The function implies damages of 4 percent of GDP per year by 2100 under the laissez-faire scenario (approximately 4 degrees C warming), and less than 1 percent under the below-2-degree scenario.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inversion filter&lt;/strong&gt;: A computationally efficient method for evaluating the likelihood function of a nonlinear dynamic model (Fair and Taylor 1983; Guerrieri and Iacoviello 2017; Atkinson et al. 2020). Instead of particle-filter simulation, it analytically recovers the sequence of structural shocks by inverting the observation equations for a given set of initial conditions and parameter values. Combined with the extended-path solution, it allows Bayesian estimation of the nonlinear E-DSGE model on world data.&lt;/p&gt;</description></item><item><title>Forecasting with Feedback</title><link>https://macropaperwarehouse.com/papers/forecasting-with-feedback/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/forecasting-with-feedback/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper develops a strategic model of point forecast production in environments where the forecast itself influences the outcome being predicted — what the authors call &amp;ldquo;forecasting with feedback.&amp;rdquo; The canonical example is Federal Reserve staff (Greenbook) inflation forecasts: these forecasts guide FOMC interest rate decisions, and those rate decisions in turn affect realized inflation. The central theoretical claim, proved formally, is that even a forecaster with purely quadratic (mean-squared-error) loss will optimally produce biased forecasts in such environments, provided there is some uncertainty about how strongly the decision maker (DM) will react to the forecast. This finding offers a third interpretation of observed forecast biases — beyond the two dominant explanations in the prior literature, namely forecaster irrationality and asymmetric loss functions.&lt;/p&gt;
&lt;p&gt;The model has three components. First, an outcome equation: y_{t+1} = theta_t + a_t + epsilon_{t+1}, where theta_t is a private signal (the state of the economy) observed only by the forecaster, a_t is the DM&amp;rsquo;s action, and epsilon_{t+1} is unforecastable noise. Second, a DM reaction function: a_t = x_t * [y_T - E(theta_t | f_t)], analogous to a Taylor rule, where y_T is a known target, and x_t is a strength-of-reaction multiplier drawn from a distribution with mean mu and variance tau^2; x_t is the DM&amp;rsquo;s private information. Third, the forecaster minimizes expected squared error, anticipating the DM&amp;rsquo;s endogenous response. The model is linear and closed-form solutions are derived.&lt;/p&gt;
&lt;p&gt;The key mechanism is a bias-variance tradeoff. Because the DM&amp;rsquo;s action responds to the forecast, the variance of the realized outcome itself becomes a function of the forecast. When the DM&amp;rsquo;s reaction strength x_t is uncertain (tau^2 &amp;gt; 0), this variance-of-outcome term is not trivially minimized by an unbiased forecast. The forecaster reduces outcome volatility by attenuating the sensitivity of the forecast to the state — shrinking the forecast slope toward zero relative to what an unbiased forecast would require — at the cost of introducing systematic bias. When tau^2 = 0 (no uncertainty about the DM&amp;rsquo;s reaction), the forecaster can perfectly anticipate and correct for the DM&amp;rsquo;s response, and the optimal forecast is unbiased. Feedback alone, without uncertainty, does not produce bias.&lt;/p&gt;
&lt;p&gt;The paper derives equilibrium forecasts in a Perfect Bayesian Equilibrium where the DM holds correct (rational) beliefs about the forecasting rule. Key analytical results include: (i) the equilibrium exists when tau^2 &amp;lt;= 1/4; (ii) the equilibrium conditional bias equals [(1 - sqrt(1 - 4*tau^2))/2] * (theta_t - y_T), which changes sign depending on whether the state is above or below the target — the forecaster gravitates toward the target; (iii) the Mincer-Zarnowitz (MZ) regression slope (the slope from regressing realized outcomes on forecasts) can be large and positive, close to zero, or even negative, depending on mu and tau^2; (iv) when mu = 1 (the DM on average fully closes the gap to the target), the equilibrium MZ slope is exactly zero for any tau^2 value.&lt;/p&gt;
&lt;p&gt;The paper motivates these results with two documented empirical patterns in Greenbook 4-quarter-ahead inflation forecasts from 1980q1 to 2019q4. First, using 40-quarter rolling windows, bias in Greenbook forecasts is persistent but sign-changing over time — a pattern consistent with the model&amp;rsquo;s prediction that the sign of bias tracks whether the state theta_t is above or below the inflation target y_T. Second, the MZ slope (from 40-quarter rolling-window regressions) hovers near unity in the mid-1980s through early 1990s, returns to unity by the late 1990s, then drops sharply to significantly negative territory by the mid-2000s, before becoming indistinguishable from zero in the final portion of the sample — a pattern consistent with the model&amp;rsquo;s prediction that the MZ slope shifts radically with changes in mu and tau^2. Both facts are computed using the last revision of the GDP deflator.&lt;/p&gt;
&lt;p&gt;The policy and methodological implications are significant. Standard forecast rationality tests (Mincer-Zarnowitz regressions, bias tests) are designed to detect irrationality or asymmetric loss, but in feedback environments these same test statistics can indicate &amp;ldquo;failure&amp;rdquo; even when the forecaster is fully rational under quadratic loss. Studies conducting rationality tests or estimating loss functions must either explicitly assume away feedback (and justify that assumption) or account for the feedback mechanism.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-identification"&gt;Q1. What is the identification strategy, and what are the main threats to identification?&lt;/h3&gt;
&lt;p&gt;The paper is primarily theoretical: it derives closed-form equilibrium forecasting rules and forecast statistics from first principles within a stylized game-theoretic model. There is no econometric identification exercise. The Greenbook evidence is descriptive and motivational — rolling-window bias estimates and MZ slope estimates are presented as stylized facts consistent with the theory, not as causal identification. The main caveat the authors themselves make is that the model is not claimed to be an exclusive or exhaustive explanation of the documented GB forecast patterns. Inflation forecasting is complex, and many other factors (learning, structural breaks, regime changes in monetary policy, data revisions) could contribute to the observed patterns. The authors explicitly disclaim any claim to exclusivity.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-core-mathematical-mechanism-and-how-does-uncertainty-play-a-necessary-role"&gt;Q2. What is the core mathematical mechanism, and how does uncertainty play a necessary role?&lt;/h3&gt;
&lt;p&gt;The forecaster&amp;rsquo;s MSE decomposes into a conditional variance term and a squared-bias term: MSE = Var[a*(f_t) | theta_t] + bias^2(f_t | theta_t) + sigma^2. The critical insight is that when x_t (the reaction-strength multiplier) is uncertain, the variance of the DM&amp;rsquo;s action — and hence of the outcome — depends on the level of the forecast itself. Specifically, Var[a*(f_t) | theta_t] = tau^2 * (y_T - f_t/c + b/c)^2. So choosing a larger or smaller forecast changes not just the bias term but also the variance term. The optimal resolution of this tradeoff requires an attenuated (biased) forecast slope. When tau^2 = 0 (no uncertainty), the variance term vanishes entirely and the forecaster can correct for feedback in full by solving a fixed-point problem, producing an unbiased forecast. The paper explicitly proves (taking limits as tau^2 to 0 in the bias and MZ slope formulas) that both return to zero and one respectively, confirming that uncertainty is a necessary condition for bias.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-equilibrium-concept-and-what-are-its-properties"&gt;Q3. What is the equilibrium concept and what are its properties?&lt;/h3&gt;
&lt;p&gt;The equilibrium is a linear Perfect Bayesian Equilibrium (PBE). The DM conjectures that the forecast is a linear function f_t = b + c*theta_t, uses that conjecture to form expectations E(theta_t | f_t) = (f_t - b)/c, and chooses her action optimally. Equilibrium requires that the DM&amp;rsquo;s conjectured intercept and slope (b, c) coincide with those actually used by the forecaster. The paper shows (Corollary 1) that such a linear PBE exists when tau^2 &amp;lt;= 1/4, and that the equilibrium is fully revealing — the DM can learn the true state theta_t from the forecast because the forecast is a one-to-one function of the state. Two linear equilibria exist: the paper focuses on the Pareto-preferred one (lower forecaster loss, lower absolute bias), which is also the one whose limit as tau^2 approaches 0 corresponds to the natural optimal forecast.&lt;/p&gt;
&lt;h3 id="q4-what-sign-and-magnitude-patterns-does-the-equilibrium-bias-exhibit"&gt;Q4. What sign and magnitude patterns does the equilibrium bias exhibit?&lt;/h3&gt;
&lt;p&gt;From Corollary 2(a), the conditional equilibrium bias is: E(y_{t+1} - f_t^dagger | theta_t) = [(1 - sqrt(1 - 4&lt;em&gt;tau^2)) / 2] * (theta_t - y_T). The multiplier (1 - sqrt(1 - 4&lt;/em&gt;tau^2))/2 is always positive (for tau^2 in (0, 1/4]), so the sign of the bias is determined entirely by the sign of (theta_t - y_T). When theta_t &amp;gt; y_T (state above target), bias is positive — the forecaster underpredicts, shrinking the forecast toward the target. When theta_t &amp;lt; y_T, bias is negative — the forecaster overpredicts, again gravitating toward the target. This sign-change mechanism, driven by changing economic conditions relative to a fixed target, is cited as consistent with the persistent but sign-changing bias observed in Greenbook inflation forecasts from 1980 to 2019.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-model-predict-about-the-mincer-zarnowitz-slope-and-how-variable-can-it-be"&gt;Q5. What does the model predict about the Mincer-Zarnowitz slope, and how variable can it be?&lt;/h3&gt;
&lt;p&gt;From Corollary 2(b), the MZ slope in equilibrium is a highly nonlinear function of mu and tau^2. Figure 3 in the paper (discussed in the text) shows that the slope can be large and positive, positive but close to zero, negative, or even very steeply negative, for different combinations of mu and tau^2. A key special case: when mu = 1 (DM fully closes the gap to target on average), E(y_{t+1} | f_t^dagger) = y_T for all values of the forecast, giving an MZ slope of exactly zero and intercept equal to y_T. The authors note that when mu is close to 1 and tau^2 is small, even small deviations of mu from unity can produce large positive or negative MZ slopes. The model can thus account for the dramatic shift in the GB MZ slope documented in the paper — from around unity in the 1980s-1990s, to significantly negative territory in the mid-2000s, to approximately zero thereafter.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-relationship-between-the-dms-reaction-function-and-the-taylor-rule-and-how-is-it-microfounded"&gt;Q6. What is the relationship between the DM&amp;rsquo;s reaction function and the Taylor rule, and how is it microfounded?&lt;/h3&gt;
&lt;p&gt;The DM&amp;rsquo;s reaction function is a_t* = x_t * [y_T - E(theta_t | f_t)], directly analogous in spirit to a Taylor rule (Taylor, 1993). Online Appendix A provides a formal microfoundation: if the DM minimizes a quadratic loss in (y_{t+1} - y_T)^2 plus a quadratic adjustment cost w_t * a_t^2 — where w_t is a private, randomly drawn adjustment cost parameter — then the optimal action is precisely a_t* = x_t * [y_T - E(theta_t | f_t)] with x_t = 1/(1 + w_t). This microfoundation connects the model to the literature on central bank optimal control and provides a rational justification for the reaction function structure used throughout the paper.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-the-crawford-sobel-1982-cheap-talk-model"&gt;Q7. How does this paper relate to and differ from the Crawford-Sobel (1982) cheap talk model?&lt;/h3&gt;
&lt;p&gt;The paper borrows the sender-receiver communication game structure from Crawford and Sobel (1982), with the forecaster as sender and the DM as receiver. However, it departs in two important ways. First, in Crawford-Sobel, the sender&amp;rsquo;s payoff depends only on the state and the action, not directly on the message (the forecast). In this paper, the forecast enters the forecaster&amp;rsquo;s loss function directly through the outcome equation (y = theta + a + epsilon, and the forecast determines a which determines y which enters the loss), making it a model of &amp;lsquo;costly talk&amp;rsquo; in the sense of Kartik, Ottaviani, and Squintani (2007). Second, in standard communication games the realized outcome is exogenous — the DM&amp;rsquo;s action affects only her own payoff but not the variable being forecast. Here, the DM&amp;rsquo;s action causally determines the realized outcome that the forecaster was trying to predict. This feedback causality is absent in the standard setup and is the source of the paper&amp;rsquo;s novel results.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-bernanke-and-woodford-1997"&gt;Q8. How does this paper relate to Bernanke and Woodford (1997)?&lt;/h3&gt;
&lt;p&gt;Bernanke and Woodford (1997) also study professional inflation forecasts and monetary policy in a rational expectations equilibrium framework, and raise the question of whether an informative equilibrium exists — concluding it may not. This paper differs in three respects: it assumes the forecaster has private information (state theta_t) that the DM cannot directly observe; it works in an environment with uncertainty about the DM&amp;rsquo;s reaction (x_t is random); and rather than focusing on equilibrium existence, it derives the statistical properties of equilibrium forecasts — the bias formula, MZ regression coefficients — which Bernanke and Woodford do not. The authors describe their work as providing &amp;rsquo;the first formal treatment of the statistical properties of forecasts&amp;rsquo; in feedback environments.&lt;/p&gt;
&lt;h3 id="q9-what-heterogeneity-and-parameter-sensitivity-is-documented"&gt;Q9. What heterogeneity and parameter sensitivity is documented?&lt;/h3&gt;
&lt;p&gt;The paper documents sensitivity of forecast properties to mu (mean policy reaction strength) and tau^2 (variance of policy reaction strength). The DM&amp;rsquo;s average aggressiveness mu affects both the sign and magnitude of the MZ slope: for cautious DMs (mu near 0.1), the equilibrium MZ slope is relatively close to unity; for aggressive DMs (mu near 1), the slope can flatten toward zero; for moderate but increasing mu (with tau^2 above a threshold of approximately 0.05), the slope flattens monotonically. A higher tau^2 at given mu generally attenuates the slope toward zero, but the relationship is nonlinear. When mu is precisely one, the MZ slope is exactly zero regardless of tau^2. The equilibrium bias magnitude scales with [(1 - sqrt(1 - 4*tau^2))/2], which increases in tau^2. The sign of bias is determined by the direction of (theta_t - y_T). The paper does not present cross-sectional or time-series panel heterogeneity — the parametric sensitivity analysis in Figure 3 constitutes the heterogeneity exercise.&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-are-run-for-the-greenbook-empirical-patterns"&gt;Q10. What robustness checks are run for the Greenbook empirical patterns?&lt;/h3&gt;
&lt;p&gt;The authors state (in a footnote) that the documented patterns — persistent but sign-changing bias in 4-quarter-ahead GB inflation forecasts from 1980q1 to 2019q4 — are robust to using the second release of the GDP deflator rather than the last release. The main results use the last release. The choice of 40-quarter (10-year) rolling window is applied uniformly for both the bias plot and the MZ slope plot. No additional robustness checks (alternative window lengths, alternative forecast horizons, formal structural break tests) are explicitly documented in the paper, though the authors cite Rossi and Sekhposyan (2016), who use formal rationality tests and confirm that GB forecast rationality breaks down around 2005 — consistent with the pattern the authors document via the rolling MZ slope.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-model-say-about-the-forecasters-inability-to-commit-and-could-commitment-help"&gt;Q11. What does the model say about the forecaster&amp;rsquo;s inability to commit, and could commitment help?&lt;/h3&gt;
&lt;p&gt;In the baseline model, the forecaster cannot commit to a fixed forecasting rule ex ante because the state theta_t is not directly observable by the DM. The authors note in Section 3.3 that modeling forecasters with commitment is a straightforward extension, and that commitment can actually increase forecaster welfare in equilibrium. However, this extension is not formally developed in the paper. The intuition is that if the forecaster could credibly commit to a more informative forecast rule, the DM could react more precisely, reducing the variance of outcomes; but without commitment, the strategic equilibrium involves an attenuated (biased) forecast.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-implications-for-forecast-rationality-tests-and-loss-function-estimation"&gt;Q12. What are the implications for forecast rationality tests and loss function estimation?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s central methodological warning is that standard forecast rationality tests (MZ regression tests for zero intercept and unit slope; bias tests) and loss function estimation exercises are contaminated in environments with policy feedback. If feedback is present and x_t is uncertain, a fully rational forecaster with quadratic loss will produce forecasts that fail standard rationality tests — showing nonzero bias, non-unit MZ slopes (potentially even negative), and forecast errors correlated with the forecaster&amp;rsquo;s own information. Researchers conducting such tests must either: (a) explicitly assume no feedback applies (and justify this assumption in their specific application), or (b) carefully model the feedback mechanism and account for it. Studies that interpret GB forecast irrationality (e.g., Rossi and Sekhposyan 2016) or asymmetric loss (e.g., Capistran 2008) as the explanation for observed GB forecast properties may be confounded by the feedback mechanism identified in this paper.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-conditions-under-which-a-linear-equilibrium-does-or-does-not-exist"&gt;Q13. What are the conditions under which a linear equilibrium does or does not exist?&lt;/h3&gt;
&lt;p&gt;From Corollary 1 and Remark 3 following it: a linear PBE exists if and only if tau^2 &amp;lt;= 1/4. When tau^2 &amp;gt; 1/4, the forecaster always wants to attenuate the slope more than the DM expects, so no fixed-point equilibrium in linear strategies exists. The paper also notes a sufficient condition for equilibrium existence: if the support of x_t is contained in [0, 1] (the DM never overreacts and never underreacts by more than half), then tau^2 &amp;lt;= 1/4 is automatically satisfied and an equilibrium always exists. Two linear equilibria exist when tau^2 &amp;lt;= 1/4, but the paper focuses on the Pareto-preferred one, which has lower forecaster loss, lower absolute bias, and a natural limiting behavior as tau^2 approaches 0.&lt;/p&gt;
&lt;h3 id="q14-what-scope-conditions-limit-the-applicability-of-the-results"&gt;Q14. What scope conditions limit the applicability of the results?&lt;/h3&gt;
&lt;p&gt;Several scope conditions are made explicit: (1) The outcome equation is linear; nonlinear outcome determination would change quantitative results but the feedback mechanism would persist qualitatively. (2) The model is a single-period (point-in-time) game, not a multi-period learning model — it does not analyze how beliefs about mu and tau^2 evolve over time. (3) The independence assumption between x_t and theta_t is a benchmark; if policy aggressiveness varies with economic conditions, additional effects arise. (4) The focus on linear equilibria rules out non-linear forecasting strategies. (5) The results apply to unconditional forecasts (where the forecaster anticipates the DM&amp;rsquo;s response); conditional forecasts (conditioned on a pre-specified action) behave differently. (6) The empirical Greenbook evidence is illustrative, not a formal test of the model — the authors explicitly state they do not claim their model provides an exclusive explanation of GB forecast properties.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Forecasting with feedback&lt;/strong&gt;: A forecasting environment in which the DM&amp;rsquo;s action — taken in response to the forecast — causally affects the realized value of the variable being forecast, so that the forecast influences its own target outcome. Distinguished from no-feedback environments (e.g., weather forecasting) where decisions made on the basis of the forecast do not affect the outcome.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Unconditional forecast&lt;/strong&gt;: A forecast that anticipates and factors in the expected response of the decision maker to the forecast itself, rather than being conditioned on a pre-specified (potentially counterfactual) action. The paper&amp;rsquo;s model produces unconditional forecasts; conditional forecasts (conditioned on a given policy path) are a distinct and narrower concept.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bias-variance tradeoff (in feedback forecasting)&lt;/strong&gt;: The tradeoff that arises when the DM&amp;rsquo;s reaction to the forecast is uncertain: a less informative (attenuated) forecast reduces the variance of the outcome (by inducing a less volatile policy action) but introduces systematic bias. The optimal forecast under quadratic loss resolves this tradeoff by attenuating the forecast slope below what an unbiased forecast would require, producing an optimally biased forecast.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reaction function (DM&amp;rsquo;s)&lt;/strong&gt;: The rule by which the decision maker translates a forecast into a policy action: a_t* = x_t * [y_T - E(theta_t | f_t)], analogous to a Taylor rule. The multiplier x_t captures the strength of the policy response and is drawn from a distribution with mean mu and variance tau^2; it is the DM&amp;rsquo;s private information and a key source of the forecaster&amp;rsquo;s uncertainty.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mincer-Zarnowitz (MZ) regression&lt;/strong&gt;: The linear regression of the realized outcome on the forecast: y_{t+1} = alpha + beta * f_t + error. Under the canonical null of rational forecasting with quadratic loss and no feedback, the intercept alpha should be zero and the slope beta should be one. The paper shows that under optimal forecasting with feedback, alpha and beta can take a wide range of values, including negative beta, even when the forecaster is rational.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Equilibrium forecast slope (c-dagger)&lt;/strong&gt;: The slope of the linear forecasting rule in Perfect Bayesian Equilibrium, given by c^dagger = (1/2) - mu + sqrt(1 - 4*tau^2)/2. This slope is less than one and can be negative depending on mu and tau^2, reflecting the attenuation of the forecast toward the policy target that arises from the bias-variance tradeoff under uncertain DM reactions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Greenbook (GB) inflation forecasts&lt;/strong&gt;: Inflation forecasts produced by Federal Reserve staff (now called Tealbook forecasts), used as empirical motivation in the paper. The paper documents two stylized facts for 4-quarter-ahead GB forecasts from 1980q1 to 2019q4: (i) persistent but sign-changing bias in rolling 40-quarter windows, and (ii) a dramatic shift in the rolling MZ slope from approximately unity in the 1980s-1990s to significantly negative in the mid-2000s and approximately zero in the final part of the sample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy feedback (as a confound for rationality tests)&lt;/strong&gt;: The paper&amp;rsquo;s use of this term to describe the mechanism by which the presence of feedback invalidates the standard interpretation of forecast rationality test outcomes: a forecaster who is fully rational (quadratic loss, no private agenda) and operating in a feedback environment will systematically produce forecasts that fail standard MZ-based rationality tests, not because of irrationality or asymmetric loss, but because of the optimal bias-variance tradeoff induced by uncertain policy reactions.&lt;/p&gt;</description></item><item><title>General Equilibrium Effects in Space: Theory and Measurement</title><link>https://macropaperwarehouse.com/papers/general-equilibrium-effects-in-space-theory-and-measurement/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/general-equilibrium-effects-in-space-theory-and-measurement/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;How do international trade shocks propagate through spatially connected regional labor markets, and how large are the general equilibrium effects that standard shift-share specifications miss? Adão, Arkolakis, and Esposito address this question by extending shift-share empirical designs to incorporate general equilibrium (GE) effects arising from spatial links between markets. Their motivation is that the difference-in-difference logic of standard shift-share regressions recovers only the differential response of treated versus control regions, not the level response that includes indirect (spillover) effects propagating through trade, labor supply, and agglomeration links. Ignoring these indirect effects biases estimates of trade shocks&amp;rsquo; aggregate labor market consequences.&lt;/p&gt;
&lt;p&gt;The theoretical framework is a multi-sector general equilibrium spatial model with N markets linked through three channels: (i) gravity-type trade demand, (ii) endogenous labor supply that depends on wages and price indices in all markets, and (iii) local labor productivity that depends on employment (agglomeration). The key theoretical result is that wage and employment responses to trade shocks decompose into two shift-share exposure vectors — a revenue exposure (proportional to the ADH import penetration measure, weighted by sectoral employment shares) and a consumption cost exposure (weighted by sectoral spending shares) — multiplied by bilateral reduced-form elasticity matrices (βij and φij). These elasticities are sufficient statistics for GE aggregation and can be expressed as a series expansion of the &amp;ldquo;spatial links&amp;rdquo; matrix, which is itself a function of trade demand substitution, labor supply substitution, and agglomeration elasticities. When demand substitution dominates (gross substitution property holds), indirect effects reinforce direct effects: a negative revenue shock in one CZ reduces demand for goods from other CZs, propagating wage and employment losses outward.&lt;/p&gt;
&lt;p&gt;The authors apply the framework to the China shock, using 722 U.S. Commuting Zones (CZs) over 1990–2007, following Autor, Dorn, and Hanson (2013) (ADH). The revenue exposure measure is identical to the ADH instrumental variable (employment-share-weighted Chinese export growth to non-U.S. developed countries); the consumption exposure is analogously constructed using sectoral spending shares from input-output tables. Structural parameters are estimated using a Model-implied Optimal IV (MOIV) two-step GMM estimator derived from Chamberlain (1987).&lt;/p&gt;
&lt;p&gt;Main quantitative findings: (1) In a simple extension of ADH, the indirect revenue spillover effect on neighboring CZs is roughly three times larger in magnitude than the direct effect of a CZ&amp;rsquo;s own import competition exposure — an increase of $1,000 in Chinese imports per U.S. worker in nearby CZs is associated with 1.3 log-point lower employment growth and 1.0 log-point lower wage growth in a given CZ. (2) Consumption cost shifts (cheaper imports) have no statistically significant direct or indirect effect on employment or wages, consistent with a weak price elasticity of labor supply relative to the wage elasticity. (3) Structural parameter estimates yield: labor productivity–employment elasticity ψ = 0.56 (agglomeration), labor supply–wage elasticity φw = 2.11, labor supply–price elasticity φp = −1.36, trade elasticity ε = 3.94. (4) In GE aggregation, the China shock reduced average U.S. CZ wages by approximately 4.0 log-points and employment by approximately 2.8 log-points between 1990 and 2007, with the indirect revenue channel (−4.24 log-points for wages, −4.95 log-points for employment) dominating the direct revenue effect (−0.81 and −1.94 respectively) and being partially offset by positive consumption cost effects (+0.98 wages, +3.18 employment). Average real wages rose by 0.16 log-points on net, but 39% of CZs experienced real wage declines. Standard deviations of responses were 1.30 for wages, 3.31 for employment, and 1.75 for real wages, indicating large cross-CZ heterogeneity. (5) Model fit: the baseline estimated model yields fit coefficients close to 1 (0.67 for wages, 0.90 for employment), whereas quantitative models calibrated with Ricardian/standard parameters yield fit coefficients of 3.56 to 10.42, indicating their predicted responses are too small by factors of 4–10. Simple aggregation of the ADH specification implies employment losses of only 1.5 log-points — less than half the authors&amp;rsquo; baseline estimate.&lt;/p&gt;
&lt;p&gt;The key mechanism driving the amplification is strong agglomeration (ψ ≈ 0.56), which roughly doubles typical calibrations from Krugman-type models and is absent in Ricardian frameworks. Demand-side trade links propagate revenue shocks across CZs with similar sectoral composition and trade partners. The policy implication is that analyses of trade shocks using standard shift-share regressions — which absorb common indirect effects in time fixed effects — systematically understate aggregate employment and wage losses.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Identification rests on the same orthogonality condition used by ADH and Kovak (2013): observed shock exposure (revenue and consumption shift-share measures) is mean-independent of unobserved residuals. This is implied by independence between the observed Chinese export shock and unobserved trade cost shocks, given the initial trade matrix. The authors use the ADH instrument (Chinese export growth to non-U.S. developed countries) to construct exogenous sectoral shifts, exploiting cross-CZ variation in initial industry composition. The main threats are: (i) unobserved shocks correlated with pre-existing industry composition (e.g., concurrent automation), addressed by controlling for lagged population growth (following Greenland et al. 2019) and the full ADH control set; (ii) spatial correlation of residuals, addressed by clustering standard errors at the state level and by robustness using the inference procedure in Adão et al. (2019); (iii) simultaneity, since the MOIV estimator instruments the non-linear functions of shock exposure with model-implied moment functions that are functions of the observed shifts only.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-shift-share-exposure-measures-and-how-do-they-differ"&gt;Q2. What are the two shift-share exposure measures and how do they differ?&lt;/h3&gt;
&lt;p&gt;The revenue exposure (IPW) is the standard ADH shift-share variable: the product of Chinese export growth to other developed countries and the CZ&amp;rsquo;s initial employment share in each sector, summed across sectors. It captures the shock to the demand for a CZ&amp;rsquo;s goods. The consumption cost exposure (IPC) is an analogous variable where the share is the CZ&amp;rsquo;s sectoral spending share (including intermediate inputs, constructed using national input-output tables interacted with regional employment shares) rather than employment share. It captures the shock to the CZ&amp;rsquo;s cost of living and input costs. The two measures have a spatial correlation of 0.34. Standard deviations across CZs are 2.52 for IPW and 1.22 for IPC.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q3. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;Three spatial channels determine GE reduced-form elasticities: (1) trade demand links — markets with similar sectoral composition and trade partners are closer substitutes, so a revenue shock in one CZ propagates negatively to CZs competing for the same export destinations; (2) labor supply links — employment responses in one CZ to wage/price changes in another, captured through migration (parametrized by bilateral birth-state shares) and the local wage and price elasticities of labor supply; (3) agglomeration — local labor productivity responds positively to local employment, amplifying both direct and indirect effects. Empirically, the authors distinguish these by estimating separate parameters (ψ for agglomeration, φw for wage elasticity of labor supply, φp for price elasticity, φm for migration links, ε for trade elasticity), with identification coming from cross-CZ heterogeneity in bilateral trade shares, sector specialization, and migration shares. The weak IPC effect (statistically insignificant) points to a small φp, while the large employment and wage responses to IPW point to large φw and ψ.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-estimated-structural-parameters-and-how-do-they-compare-to-existing-literature"&gt;Q4. What are the estimated structural parameters and how do they compare to existing literature?&lt;/h3&gt;
&lt;p&gt;Panel A estimates (without migration): ψ = 0.56 (s.e. 0.07), φw = 2.11 (s.e. 0.25), φp = −1.36 (s.e. 0.24), ε = 3.94 (s.e. 0.41). Panel B (with migration): nearly identical point estimates but standard errors two to five times larger due to high collinearity of bilateral migration and trade shares; φm = −0.06 (s.e. 0.05), not statistically significant. The agglomeration elasticity ψ = 0.56 is roughly twice the Krugman (1980) implied value (~0.2) used by Monte et al. (2018) and far above zero (used in Ricardian frameworks by Galle et al. 2017, Caliendo et al. 2018, 2019). It is closer to Kline and Moretti (2014)&amp;rsquo;s estimate of ~0.4 from regional demand shocks. The labor supply elasticity φw = 2.11 is three times the median micro-estimate in Chetty et al. (2013) and is consistent with aggregate employment responses. The trade elasticity ε ≈ 4 is within standard literature ranges.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-in-spatial-effects-is-documented"&gt;Q5. What heterogeneity in spatial effects is documented?&lt;/h3&gt;
&lt;p&gt;There is substantial heterogeneity in both direct and indirect reduced-form elasticities across CZs. For revenue shifts, the 10th/50th/90th percentiles of direct wage elasticities are 0.44/0.67/1.67, and for employment 0.92/1.46/3.97. For indirect effects, median values are 0.002 (wages) and 0.003 (employment), but the 90th percentile is 0.021 and 0.039 respectively. The simple gravity proxy zij (inverse distance weighted by population) explains only a small fraction of variation in indirect effects; instead, the elements of the full spatial links matrix (bilateral revenue shares yij and trade demand substitutability χij) explain roughly 50% of variation in indirect effects across CZ pairs. Both manufacturing and non-manufacturing employment show significant indirect effects; wage responses are mainly driven by the non-manufacturing sector (consistent with ADH). 39% of CZs experienced real wage declines despite a small average real wage gain.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;For the simple ADH extension (Table 1): (i) varying the distance decay parameter δ ∈ (1,8); (ii) using CZ size vs. no size weighting in zij; (iii) restricting to same-state CZs for indirect effects; (iv) weighting CZs by 1990 population; (v) using the Adão et al. (2019) inference procedure; (vi) alternative spending share constructions. For the structural estimation: (i) allowing for trade imbalances (following Dekle et al. 2007); (ii) calibrating migration links from external estimates; (iii) alternative numeraire for labor supply homogeneity (national vs. world price index). In all cases, indirect effects remain negative and significant, and reduced-form elasticities are highly correlated with baseline estimates. Counterfactual employment losses range from −0.5 to −5.4 log-points depending on the labor supply normalization and migration specification, with average wage decline remaining close to 4 log-points across specifications. The NTR gap (Pierce and Schott 2016) as the sector-level shifter also yields qualitatively similar results.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-paper-evaluate-the-fit-of-quantitative-spatial-models"&gt;Q7. How does the paper evaluate the fit of quantitative spatial models?&lt;/h3&gt;
&lt;p&gt;The authors propose regressing actual changes in CZ employment/wages on model-predicted responses (equation 39) and checking whether the slope coefficient ρ is close to 1. A coefficient much greater than 1 means the model&amp;rsquo;s predicted responses are too small relative to actual cross-CZ variation. The baseline structural estimates yield fit coefficients of 0.67 (wages) and 0.90 (employment) — close to 1. Alternative calibrations from quantitative frameworks yield coefficients of 3.56–10.42 for wages and 6.60–10.42 for employment, indicating those models underpredict differential responses by factors of 4–10. The main driver is weak agglomeration forces: setting ψ = 0 (Ricardian) vs. ψ = 0.56 (baseline) dramatically degrades fit. Setting φw = −φp (labor supply responding to real wages only, as in Caliendo et al. 2019) makes employment fit estimates very imprecise because the consumption price channel becomes too strong relative to its empirical counterpart.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-quantitative-ge-impact-of-the-china-shock-on-average-us-cz-wages-and-employment-and-how-does-it-decompose"&gt;Q8. What is the quantitative GE impact of the China shock on average U.S. CZ wages and employment, and how does it decompose?&lt;/h3&gt;
&lt;p&gt;Over 1990–2007: average wage fell by 3.98 log-points (s.d. 1.30), average employment fell by 2.78 log-points (s.d. 3.31), average real wage rose by 0.16 log-points (s.d. 1.75). Decomposition of wage change: direct revenue effect −0.81 (s.d. 1.79), direct consumption cost effect +0.98 (s.d. 1.36), indirect revenue effect −4.24 (s.d. 1.71), indirect consumption cost effect +0.09 (s.d. 1.18). The indirect revenue channel dominates; consumption gains are not large enough to offset revenue losses. For real wages, the main components are: terms-of-trade loss from wage decline (−0.98, s.d. 2.53), productivity/efficiency gains (+3.14, approximately), and consumption cost gains. Most impact occurred in the 2000–2007 sub-period after China&amp;rsquo;s WTO accession.&lt;/p&gt;
&lt;h3 id="q9-how-do-these-ge-estimates-compare-to-estimates-from-the-existing-literature"&gt;Q9. How do these GE estimates compare to estimates from the existing literature?&lt;/h3&gt;
&lt;p&gt;Simple aggregation of the ADH specification (ignoring GE indirect effects) implies average wage losses of 1.17 log-points and employment losses of 1.50 log-points — less than half the authors&amp;rsquo; GE estimates. Including intuitive distance-weighted indirect effects (ADH extension in Table 1 column 3) brings employment estimates closer (−4.51 log-points) but with correlation below 0.5 with baseline cross-CZ heterogeneity predictions. Quantitative spatial models calibrated with standard parameters (Ricardian, weak agglomeration) generate average responses near zero and are often uncorrelated with actual CZ outcomes. The key reason quantitative models underperform is that they specify agglomeration forces as too weak (ψ ≈ 0 versus the estimated 0.56) and labor supply sensitivity to import prices as too strong relative to wage sensitivity.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-role-of-the-consumption-cost-ipc-channel-and-why-does-it-matter-less-than-the-revenue-channel"&gt;Q10. What is the role of the consumption cost (IPC) channel and why does it matter less than the revenue channel?&lt;/h3&gt;
&lt;p&gt;The IPC captures the welfare gain from cheaper Chinese imports: as Chinese productivity rises, import prices fall, increasing real purchasing power and potentially stimulating labor supply. However, the estimated labor supply price elasticity (φp = −1.36) is substantially smaller in absolute value than the wage elasticity (φw = 2.11), so the positive employment and wage response to lower import prices is weaker than the negative response to falling demand for local output. Empirically, both the direct and indirect effects of IPC are statistically insignificant in the simple ADH extension (Table 1, columns 2 and 4), consistent with weak φp. The structural estimation exploits all channels to pin down φp precisely. Input-output linkages (CZs using inputs from sectors with stronger Chinese export growth) are incorporated in IPC and are also found to have no significant employment effect, consistent with Pierce and Schott (2016) and Acemoglu et al. (2016).&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-connect-to-the-shift-share-and-market-access-literatures"&gt;Q11. How does the paper connect to the shift-share and market access literatures?&lt;/h3&gt;
&lt;p&gt;The paper generalizes standard shift-share designs (Bartik 1991, Blanchard and Katz 1992, ADH 2013, Kovak 2013) in two ways: it adds a consumption cost shift-share (spending shares instead of employment shares) and it adds indirect exposure from other CZs&amp;rsquo; shift-share measures, weighted by model-implied bilateral reduced-form elasticities. Unlike standard designs, time fixed effects in the authors&amp;rsquo; estimating equation absorb only the mean unobserved shock, not any GE indirect effects (since the latter are heterogeneous across CZ pairs). The paper connects to the market access approach (Redding and Venables 2004; Donaldson and Hornbeck 2016) by showing that the authors&amp;rsquo; revenue and consumption exposure measures are partial-equilibrium versions of producer and consumer market access, holding wages and employment constant. The key advantage is that the authors&amp;rsquo; measures can be constructed from initial-equilibrium data without solving the full GE model.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper implies that trade shock analyses ignoring GE spillovers substantially understate aggregate employment and wage losses for U.S. workers. The gross substitution condition (trade demand links dominating labor supply links) is required for indirect effects to reinforce rather than attenuate direct effects; this is consistent with the empirical evidence but could fail in settings with very mobile labor markets. The real wage calculation shows that, on average, cheaper imports provide a small net welfare gain (+0.16 log-points), but 39% of CZs experienced net real wage losses, pointing to substantial distributional consequences within the U.S. The framework&amp;rsquo;s scope is first-order (linearization around initial equilibrium), so it is a good approximation for moderate shocks; large shocks require integrating over the adjustment path. The methodology is applicable beyond the China shock to any trade policy with measurable regional exposure variation.&lt;/p&gt;
&lt;h3 id="q13-what-is-the-moiv-estimator-and-why-is-it-efficient"&gt;Q13. What is the MOIV estimator and why is it efficient?&lt;/h3&gt;
&lt;p&gt;The Model-implied Optimal IV (MOIV) is a two-step feasible implementation of the Chamberlain (1987) efficient GMM estimator. The class of consistent GMM estimators for the spatial link parameters θ = (φw, φp, φm, ψ, ε) differs only in how they weight the observed exposure of different markets. The optimal weighting function H*i assigns more weight to markets whose reduced-form elasticities (βij and φij) are most sensitive to changes in the parameter being estimated — i.e., markets that provide the most information about a given parameter. In step 1, an arbitrary initial θ0 is used to obtain a consistent but non-optimal first-stage estimate. In step 2, the consistent estimate is used to compute the optimal instrument, and a second-stage GMM is run. The MOIV is asymptotically equivalent to the Chamberlain efficient estimator. The paper&amp;rsquo;s contribution is to derive the optimal moment conditions for a flexible spatial GE model with non-linear parameter-dependent elasticities.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Spatial Links Matrix&lt;/strong&gt;: The Jacobian of the excess labor demand system with respect to wages, denoted γ-bar, summarizing the combined effect of trade demand substitution (how wage changes in one market shift demand from other markets) and supply substitution (how wage changes affect labor supply across markets, amplified by agglomeration). It governs the propagation of partial equilibrium excess demand shifts to general equilibrium wage and employment responses, and determines the sign and heterogeneity of indirect effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bilateral Reduced-Form Elasticity&lt;/strong&gt;: The element βij (for wages) or φij (for employment) measuring how much market i&amp;rsquo;s outcome responds to a unit shift in market j&amp;rsquo;s excess labor demand, after all GE adjustment rounds. It is a series expansion of the spatial links matrix and is larger for market pairs with stronger bilateral or third-market spatial connections. These elasticities are sufficient statistics for aggregating regional shock exposures to compute GE impact.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Revenue Exposure (IPW)&lt;/strong&gt;: The shift-share variable capturing a CZ&amp;rsquo;s partial equilibrium revenue shift from a foreign productivity shock: the employment-share-weighted average of sectoral export growth shocks. Identical to the ADH instrument. Measures how much a CZ&amp;rsquo;s producer revenues (and thus labor demand) fall when Chinese costs decline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption Cost Exposure (IPC)&lt;/strong&gt;: A novel shift-share variable capturing the partial equilibrium consumption cost shift: the spending-share-weighted average of sectoral export growth shocks, constructed using national input-output tables interacted with regional employment. Measures how much cheaper Chinese imports reduce the cost of living and inputs in a CZ, with a positive effect on real wages and labor supply.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model-Implied Optimal IV (MOIV)&lt;/strong&gt;: A two-step feasible GMM estimator that achieves the Chamberlain (1987) efficiency bound for estimating the vector of structural spatial link parameters θ. In the first step any consistent estimator is used; in the second step the first-step estimates are used to compute the optimal moment function — which places more weight on CZs whose reduced-form elasticities are most sensitive to changes in the parameter being estimated — and a second-stage GMM yields the efficient estimate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gross Substitution Property&lt;/strong&gt;: A condition on the spatial links matrix (γij &amp;lt; 0 for all off-diagonal pairs) under which all bilateral reduced-form elasticities βij are positive, so indirect effects of excess demand shifts always reinforce direct effects. The condition is satisfied when trade demand substitution dominates labor supply substitution in the spatial links matrix. Empirically supported for U.S. CZs: negative revenue shocks spread negatively to other CZs rather than triggering offsetting employment inflows.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agglomeration Elasticity (ψ)&lt;/strong&gt;: The elasticity of local labor productivity to local employment in the production function, governing the feedback of employment changes on production costs and thus on excess labor demand. The authors estimate ψ = 0.56 for U.S. CZs — roughly twice the Krugman (1980) value and far above the zero assumed in Ricardian frameworks — and show it is the key parameter that amplifies both direct and indirect responses to trade shocks and determines model fit.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous Fixed Effect&lt;/strong&gt;: A common component of GE indirect effects that arises when spatial links are identical across markets (Corollary 2). In this special case all indirect effects collapse to a common term absorbed by time fixed effects in standard regressions, making those regressions unable to separately identify the indirect effect from aggregate time trends. In the general case with heterogeneous spatial links, indirect effects differ across CZ pairs and are not absorbed by time fixed effects.&lt;/p&gt;</description></item><item><title>How Costly Are Cartels?</title><link>https://macropaperwarehouse.com/papers/how-costly-are-cartels/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/how-costly-are-cartels/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Moreau and Panon ask how much cartels cost the aggregate economy — in terms of both total factor productivity and welfare — and find the losses are considerably larger than the received wisdom from Harberger (1954) would suggest. The paper&amp;rsquo;s motivation is the mounting evidence that markups are large and growing, combined with a near-total absence of macroeconomic quantification of collusion as one micro-origin of those markups.&lt;/p&gt;
&lt;p&gt;The empirical foundation is an original firm-level database for France covering the period 1994–2007, assembled by scraping all written decisions of the French Competition Authority (ADLC). The final dataset contains 174 cartels and more than 1,000 firms before matching. These cartel records are merged to administrative balance-sheet and income-statement data covering the universe of French firms (BRN and RSI regimes). Key facts documented: average cartel duration is 4.5 years (median 3 years); average cartel size is 6.3 members (median 4); cartels are prevalent across construction, manufacturing, wholesale, retail, and transportation. Crucially, cartel members are empirically shown to be dramatically larger than non-members even within narrowly defined 4-digit industries — roughly 1,900% more sales, a market share premium of 4 percentage points, 1,150% more employment, and 37% higher labor productivity. Firms within a cartel are also substantially more homogeneous in productivity than the overall within-industry distribution: the interquartile productivity ratio across cartel members is only 1.4-to-1, versus 2-to-1 across all non-cartel firms in the same industry.&lt;/p&gt;
&lt;p&gt;The theoretical framework extends the static heterogeneous-firm oligopoly model of Atkeson and Burstein (2008) by introducing collusion microfounded via the cross-ownership framework of O&amp;rsquo;Brien and Salop (1999). A single collusion-intensity parameter κ ∈ [0,1] governs how much each cartel member internalizes the profits of other members. When κ = 0 the model reduces to competitive Cournot oligopoly; when κ = 1 all cartel members jointly maximize profits. In equilibrium, markups rise with firm market share, generating endogenous markup dispersion. Adding collusion causes cartel members to face a lower effective demand elasticity — their own market share augmented by the weighted market shares of co-conspirators — and to charge supracompetitive markups (overcharges). Critically, the effect of cartels on aggregate productivity is theoretically ambiguous: the output contraction of colluding firms redirects demand toward non-colluding firms. If the cartel is composed of the largest (most productive) firms, demand shifts toward less productive non-members, reducing productivity. If the cartel is composed of the least efficient firms, demand shifts toward large non-members, potentially improving allocation.&lt;/p&gt;
&lt;p&gt;The model is calibrated to match six moments from French data in 2007 — aggregate markup, cartel overcharge, the slope of the inverse-markup-on-HHI regression, the median number of firms per sector, the median number of cartel members, and the distribution of relative sales. The key calibrated parameters are: within-sector elasticity of substitution ρ = 10.19; across-sector elasticity η = 1.86; collusion intensity κ = 0.79. The cartel overcharge target is set to 10%, consistent with the OECD benchmark used by antitrust authorities and with Laborde (2021).&lt;/p&gt;
&lt;p&gt;Main quantitative findings (baseline calibration, cartels composed of top producers):&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Eliminating all cartels raises aggregate TFP by 1.1%.&lt;/li&gt;
&lt;li&gt;The productivity cost of markups with respect to the efficient allocation is 70% higher in the model with collusion (3.67%) than in the calibrated competitive oligopoly (2.16%), because collusion generates additional markup dispersion on top of the dispersion inherent in firm heterogeneity.&lt;/li&gt;
&lt;li&gt;Eliminating cartels brings the economy 30% closer to the efficient allocation.&lt;/li&gt;
&lt;li&gt;The aggregate markup falls by approximately 1.5 percentage points when cartels are eliminated.&lt;/li&gt;
&lt;li&gt;Consumption-equivalent welfare gains from eliminating cartels equal 2%.&lt;/li&gt;
&lt;li&gt;Larger cartels (market share above median) account for roughly 80% of the productivity gains; dismantling only large cartels yields a 0.88% TFP gain and 1.97% consumption-equivalent welfare gain; smaller cartels yield 0.23% TFP and 0.54% welfare.&lt;/li&gt;
&lt;li&gt;Umbrella pricing — non-cartel members raise their markups because the cartel&amp;rsquo;s higher prices provide cover — dampens aggregate gains quantitatively but only slightly: fixing non-members&amp;rsquo; markups yields 1.14% productivity gain versus 1.11% in the benchmark.&lt;/li&gt;
&lt;li&gt;Reducing collusion intensity from κ = 0.79 to κ ≈ 0.4 (roughly a 50% reduction) still generates TFP gains of 0.54% and welfare gains of 0.85%, demonstrating that tougher antitrust enforcement at the intensive margin (forcing cartels to soften, not dissolve) yields substantial gains.&lt;/li&gt;
&lt;li&gt;These estimates are one order of magnitude above Harberger&amp;rsquo;s (1954) 0.1% dead-weight loss estimate; the paper shows this discrepancy arises because Harberger uses sectoral data and near-unit demand elasticities, both of which suppress markup dispersion within sectors.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The paper&amp;rsquo;s scope conditions are explicit: results reflect the static cost of cartels; dynamic effects (entry deterrence, innovation incentives) are acknowledged but not quantified; only domestic, detected cartels are covered, so estimates likely understate the true cost; the channel through geographic markup dispersion is excluded.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-primary-identification-strategy-and-what-are-its-main-limitations"&gt;Q1. What is the paper&amp;rsquo;s primary identification strategy, and what are its main limitations?&lt;/h3&gt;
&lt;p&gt;The paper does not rely on a natural experiment or difference-in-differences design. Instead, it uses a structural calibration approach: a heterogeneous-firm oligopoly model with collusion is calibrated to match French data moments, and the cost of cartels is computed as the difference between the calibrated cartel equilibrium and a counterfactual competitive Nash-Cournot equilibrium. The main threats to this strategy are: (1) the sample of cartels consists only of detected cartels, which may not be representative of the latent population — discovered cartels could be either more or less severe than undiscovered ones; (2) no firm-level price data are available, so markups cannot be estimated directly; (3) the counterfactual is a calibrated competitive model rather than an empirically observed post-cartel state; (4) the model abstracts from entry and exit, which may dampen or amplify the true gains from cartel dissolution.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-through-which-cartels-affect-aggregate-productivity-and-how-are-they-distinguished"&gt;Q2. What are the main mechanisms through which cartels affect aggregate productivity, and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;Two channels operate simultaneously. First, the direct price effect: cartel members raise markups above the competitive level (overcharges), reducing their output. In the presence of markup dispersion, this disproportionately contracts output from high-markup (high-productivity) firms, increasing misallocation. Second, the demand reallocation effect: as cartel members contract output and raise prices, non-cartel members gain market share and increase their markups via the umbrella pricing mechanism. The net effect on productivity depends on which firms gain market share. When cartels consist of top producers, reallocation goes toward less productive non-members, reducing aggregate TFP. When cartels consist of the least efficient firms, reallocation goes toward larger non-members, potentially improving allocation. The two channels are not empirically separated in the data; rather, the model disentangles them analytically and then disciplines the net effect via calibration to observed cartel overcharges.&lt;/p&gt;
&lt;h3 id="q3-why-do-the-authors-assume-cartels-are-composed-of-the-most-productive-firms-and-what-is-the-evidence-for-this"&gt;Q3. Why do the authors assume cartels are composed of the most productive firms, and what is the evidence for this?&lt;/h3&gt;
&lt;p&gt;The assumption is motivated by three pieces of evidence. First, empirical regressions on the matched administrative data show that cartel members within their 4-digit industries have roughly 1,900% more sales, 1,150% more employment, and 37% higher labor productivity than non-members. Second, firms within a cartel are much more homogeneous than the overall within-industry distribution: the interquartile productivity ratio within a cartel is 1.4-to-1, versus approximately 2-to-1 for all non-cartel firms in the same industry, and the 90-10 ratio is 1.7-to-1 within a cartel versus over 4-to-1 across the industry. Third, only the top-producer composition assumption, combined with a collusion intensity κ = 0.79, can generate a cartel overcharge of 10% consistent with the calibration target. All other composition configurations (least efficient, all-inclusive, random top-10%) yield either implausibly small overcharges or implausibly large ones.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-umbrella-pricing-effect-and-how-large-is-it-quantitatively"&gt;Q4. What is the umbrella pricing effect and how large is it quantitatively?&lt;/h3&gt;
&lt;p&gt;Umbrella pricing refers to the mechanism by which cartel members&amp;rsquo; higher prices raise the sectoral price index, allowing non-cartel members to expand output and raise their own markups without reducing their market share. Proposition 1 of the model shows that collusion increases the markups of all firms — cartel and non-cartel — with non-cartel members experiencing markup increases that are larger for larger non-members. Quantitatively, when non-cartel members are held to fixed markups (so the umbrella effect is turned off), the aggregate TFP gain from eliminating cartels rises from 1.11% to 1.14% — a difference of 0.03 percentage points, or less than 3% of the total effect. The welfare effect is similarly small: 2.01% versus 2.00%. The umbrella pricing channel thus dampens aggregate gains but is quantitatively minor.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-in-cartel-effects-is-documented"&gt;Q5. What heterogeneity in cartel effects is documented?&lt;/h3&gt;
&lt;p&gt;Three dimensions of heterogeneity are explored. First, cartel size matters: large cartels (those with cumulated market share above the median) account for roughly 80% of the aggregate TFP gain from eliminating all cartels (0.88 percentage points out of 1.11%), while small cartels account for only 0.23 percentage points. Second, cartel composition is critical: top-producer cartels amplify misallocation, all-inclusive cartels generate very large overcharges and dramatically higher misallocation, least-efficient-firm cartels barely affect allocation, and random-top-10% cartels can slightly improve allocation. Third, collusion intensity matters monotonically: across the range κ = 0.1 to κ = 0.4, TFP gains from elimination fall from 0.99% to 0.54%, and welfare gains fall from 1.70% to 0.85%.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run-and-how-do-the-results-change"&gt;Q6. What robustness checks are run, and how do the results change?&lt;/h3&gt;
&lt;p&gt;The paper runs six main robustness experiments, all recalibrating the model: (1) Alternative overcharge target of 15% (versus 10% baseline): requires κ = 1.28, yields TFP gains of 1.63% and welfare gains of 2.77%. (2) Low aggregate markup target M = 1.1: TFP gain of 1.37%, welfare gain of 2.07%. (3) High aggregate markup target M = 1.3: TFP gain of 0.90%, welfare gain of 1.96%. (4) Bertrand rather than Cournot competition: TFP gain of 0.55%, welfare gain of 1.35% — smaller because Bertrand generates less markup dispersion, though the reduction in distance to the efficient allocation is larger (39%). (5) Heterogeneous κ across cartels drawn from a truncated normal with four variance levels: TFP gains range from 0.84% to 1.11% and welfare gains from 1.53% to 1.99%, close to the benchmark of 1.11% and 2.00%. (6) The cartel screen regression yields an estimated κ of 0.70 from data on colluding firms, close to the calibrated benchmark of 0.79.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-model-generate-a-cartel-detection-screen-and-what-does-it-find"&gt;Q7. How does the model generate a cartel detection screen, and what does it find?&lt;/h3&gt;
&lt;p&gt;The model&amp;rsquo;s equilibrium first-order conditions imply a regression of a cartel member&amp;rsquo;s labor share (a proxy for the inverse markup under log-linear production) on its own market share and the total cartel market share. The ratio of the estimated coefficient on cartel market share to the sum of both coefficients recovers the collusion intensity κ. Running this regression on the sample of detected cartel firms, the authors find a coefficient on own market share of -0.53 and an intercept of 0.70, both significant at 1%. Adding the cartel joint market share, its coefficient is negative and significant at 1%; the estimated κ from this specification is 0.70, close to the benchmark of 0.79. Results are qualitatively robust to including year fixed effects, though estimates become slightly noisier.&lt;/p&gt;
&lt;h3 id="q8-how-do-the-authors-explain-the-large-discrepancy-with-harberger-1954"&gt;Q8. How do the authors explain the large discrepancy with Harberger (1954)?&lt;/h3&gt;
&lt;p&gt;Harberger&amp;rsquo;s classic estimate of the deadweight loss from monopoly is approximately 0.1% of GDP. The authors show that their model can reproduce estimates close to this when (a) the model is aggregated to the sectoral level, eliminating within-sector markup dispersion — in that case, the TFP gain from eliminating cartels falls to 0.08%; or (b) demand elasticities are set close to unity as in Harberger&amp;rsquo;s sectoral data — the TFP gain falls to 0.24%. The key reason for the discrepancy is that Harberger&amp;rsquo;s framework suppresses both the within-sector dispersion of markups (which in the baseline model amplifies allocative losses) and the endogenous markup response to market share changes (which is large when ρ is substantially greater than 1). Using disaggregated firm-level data and calibrated high-within-sector elasticities restores the large estimated costs.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper implies that antitrust enforcement against horizontal price-fixing cartels can yield aggregate TFP gains of 1.1% and welfare gains of 2% in consumption-equivalent terms — figures the authors describe as conservative, because (i) the estimate is static (no dynamic gains from entry or innovation effects are included), (ii) only domestic detected cartels are captured and international cartels are excluded, (iii) geographic markup dispersion is abstracted from, and (iv) the calibration uses a conservative overcharge target of 10%. Importantly, the gains from targeting the intensive margin (forcing cartels to reduce overcharges rather than dissolving them entirely) are also substantial: a 50% reduction in κ still yields 0.54% TFP and 0.85% welfare gains. The results further imply that industrial policy and trade liberalization reforms that ignore competition enforcement may be partially undermined if new market power enables cartelization. The scope condition most critical to the quantitative magnitude is cartel composition: results depend on cartels being composed of top producers; the sign and magnitude of productivity effects can flip for alternative compositions. The authors also note that if cartels spur long-run innovation (through higher profits), their static welfare cost estimates would overstate the net social cost.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-differ-from-edmond-midrigan-and-xu-2022-and-baqaee-and-farhi-2020"&gt;Q10. How does this paper differ from Edmond, Midrigan, and Xu (2022) and Baqaee and Farhi (2020)?&lt;/h3&gt;
&lt;p&gt;Edmond et al. (2022) and Baqaee and Farhi (2020) quantify the total welfare and productivity cost of markups relative to the efficient allocation — the gap between the current economy (with all its markup dispersion from firm heterogeneity) and the first-best. Moreau and Panon instead isolate the cost of one specific, policy-relevant source of excess markup dispersion — collusion — by computing the gap between the cartel equilibrium and the competitive (but still imperfect) Nash-Cournot equilibrium. They also show that competitive oligopoly models of the Edmond et al. type understate the total misallocation cost of markups by approximately 70% when cartels are present and composed of top producers, because competitive models are calibrated to match the same aggregate markup data but attribute all markup dispersion to firm heterogeneity rather than to collusion. The papers are thus complementary: Edmond et al. bound the full cost of all markup distortions, while Moreau and Panon bound the portion attributable to cartels and amenable to competition enforcement.&lt;/p&gt;
&lt;h3 id="q11-what-caveats-and-limitations-do-the-authors-acknowledge"&gt;Q11. What caveats and limitations do the authors acknowledge?&lt;/h3&gt;
&lt;p&gt;The authors flag several important limitations. (1) The analysis is static: dynamic effects — including entry deterrence by cartels, barriers to exit for inefficient firms, and the innovation-competition relationship — are not modeled. The relationship between competition and innovation is hump-shaped (Aghion et al., 2005), so cartels could in principle spur or dampen innovation; the authors treat their estimates as an upper bound if cartels raise innovation. (2) Only detected French domestic cartels are in the sample; international cartels (investigated by the European Commission) and undetected cartels are excluded, likely causing understatement of total costs. (3) The selection of detected cartels is non-random: the direction of bias from using only discovered cartels is unclear — discovered cartels may be unusually large (biasing costs upward) or undiscovered large cartels may exist (biasing costs downward). (4) The model abstracts from geographic markup dispersion and from vertical arrangements across industries. (5) The model has no entry or exit of firms, which could amplify or dampen transition dynamics. (6) Firm-level prices are unavailable, so markups cannot be directly measured and must be inferred from the model or from labor shares.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Collusion intensity parameter (κ)&lt;/strong&gt;: A scalar in [0,1] that governs the weight each cartel member assigns to co-conspirators&amp;rsquo; profits when choosing output. When κ = 0, behavior is competitive Cournot; when κ = 1, members jointly maximize aggregate cartel profits. In the baseline calibration κ = 0.79, chosen to match a 10% median cartel overcharge in French data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cartel overcharge&lt;/strong&gt;: The percentage difference in cartel members&amp;rsquo; average markups between the cartel equilibrium and the competitive Nash-Cournot equilibrium. Computed as the median overcharge across cartels in the model. In the baseline calibration it is 10%, consistent with the OECD benchmark and Laborde (2021). The overcharge increases with both collusion intensity (κ) and the cartel&amp;rsquo;s total market share.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Umbrella pricing&lt;/strong&gt;: The mechanism by which a cartel&amp;rsquo;s higher prices raise the sectoral price index, enabling non-cartel members to expand demand, gain market share, and charge higher markups than they would in the absence of the cartel. In the model, umbrella pricing implies that the introduction of collusion increases the markups of all firms in cartelized sectors, not just cartel members; quantitatively, the effect dampens but does not reverse the aggregate productivity gains from cartel dissolution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distance to efficient allocation&lt;/strong&gt;: The ratio of the productivity gain from eliminating cartels (Acartel → Acomp) to the total productivity gain from eliminating all markup dispersion (Acomp → Aeff or equivalently from Acartel → Aeff). In the baseline, eliminating cartels reduces this distance by 30%, meaning cartels are responsible for roughly 30% of the gap between the actual economy and the first-best efficient allocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous markups (size-related)&lt;/strong&gt;: In the Atkeson-Burstein framework embedded in this model, a firm&amp;rsquo;s equilibrium markup is a harmonic average of within- and between-sector demand elasticities weighted by the firm&amp;rsquo;s own market share. More productive firms endogenously hold larger market shares and thus face lower demand elasticities, charging higher markups. Collusion further distorts this by augmenting the effective market share with co-members&amp;rsquo; shares, yielding supracompetitive overcharges.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cartel composition&lt;/strong&gt;: The identity of firms within a cartel — specifically, where they sit in the within-industry productivity distribution. The paper shows this is the single most important determinant of whether cartels amplify or dampen aggregate misallocation. Empirically, discovered French cartels are composed of the largest, most productive firms (nearly 1,900% more sales than non-members), and this is the only composition configuration that can match observed 10% overcharges in the calibrated model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive versus extensive margin of cartel policy&lt;/strong&gt;: The extensive margin refers to whether a cartel exists (zero versus positive κ); the intensive margin refers to the degree of collusion among existing cartel members (high versus low κ). The paper shows both margins are quantitatively important: breaking down all cartels (extensive margin) yields 1.11% TFP gain, while halving κ without dissolution (intensive margin) yields 0.54% TFP gain and 0.85% welfare gain.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cartel screen&lt;/strong&gt;: A regression of cartel members&amp;rsquo; labor shares on their own market share and the joint cartel market share, derived directly from the model&amp;rsquo;s equilibrium first-order conditions. The collusion intensity κ can be recovered as the ratio of the joint market share coefficient to the sum of both market share coefficients. Applied to French data on detected cartel firms, this screen yields κ̂ = 0.70, close to the calibrated value of 0.79.&lt;/p&gt;</description></item><item><title>Identifying Monetary Policy Shocks: A Natural Language Approach</title><link>https://macropaperwarehouse.com/papers/identifying-monetary-policy-shocks-a-natural-language-approach/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/identifying-monetary-policy-shocks-a-natural-language-approach/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: To study how monetary policy affects the economy, macroeconomists must isolate &amp;ldquo;shocks&amp;rdquo; — changes in interest rates that are not systematic responses to economic conditions. The paper proposes a new identification method that captures the Federal Reserve&amp;rsquo;s information set far more comprehensively than prior approaches, using the natural-language text of documents Fed staff prepare for FOMC meetings, not just numerical forecasts.&lt;/p&gt;
&lt;p&gt;Method and data: The approach extends Romer and Romer (2004), who regress changes in the Federal Funds Rate (FFR) target on Greenbook forecasts and take the residual as the shock. The authors instead convert the text of FOMC documents into many &amp;ldquo;aspect-based&amp;rdquo; sentiment time series and predict the FFR change with both these sentiments and an expanded forecast set. They process 772 PDF files for 276 meetings (630 files for 210 meetings before the zero lower bound), covering Greenbook 1/2, Tealbook A, Redbook, and Beigebook documents, starting October 5, 1982 (when the Fed began targeting the FFR per Thornton 2006). Most documents are released with a 5-year lag, so the latest is from end-2016. They extract the most frequently mentioned economic terms, yielding 296 single/multi-word concepts (e.g., &amp;ldquo;inflation,&amp;rdquo; &amp;ldquo;economic activity&amp;rdquo;). For each concept they build a sentiment indicator by scoring positive (+1) and negative (-1) words within a 10-word window, using an augmented Loughran-McDonald (2011) dictionary of 2,882 classified words. The empirical model (equation 3) includes 132 forecast series, 296 sentiment indicators with 4 lags, and quadratic terms — 3,226 regressors total — far exceeding the 210 FOMC-meeting observations over October 1982 to October 2008. They estimate it with a ridge regression, choosing the penalty by 10-fold cross-validation; the shock is the residual.&lt;/p&gt;
&lt;p&gt;Main quantitative findings: (1) Fit/systematic share: the original Romer-Romer OLS specification yields R-squared of 0.50 (so 50% of FFR variation is attributed to shocks), while the preferred nonlinear ridge with forecasts and sentiments yields R-squared of 0.94 — cutting the exogenous shock share from 50% to 6%, an almost ten-fold reduction. Lags 0–4 give R-squared of 0.75, 0.81, 0.90, 0.92, 0.94. (2) Information content: text-based sentiments predict Greenbook unemployment-rate forecast errors; a one-standard-deviation increase in the sentiment first principal component is associated with an almost 0.5 percentage-point negative 1-year-ahead forecast error (R-squared up to 0.25), supporting the view that staff forecasts are modal, not mean, predictions. (3) Comparison to high-frequency surprises: correlation with Swanson (2021) FFR surprises (1991–2008) is 0.49 (vs. 0.36 for Romer-Romer); 0.77 for the top-10 shocks (vs. 0.61) and 0.51 for the top-10 surprises (vs. 0.18). The estimated shocks have lower autocorrelation (0.066 vs. 0.204 for Romer-Romer). (4) IRFs (BVAR with shock as external instrument, IRF sample 1984:02–2016:12): a tightening produces a persistent yield rise (about 20 months), a fall in real output and rise in unemployment materializing after about a year, a sluggish decline in the price level (mild initial &amp;ldquo;price puzzle,&amp;rdquo; visibly negative after about 18 months, significantly negative after 30 months), a sharp rise in the excess bond premium, and a fall in stock prices — all consistent with theory. By contrast, Romer-Romer OLS residuals imply flat output/unemployment responses, an insignificant EBP response, and positive stock-price/rate comovement, at odds with theory.&lt;/p&gt;
&lt;p&gt;Implications: Including text-based information is essential for clean identification — even for the original method to correctly recover responses (especially of unemployment). A Beigebook-only version extends the method to recent meetings, implying the 2022–2023 tightening (525 bp total) carried only about 21 bp of contractionary shock.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-exactly-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What exactly is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Monetary policy shocks are defined (equation 1) as the residual after orthogonalizing the FFR target change against the central bank&amp;rsquo;s information set. The authors proxy that information set with the full numerical-forecast set plus 296 text-derived sentiment indicators (with 4 lags and quadratic terms), and estimate the prediction via ridge regression with 10-fold cross-validation. The shock is the residual. Two key assumptions inherited from Romer-Romer are threats: (i) the included variables must be a good proxy for the true information set — the paper argues forecasts alone are insufficient because they are modal, not mean, predictions and assume a specific policy path (Faust-Wright 2008), which is why text is required; and (ii) the mapping from information to decisions must be well-specified — they relax linearity by adding quadratic terms. A residual concern is that even the large information set may not capture truly idiosyncratic considerations, but they argue this is exactly what should remain in the shock.&lt;/p&gt;
&lt;h3 id="q2-why-are-text-sentiments-necessary-beyond-numerical-forecasts--what-is-the-cochrane-critique-and-how-do-they-answer-it"&gt;Q2. Why are text sentiments necessary beyond numerical forecasts — what is the Cochrane critique and how do they answer it?&lt;/h3&gt;
&lt;p&gt;Cochrane (2004) argued that to study the effect of policy on a given variable, it suffices to orthogonalize the FFR against the Fed&amp;rsquo;s forecast of that variable alone, since an efficient forecast incorporates all relevant information. This holds only if Greenbook forecasts equal the conditional mean. The authors show, via FOMC transcripts (Appendix D, spanning 1985–2016) and econometrics, that staff produce MODAL forecasts accompanied by verbal descriptions of asymmetric risks. Their sentiment indicators predict Greenbook unemployment forecast errors (Table 2): the first PC and even the single &amp;rsquo;economic activity&amp;rsquo; sentiment are significant at multiple horizons (R-squared up to 0.25; a 1-sd PC increase implies an almost 0.5 pp negative 1-year error). After orthogonalizing forecast errors on sentiment, the error distribution becomes more symmetric and centered on zero (Figure 3). Hence at least some text information is required even for the original Romer-Romer method to recover the true unemployment response.&lt;/p&gt;
&lt;h3 id="q3-why-ridge-regression-rather-than-lasso-or-ols"&gt;Q3. Why ridge regression rather than LASSO or OLS?&lt;/h3&gt;
&lt;p&gt;OLS is infeasible (3,226 regressors vs. 210 observations). Ridge minimizes residual sum of squares plus a penalty on squared coefficients (shrinkage toward zero), equivalent to Bayesian OLS with a normal prior centered at zero. Unlike LASSO (which produces sparse models), ridge keeps all regressors (a dense model), more akin to factor models/PCA. The authors prefer dense methods because economic data have many correlated regressors and few observations; Giannone, Lenza, and Primiceri (2022) (&amp;rsquo;the illusion of sparsity&amp;rsquo;) find sparse methods become unstable under high collinearity — clearly present across forecasts and sentiments here. The penalty lambda is chosen by 10-fold cross-validation, so the high R-squared is not purely mechanical.&lt;/p&gt;
&lt;h3 id="q4-how-do-the-authors-interpret-what-the-shocks-capture-and-what-case-studies-support-this"&gt;Q4. How do the authors interpret what the shocks capture, and what case studies support this?&lt;/h3&gt;
&lt;p&gt;They inspect FOMC discussions in meetings with the largest estimated shocks. November 7, 1984: largest shock in absolute value — a 75 bp FFR decline of which staff forecasts/sentiments predict 53 bp, leaving a -22 bp easing shock, driven by FOMC participants finding the staff forecast too optimistic. November 15, 1994: a 75 bp hike of which 21 bp is a contractionary shock — Greenspan argued &amp;lsquo;a mild surprise would be of significant value&amp;rsquo; for credibility, and the 75-vs-50 bp gap between his decision and the staff&amp;rsquo;s option almost exactly matches the estimated 21 bp. The interpretation: shocks are FFR decisions that are &amp;lsquo;surprises&amp;rsquo; to the Fed staff — orthogonal to the staff&amp;rsquo;s information set. They note their interpretation is narrower than Romer-Romer&amp;rsquo;s (which included target-definition changes and political pressure, both pre-1982 phenomena per Drechsel 2023). Systematic credibility concerns would be absorbed into systematic policy; only nonsystematic ones become shocks.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-three-interpretations-of-why-romer-romer-irfs-go-wrong-and-how-are-they-distinguished"&gt;Q5. What are the three interpretations of why Romer-Romer IRFs go wrong, and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;(1) Unemployment: because Greenbook unemployment forecasts are modal and text-sentiment predicts their errors, the Romer-Romer OLS cannot fully absorb asymmetric risk shifts, producing a spurious correlation (easing shocks estimated when unemployment rises) and thus a flat/incorrect unemployment IRF (Figure 6). (2) Stock prices: the Fed systematically reacts to equities (Cieslak and Vissing-Jorgensen 2020); failing to control for this leaves spurious positive rate/stock comovement. They test this by adding HF S&amp;amp;P500 surprises as a second instrument with Jarocinski-Karadi (2020) sign restrictions (negative rate/stock comovement for policy shocks): their measure already satisfies the restrictions (Panel a barely changes), whereas the Romer-Romer IRFs change drastically once imposed, &amp;lsquo;correcting&amp;rsquo; activity/price/EBP responses (Figure 7). (3) Credit spreads: Romer-Romer residuals retain endogenous credit-spread variation; the authors&amp;rsquo; sentiments include &amp;lsquo;spreads,&amp;rsquo; &amp;lsquo;credit standards,&amp;rsquo; &amp;lsquo;credit quality.&amp;rsquo; Caldara and Herbst (2019) show that ignoring the Fed&amp;rsquo;s credit-spread reaction attenuates IRFs, supporting this channel.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;(1) 5-word vs. 10-word sentiment windows give nearly identical R-squared (0.95 vs. 0.94 in the top spec). (2) Sentence-based sentiment construction is highly correlated with the window-based version (0.875 for employment, 0.959 for credit; Appendix C). (3) Lag structure: 0–4 lags raise R-squared 0.75→0.94 with diminishing gains past 4 lags. (4) FOMC composition controls (governor/bank-rep attendance, voting status, appointing president, female attendance) raise R-squared by less than 0.1% — personal dynamics do not drive FFR changes. (5) Alternative nonlinear forms: cubic residuals 99% correlated with quadratic; a ~40,000-variable full-interaction spec yields residuals 96% correlated with quadratic. (6) Forecast-error predictability holds for output and inflation too (Appendix E), and using first-release vs. final-vintage data gives similar results. (7) Local projections (Jorda 2005) confirm the BVAR results, with Romer-Romer again off-theory. (8) IRFs built from only the 10 largest shocks reproduce the main pattern. (9) The extended-forecast ridge (no sentiments) already corrects the IRFs, though the authors stress theory-consistent IRFs are necessary but not sufficient for a good shock measure.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-beigebook-only-extension-work-and-what-does-it-find"&gt;Q7. How does the Beigebook-only extension work and what does it find?&lt;/h3&gt;
&lt;p&gt;Tealbooks/forecasts are released with a 5-year lag, but Beigebooks are public before each meeting. Over 1982–2008, building sentiments from Beigebooks alone gives indicators strongly correlated with the baseline (e.g., &amp;rsquo;economic activity&amp;rsquo;, Figure 8), an R-squared of 0.68 (vs. 0.94 with full documents), and shocks correlated 0.92 with the baseline shocks, with qualitatively similar IRFs. As a proof of concept over December 2015–October 2023 (excluding the March 2020–December 2021 ZLB period), the R-squared is 0.98. Inflation sentiment dropped more than 6 standard deviations in late 2021/early 2022 (driven by &amp;lsquo;concern&amp;rsquo; near &amp;lsquo;inflation&amp;rsquo;). The 2022–2023 tightening of 525 bp total implies only about 21 bp of cumulative contractionary shock — i.e., mostly systematic tightening. This extension is impossible for Romer-Romer because Beigebooks contain no numerical forecasts.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q8. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It contributes to three literatures. (1) Monetary-shock identification: builds directly on Romer-Romer (2004) but adds NLP/ML and a much larger information set; contrasts with SVAR and high-frequency approaches (Gurkaynak et al. 2005, Gertler-Karadi 2015, Swanson 2021, Bauer-Swanson). (2) Text/ML on Fed documents: unlike Sharpe-Sinha-Hollrah (2020), who build a single sentiment index, the authors build aspect-based sentiments per concept; closest are Handlan (2020), who builds a &amp;rsquo;text shock&amp;rsquo; separating forward guidance from current assessment since 2005, and Ochs (2021), who extracts surprises from the private agents&amp;rsquo; viewpoint — the authors instead orthogonalize against the Fed&amp;rsquo;s internal information set, staying closer to Romer-Romer. (3) Greenbook-forecast literature (Romer-Romer 2000, Faust-Wright, Nakamura-Steinsson 2018): they emphasize the modal nature of forecasts and show sentiments explain forecast errors on average.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policyresearch-implications-and-their-scope-conditions"&gt;Q9. What are the policy/research implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The method delivers a cleanly identified, &amp;lsquo;all-purpose&amp;rsquo; shock series usable for any macro variable — including ones without Fed forecasts (e.g., credit spreads). It spans a longer period than HF measures (which begin in the early 1990s due to futures-data availability and the fact that the FOMC did not announce rate changes publicly before 1994). Scope conditions: the preferred (Tealbook-based) measure requires the 5-year document lag, so recent meetings need the lower-fidelity Beigebook-only version (R-squared 0.68 in-sample); the main estimation sample ends October 2008 to avoid the ZLB. The method relies on the structured, consistent wording of Fed-staff documents, making dictionary-based sentiment particularly applicable. The authors recommend using the baseline measure whenever feasible, even at the cost of dropping recent observations, and resorting to Beigebook-only only when that cost is high. They also suggest combining their measure with HF surprises as multiple external instruments.&lt;/p&gt;
&lt;h3 id="q10-are-there-caveats-about-interpreting-the-models-coefficients"&gt;Q10. Are there caveats about interpreting the model&amp;rsquo;s coefficients?&lt;/h3&gt;
&lt;p&gt;Yes. The ridge is built for prediction (y-hat), not coefficient interpretation (beta-hat). With 3,226 highly collinear regressors plus lags and quadratic terms, individual coefficients cannot be cleanly interpreted — the authors invoke Mullainathan-Spiess (2017) that ML belongs in the y-hat toolbox, and a self-driving-car analogy. A potential downside of a large information set is low statistical power in the shock (since more variation becomes systematic), but they show via the BVAR IRFs that power is not a problem in practice.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Identifying the Impact of Inflation Expectations</title><link>https://macropaperwarehouse.com/papers/identifying-the-impact-of-inflation-expectations/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/identifying-the-impact-of-inflation-expectations/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Branch (2022) asks whether subjective consumer inflation expectations causally raise the inflation rate — a question whose empirical answer has been elusive despite its central role in New Keynesian theory and central bank communication. The identification problem is acute: expectations are endogenous by construction, and the standard approach of estimating a Phillips curve with aggregate data produces estimates biased sharply downward by endogeneity. OLS regressions of regional inflation on regional mean expectations, controlling for unemployment, lagged inflation, and region and time fixed effects, yield a slope of only 0.069 (Table 2 context; Figure 1b), far below the theoretical prior of near-unity pass-through.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s empirical strategy exploits a key fact: different demographic groups consume heterogeneous bundles of goods, so their inflation expectations differ systematically and reflect their own basket&amp;rsquo;s price movements. Using roughly 273,000 individual responses from the University of Michigan Survey of Consumers spanning 1978:1–2022:5, Branch classifies respondents into 160 demographic groups defined by sex, age (five categories), education (four levels), marital status, and parental status. The panel covers four U.S. Census regions, producing dimensions T = 528 months, N = 4 regions, and G = 160 groups. Regional inflation is measured from BLS CPI series for all urban consumers.&lt;/p&gt;
&lt;p&gt;The identification strategy is a shift-share (Bartik) instrument: for each region-month, the predicted regional inflation expectation is the population-weighted average of each demographic group&amp;rsquo;s national-level average inflation expectation, where the weights are the group&amp;rsquo;s share of the region&amp;rsquo;s population. Two share measures are used: (i) the January 1978 Current Population Survey (CPS78) distribution, which is time-invariant and plausibly exogenous to subsequent inflation shocks; and (ii) contemporaneous Michigan survey shares. The leave-one-out variant is the preferred construction. The instrument is relevant — first-stage F-statistic of 52.4 (significant at 0.1%) — and the Durbin-Wu-Hausman test rejects OLS consistency at the 1% level (statistic = 8.074).&lt;/p&gt;
&lt;p&gt;Main 2SLS estimates: using Michigan survey shares, a 1 percentage point increase in a region&amp;rsquo;s expected inflation raises regional inflation by 0.33 percentage points (significant at 5%; Table 2). Using CPS78 shares, the estimate rises to 0.55 percentage points (significant at 1%; Table 2). After applying the split-sample jackknife bias correction for finite-sample bias in the small-N/large-T panel, the estimates increase slightly to 0.36 and 0.60 respectively (Table 3). The paper characterizes the 60 basis point estimate as its &amp;ldquo;preferred&amp;rdquo; figure. Both are substantially above the OLS estimate of 0.069 and represent a lower bound: because time fixed effects absorb cross-regional spillovers, the aggregate pass-through is likely stronger, with the paper arguing that after accounting for spillovers the effect is plausibly in the range of 1.0–1.6, consistent with the Calvo- and Taylor-model predictions of Werning (2022), who shows pass-through should lie in [1/2, 1] or above.&lt;/p&gt;
&lt;p&gt;Sectoral decomposition reveals that the expectation effect is concentrated in non-durable goods prices (coefficient 1.74, significant at 1%; Table 7) and commodities more broadly (1.29, significant at 1%; Table 7), with no statistically meaningful effect on durables (−0.10, insignificant) and only marginal positive effects on services (0.22, marginally significant). Among services, the effect is somewhat larger when housing services are excluded.&lt;/p&gt;
&lt;p&gt;A key finding on expectations horizons: when both one-year-ahead and five-to-ten-year-ahead expectations are simultaneously instrumented using their respective Bartik shift-shares, only the short-run (one-year) expectation retains a significant positive effect on inflation. The long-horizon coefficient is small in absolute value, negative in sign, and statistically insignificant in both the joint and standalone specifications (Tables 10 and 12). After conditioning on aggregate macroeconomic factors captured by time fixed effects, long-run inflation expectations have no independent causal role in the regional inflation rate.&lt;/p&gt;
&lt;p&gt;Identification heterogeneity: using the Rotemberg weight decomposition of Goldsmith-Pinkham, Sorkin, and Swift (2020), the identifying variation derives primarily from younger, married consumers with at least a high school degree — specifically those aged 18–34 (Michigan instrument) or 25–49 (CPS78 instrument). The group-specific treatment effects (βg) for these heavily weighted groups are positive and significantly above 1. Temporally, the heaviest identification weights fall on the Great Inflation and Volcker disinflation (1978–82), the Great Recession (2007–09), and the post-pandemic inflation episode (2021–22). The impulse response function shows a significant contemporaneous positive effect of expectations on inflation that mean-reverts cyclically within approximately 12 months, though confidence bands are wide at longer horizons.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-and-what-makes-it-plausible"&gt;Q1. What is the core identification strategy and what makes it plausible?&lt;/h3&gt;
&lt;p&gt;The strategy is a differential-exposure quasi-experiment using a Bartik (shift-share) instrument. For each Census region and month, the instrument is the population-weighted average of each demographic group&amp;rsquo;s national-level mean inflation expectation, with weights equal to that group&amp;rsquo;s share of the region&amp;rsquo;s population. The key identifying assumption has two parts: (1) demographic groups have heterogeneous consumption baskets, so their inflation expectations reflect the prices in their own basket; and (2) the distribution of demographic groups across regions is exogenous to unobserved shocks driving regional inflation (as opposed to being exogenous to regional price levels, which is a weaker and separately justified claim). Plausibility is supported by the CPS78 shares having no predictive power for the other covariates of inflation over the sample, and by using a leave-one-out instrument construction to avoid mechanical correlation.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-threats-to-identification-and-how-does-the-paper-address-them"&gt;Q2. What are the main threats to identification and how does the paper address them?&lt;/h3&gt;
&lt;p&gt;The principal threat is that regional demographic composition could be endogenous to regional inflation rather than merely to regional price levels. The paper argues identification requires only exogeneity to the change in prices (inflation), not to the level. The empirical check is that CPS78 beginning-of-period shares show no statistically or economically significant correlation with the other regressors that predict regional inflation. A second threat is that groups may sort into regions based on economic conditions correlated with inflation. The paper argues the channel runs through demand from heterogeneous baskets rather than supply-side sorting. A third threat is weak instruments: this is addressed by first-stage F = 52.4. Fourth, survey measurement concerns (re-interview selection bias, outliers, endogenous prompting thresholds) are addressed through a battery of alternative specifications (first-time respondents only, outlier removal, CPS vs. survey shares, lagged shares, alternative CPI measures).&lt;/p&gt;
&lt;h3 id="q3-why-are-ols-estimates-biased-downward-and-by-how-much"&gt;Q3. Why are OLS estimates biased downward and by how much?&lt;/h3&gt;
&lt;p&gt;OLS is biased because inflation expectations are endogenous — they move with the same shocks driving inflation, so OLS conflates the causal effect with reverse causation and omitted-variable bias. The OLS estimate from the panel regression with region and time fixed effects is approximately 0.069 (Figure 1b). The 2SLS estimates using the Bartik instrument range from 0.33 to 0.55, roughly five to eight times larger than OLS, confirming substantial downward bias. The Durbin-Wu-Hausman test confirms OLS inconsistency at the 1% level.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-across-demographic-groups-is-documented"&gt;Q4. What heterogeneity across demographic groups is documented?&lt;/h3&gt;
&lt;p&gt;Women consistently report higher inflation expectations than men, particularly outside the high-inflation 1970s episode. Older respondents (50+) receive small Rotemberg identification weights, meaning their expectations contribute little to the identifying variation. Younger groups (18–34 under Michigan shares; 25–49 under CPS78 shares), married, with at least a high school education are the groups whose expectations drive the regional cross-sectional identification. The group-specific causal effects (βg) for these heavily weighted groups are uniformly positive and significantly above 1.0, ranging roughly from 1.38 to 1.91 in the top-10 groups. College-educated groups receive higher weight under the CPS78 instrument, while the Michigan shares instrument weights high school and college groups more evenly.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-sectoral-decomposition-of-the-inflation-expectations-effect"&gt;Q5. What is the sectoral decomposition of the inflation expectations effect?&lt;/h3&gt;
&lt;p&gt;Table 7 estimates separate 2SLS regressions for components of the CPI. Non-durable goods prices respond most strongly (coefficient 1.74, significant at 1%). Commodities broadly (which include non-durables and durables) also show a large effect (1.29, significant at 1%). Durable goods prices show no meaningful effect (−0.10, statistically insignificant). Services show only a marginal positive effect (0.22, marginally significant at 10%). Among services, the effect is somewhat stronger when housing services are removed. These results are consistent with prior findings that consumer grocery and non-durable prices most directly influence and reflect household inflation expectations.&lt;/p&gt;
&lt;h3 id="q6-what-do-the-long-run-expectations-results-show-and-what-is-the-interpretation"&gt;Q6. What do the long-run expectations results show and what is the interpretation?&lt;/h3&gt;
&lt;p&gt;The Michigan survey&amp;rsquo;s PX5 question elicits 5-to-10-year ahead inflation expectations. Constructing a shift-share Bartik instrument for these long-horizon expectations and including both short- and long-run instruments simultaneously, the second-stage coefficient on long-horizon expectations is small (−0.023 to −0.037 in the joint specification, Table 10), negative, and statistically insignificant in all specifications. When long-horizon expectations alone are instrumented, the second-stage coefficient is 0.005 to 0.034 (Table 12), positive but still insignificant. The interpretation is that, after controlling for time fixed effects (which capture aggregate macroeconomic factors), long-run expectations have no independent causal role in regional inflation outcomes. Only short-run (one-year ahead) expectations matter. The first stage confirms the long-run instrument is relevant for long-run expectations but orthogonal to short-run expectations.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-reported-and-what-do-they-find"&gt;Q7. What robustness checks are reported and what do they find?&lt;/h3&gt;
&lt;p&gt;Table 8 reports four alternative specifications, all using Michigan survey shares: (1) &amp;lsquo;small&amp;rsquo; — removing survey responses with absolute values above 25% — gives a coefficient of 0.66 (significant at 1%), larger than baseline, though the paper does not prefer this because large expectations may have real behavioral effects; (2) &amp;lsquo;first-only&amp;rsquo; — using only first-time respondents and dropping the 40% re-interviewed — yields a coefficient of 0.58, still positive though the standard error rises and significance falls; (3) &amp;lsquo;state-CPI&amp;rsquo; — replacing the BLS regional CPI with state-level CPIs aggregated as in Hazell et al. (2022) — gives 0.33 (significant at 5%), very close to the Michigan-shares baseline; (4) &amp;rsquo;lag Michigan shares&amp;rsquo; — instrumenting with 12-month lagged survey shares — gives 0.53 (significant at 5%), bracketed between the two baseline estimates. The jackknife bias correction (Table 3) slightly raises estimates to 0.36 and 0.60 for the two instruments.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-impulse-response-function-show"&gt;Q8. What does the impulse response function show?&lt;/h3&gt;
&lt;p&gt;Using local projections (Jordà 2005) to estimate a 2SLS impulse response function, a shock to inflation expectations produces a significant positive contemporaneous effect on regional inflation. The response is cyclical and mean-reverting, returning to near zero within approximately 12 months. Confidence intervals are wide in subsequent quarters, so the analysis cannot rule out lingering effects, but the central estimates suggest the impact dissipates within about a year. The paper notes that the lack of strong persistence may reflect the specific U.S. inflation history and suggests extending the analysis to countries with more volatile or persistent inflation histories.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-the-new-keynesian-phillips-curve-literature"&gt;Q9. How does this paper relate to the New Keynesian Phillips Curve literature?&lt;/h3&gt;
&lt;p&gt;The standard approach to measuring expectations&amp;rsquo; impact on inflation is to estimate a NKPC with an instrument for expectations under rational expectations. Mavroeidis, Plagborg-Moller, and Stock (2014) document that this approach faces severe identification and weak-instrument problems. Branch&amp;rsquo;s approach avoids these issues by not assuming rational expectations, not requiring an explicit model of expectations formation, and using a shift-share instrument whose validity rests on cross-sectional demographic heterogeneity rather than time-series moment conditions. The theoretical model in Section 3.1 permits non-rational expectations and nests &amp;lsquo;anticipated utility&amp;rsquo; or &amp;lsquo;steady-state learning&amp;rsquo; (Evans and Honkapohja 2001; Woodford 2013) as the simplifying assumption. The estimated regional coefficients are below but potentially consistent with Werning&amp;rsquo;s (2022) theoretical range of [1/2, 1] for Calvo and Taylor pricing models once spillovers are accounted for.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-paper-relate-to-the-literature-on-household-level-inflation-heterogeneity"&gt;Q10. How does the paper relate to the literature on household-level inflation heterogeneity?&lt;/h3&gt;
&lt;p&gt;The paper builds on Hobijn and Lagakos (2005), who show households consume different bundles, and Kaplan and Schulhofer-Wohl (2017), who find two-thirds of cross-household inflation variation stems from paying different prices for the same goods. D&amp;rsquo;Acunto, Malmendier, Ospina, and Weber (2021) establish that grocery store prices directly influence household inflation expectations. Branch takes these findings as given — they motivate the identifying assumption that expectations reflect basket-specific prices — and focuses on the downstream question of whether those expectations causally raise actual inflation outcomes. Earlier work on heterogeneous expectations by Branch (2004, 2007) using Michigan survey data, finding time-varying heterogeneity across forecasting rules, is also directly referenced.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-rotemberg-weight-decomposition-reveal-about-the-source-of-identifying-variation"&gt;Q11. What does the Rotemberg weight decomposition reveal about the source of identifying variation?&lt;/h3&gt;
&lt;p&gt;The Bartik estimate is a weighted average of 160 just-identified group-specific estimates. Goldsmith-Pinkham, Sorkin, and Swift (2020) show the weights (αg) measure each group&amp;rsquo;s contribution to the overall estimate and sensitivity to bias from that group&amp;rsquo;s potential endogeneity. Tables 4–5 list the top-10 weighted groups: under CPS78 shares, these are predominantly 25–49-year-olds, mostly college-educated, seven of ten married with children. Under Michigan shares, the top groups are even younger (mostly 18–24), with at least a high school degree, almost all married without children. Table 6 shows men receive slightly higher aggregate weight than women (0.53–0.57 vs. 0.43–0.47), and those aged 50+ contribute less than 15% of total weight. Figure 11 shows temporal variation: the heaviest-weighted periods are the late-1970s Great Inflation and Volcker disinflation, the Great Recession (2007–09), and the post-pandemic episode (2021–22).&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper provides empirical support for central bank attention to short-run consumer inflation expectations: a 1 percentage point increase in one-year-ahead regional expectations causally raises regional inflation by 0.33–0.55 basis points (lower bound, since spillovers are excluded). Accounting for cross-regional aggregate effects raises the likely total pass-through to above one, validating the central bank emphasis on anchoring short-run expectations. However, the null finding for long-run (5-to-10-year) expectations — controlling for aggregate time effects — suggests that &amp;lsquo;anchoring long-run expectations&amp;rsquo; may not independently prevent near-term inflation above and beyond its correlation with short-run beliefs. The scope conditions are important: the estimates come from U.S. Census regions over 1978–2022, so applicability to countries with persistently high or hyper-inflation is uncertain. The identifying variation is concentrated in high-volatility inflation episodes, suggesting potential nonlinearities in the expectations-to-inflation mapping. The empirical strategy also does not capture general equilibrium feedback from realized inflation back to expectations.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-data-limitations-and-survey-design-concerns-the-paper-acknowledges"&gt;Q13. What are the data limitations and survey design concerns the paper acknowledges?&lt;/h3&gt;
&lt;p&gt;Five limitations of the Michigan survey are acknowledged: (1) whether surveys elicit genuine expectations rather than attitudes; (2) the rotating panel structure, with roughly 40% of respondents re-interviewed after six months, creates potential selection bias if more accurate forecasters are likelier to re-participate; (3) declining telephone response rates threaten representativeness; (4) the survey prompts respondents reporting &amp;lsquo;unreasonable&amp;rsquo; expectations, with the threshold endogenously tied to recent inflation history; (5) the question wording asks about &amp;lsquo;prices going up&amp;rsquo; rather than &amp;lsquo;aggregate U.S. inflation&amp;rsquo;, making the measure closer to consumption-basket-specific expectations — which the paper treats as a feature rather than a flaw for its identifying assumption. The paper addresses concerns (1)–(4) through alternative specifications (first-time-only respondents, outlier removal, CPS vs. survey shares). The geographic dimension is limited to four Census regions because finer location identifiers are unavailable for a long panel.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Shift-share (Bartik) instrument for expectations&lt;/strong&gt;: In this paper, the instrument for regional inflation expectations is constructed by interacting each demographic group&amp;rsquo;s national-level mean inflation expectation (the &amp;lsquo;shift&amp;rsquo;) with that group&amp;rsquo;s population share in the region (the &amp;lsquo;share&amp;rsquo;). The resulting weighted average predicts how much regional expectations would be elevated purely by the region&amp;rsquo;s demographic composition reacting to aggregate group-level expectation shocks, isolating variation plausibly orthogonal to region-specific inflation supply shocks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Differential exposure quasi-experiment&lt;/strong&gt;: The identification design exploits the fact that U.S. Census regions have different demographic compositions, giving them differential exposure to aggregate shocks in group-specific inflation expectations. Regions with a higher share of a group whose expectations are rising will see a larger predicted increase in regional expectations than regions with a lower share of that group, independent of region-specific factors — this cross-regional contrast is the source of causal identification.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rotemberg weights&lt;/strong&gt;: Following Goldsmith-Pinkham, Sorkin, and Swift (2020), the Bartik 2SLS estimate is decomposed as a weighted sum of 160 just-identified group-specific estimates, where the weight αg for group g measures the sensitivity of the overall estimate to potential endogeneity in group g&amp;rsquo;s share. Groups with large αg drive identification and are the groups most important to probe for exogeneity. In this paper, the heaviest-weighted groups are younger, married consumers with at least a high school degree.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Anticipated utility / steady-state learning&lt;/strong&gt;: The paper&amp;rsquo;s theoretical model allows for non-rational subjective expectations. Firms and households are modeled as &amp;lsquo;anticipated utility&amp;rsquo; maximizers (Woodford 2013) who adjust expectations over time (&amp;rsquo;learning&amp;rsquo;) but assume for current decisions that expected inflation will remain at its present rate — termed &amp;lsquo;steady-state learning&amp;rsquo; by Evans and Honkapohja (2001). This assumption implies future prices evolve along a linear trend from current expectations, yielding a tractable closed-form link between current expectations and the sector-specific price-setting equation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneous consumption baskets as identification&lt;/strong&gt;: The paper&amp;rsquo;s core identifying assumption is that different demographic groups consume different bundles of goods across sectors, so their inflation expectations reflect the price changes in their own basket rather than a common aggregate signal. This basket heterogeneity is what makes group-level expectations differ systematically and allows the shift-share instrument to generate exogenous variation in regional inflation expectations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Lower bound interpretation of regional estimates&lt;/strong&gt;: The 2SLS estimates capture only the regional (within-country, across-region) effect of expectations on inflation, because time fixed effects absorb cross-regional spillovers — if expectations rise in one region, the increased demand for traded goods spills into other regions and raises their prices too. The paper argues the regional estimates are therefore a lower bound on the aggregate pass-through from expectations to overall U.S. inflation, consistent with the stronger aggregate correlation seen in Figure 1a.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Long-run expectations nullity&lt;/strong&gt;: The paper&amp;rsquo;s extension finds that 5-to-10 year inflation expectations, instrumented with their own shift-share Bartik and included alongside the one-year instrument, have no statistically or economically significant causal effect on regional inflation once time fixed effects control for aggregate factors. This result implies that, conditional on short-run expectations and macroeconomic controls, long-horizon expectations carry no independent causal information for the current inflation rate.&lt;/p&gt;</description></item><item><title>Illuminating the Global South</title><link>https://macropaperwarehouse.com/papers/illuminating-the-global-south/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/illuminating-the-global-south/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Satellite nighttime lights (luminosity) are the dominant remote-sensing proxy for local economic conditions in low-income countries, yet their accuracy at fine spatial scales and over time has remained contested. This paper by Chiovelli, Michalopoulos, Papaioannou, and Regan makes two linked contributions. First, it constructs a standardized, annual, global panel of nighttime lights from 1992 to 2023, integrating the legacy DMSP-OLS satellite series (1992–2013) with the higher-quality VIIRS series (2013–onward) after applying three adjustments to the noisier DMSP data: cross-sensor inter-calibration (following Li et al. 2020), top-coding correction (following Bluhm and Krause 2022, using a truncated Pareto distribution to replace pixels with Digital Number ≥ 55), and blooming correction (following Cao et al. 2019, modeling light spillover as spatial decay and subtracting predicted pseudo-light). VIIRS is then downgraded to DMSP-comparable units using an ensemble machine-learning method — extremely randomized trees trained on the single year of full overlap (2013) — yielding an out-of-sample RMSE of 1.50 versus 3.27 for the Li et al. sigmoid approach and 1.57 for the Nechaev et al. convolutional neural network; the F1 score for the binary lit/unlit classification is 0.72 versus 0.51 and 0.71 for those alternatives, with recall = 0.95 and precision = 0.58 against an actual lit-pixel share of only 8.6 percent globally. At the cross-country level — a sample of 173 countries — the adjusted series retains an elasticity of luminosity to GDP of approximately 0.85 and an R² around 0.9 in cross-section; for Africa specifically the elasticity is 0.7 and R² remains around 0.9. In long-difference panel regressions over 1992–2019, the luminosity-GDP elasticity is approximately 0.25–0.24, broadly consistent with Henderson et al. (2012)&amp;rsquo;s estimate of 0.30–0.33, while at the five-year panel frequency the elasticity is around 0.15–0.17. The second contribution is a systematic validation of the new series against multiple local development proxies across four low-income settings. Using 139 georeferenced DHS surveys from 34 African countries (gridcells of ~28km × 28km), the adjusted series yields cross-sectional coefficients of approximately 0.6 standard deviations for schooling, electricity access, and improved sanitation, and approximately 1 standard deviation for the composite wealth index, between lit and unlit gridcells; in within-gridcell panel regressions, the adjusted log-lights coefficient on schooling is approximately double that of the unadjusted series (~0.02 versus ~0.01), and lit/unlit panel coefficients are statistically significant only with the adjusted series — gridcells turning lit see schooling rise by ~0.05 standard deviations (~0.125 schooling years), wealth index rise by ~0.05 SD, and electricity access rise by ~0.05 SD. In Mozambique, using all post-civil-war censuses (1997, 2007, 2017) across 1,126 admin-4 localities, schooling and non-agricultural employment are at least 0.5 standard deviations higher in lit than unlit localities, equivalent to approximately 0.5 years of schooling and 10 percentage points of non-agricultural employment; within-locality changes in lights co-move significantly with schooling changes, with the difference in schooling gain between localities that turn lit versus stay unlit being about half a year even controlling for admin-3 fixed effects. In Indonesia, panel estimates for public goods across more than 60,000 PODES villages show the adjusted series yields a positive and significant coefficient on the composite wealth index while the unadjusted series yields a counterintuitively negative coefficient. In India, across more than 550,000 SHRUG villages and towns, the adjusted series consistently produces stronger cross-sectional and panel associations with non-farm, manufacturing, and services employment. A key empirical regularity across all settings is that the adjusted series outperforms the unadjusted one most sharply at finer spatial resolutions and in over-time (panel) comparisons, while at coarse aggregation levels (large administrative units or large grid squares) differences between the two series are minor, as spatial averaging attenuates measurement error in the unadjusted data too. Blooming correction delivers most of the improvement in the African context, where top-coding is rare (fewer than 2% of lit DMSP pixels in Africa approach the 63 DN ceiling). The paper also replicates three canonical studies — Michalopoulos and Papaioannou (2013) on precolonial ethnic institutions, Michalopoulos and Papaioannou (2014) on national institutions and split ethnic homelands, and Hodler and Raschky (2014) on regional favoritism — confirming that qualitative conclusions are robust to the data revision while documenting that the adjusted series sharpens several estimates, particularly those exploiting within-region over-time variation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper is a measurement and validation study rather than a causal identification exercise. Its core design is correlational: it regresses local development proxies on nighttime luminosity across gridcells and administrative units, conditioning on country-year fixed effects in cross-section and on unit fixed effects in panel regressions. The main threats are (a) reverse causation (luminosity and development are jointly determined), which the authors acknowledge but do not attempt to address — they are explicit that the goal is proxy validation, not causal estimation; (b) measurement error in both the luminosity variable and the development outcomes (DHS wealth index, census schooling, PODES public goods), which the paper addresses by comparing adjusted versus unadjusted luminosity series and interpreting attenuation bias reduction as evidence of improved measurement; (c) the binary transformation of luminosity (lit/unlit) produces non-classical measurement error — an explicit point drawn from econometric theory (Aigner 1973; Meyer and Mittag 2017) — which partly motivates the adjusted continuous series; and (d) spatial autocorrelation and systematic geographic patterns in prediction error, which the authors check by regressing prediction errors on latitude and longitude and find that the ERT-downgraded series reduces the latitude coefficient to 10% of its magnitude in the unadjusted VIIRS specification for log lights and to 35% for the lit indicator.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-dmsp-deficiencies-corrected-and-what-are-the-specific-methods-used"&gt;Q2. What are the three DMSP deficiencies corrected and what are the specific methods used?&lt;/h3&gt;
&lt;p&gt;Cross-sensor inter-calibration: DMSP data come from six satellites; Li et al. (2020) supply a cross-calibrated series using a second-order polynomial fitted on overlapping satellite years, which the paper adopts as its &amp;lsquo;unadjusted&amp;rsquo; baseline. Top-coding: DMSP records 8-bit Digital Numbers (DN) 0–63, so radiance above a ceiling is truncated. Pixels with DN ≥ 55 are subject to &amp;lsquo;implicit&amp;rsquo; top-coding (averages of potentially top-coded sub-readings). The correction uses the radiance-calibrated (RC) vintage available for seven years, ranks the top-coded pixels by the RC series from the nearest year, then replaces them with &amp;lsquo;structural values&amp;rsquo; drawn from a truncated Pareto distribution with parameters α = 1.5, L = 55, H = 2000. Blooming: the DMSP sensor stretches edge pixels and can be spatially displaced up to 3 km, causing light spillover. Following Cao et al. (2019), pseudo-light pixels (PLPs) — lit pixels neighboring at least one dark pixel — are identified. An OLS regression of PLP light on the inverse-squared-distance weighted sum of neighbors&amp;rsquo; light within a 7 × 7 window is estimated separately for broad global regions. The predicted blooming contribution is subtracted from each lit pixel, negative residuals are set to zero, and a local 3 × 3 mean smoothing is applied. Globally, the blooming correction raises the share of unlit pixels from 92% to 95% in 1992 and from 88% to 91% in 2012.&lt;/p&gt;
&lt;h3 id="q3-how-is-viirs-downgraded-and-harmonized-with-dmsp-and-what-does-extremely-randomized-trees-mean"&gt;Q3. How is VIIRS downgraded and harmonized with DMSP, and what does &amp;rsquo;extremely randomized trees&amp;rsquo; mean?&lt;/h3&gt;
&lt;p&gt;Because VIIRS records 14-bit DN at 15-arc-second resolution with far superior sensor quality, it is not directly comparable to the 8-bit, 30-arc-second DMSP. The authors&amp;rsquo; preferred approach downgrades VIIRS to match the DMSP scale. They use an ensemble machine-learning method called &amp;rsquo;extremely randomized trees&amp;rsquo; (Geurts et al. 2006), a variant of random forests that, instead of choosing the best splits from the training sample, picks split thresholds randomly, which further reduces variance and improves computational efficiency. Features used to predict DMSP-like values from VIIRS include: pixel statistics (mean, median, min, max of the four VIIRS sub-pixels within each DMSP 30-arc-second cell), statistics of neighboring pixels within windows of 3, 4, 7, 9, 11, 13, 17, and 21 pixel widths, and regional dummies for broad world regions. The model is trained on 2013 (the one full year of DMSP-VIIRS overlap) and its out-of-sample performance is assessed by retraining on 2012 and predicting 2013. Four merged series are produced corresponding to the four versions of DMSP (unadjusted; blooming only; top-coding only; both). The authors&amp;rsquo; approach outperforms both the Li et al. (2020) sigmoid-function method (RMSE 3.27 globally vs. 1.50) and the Nechaev et al. (2021) CNN approach (RMSE 1.57), especially in the low-to-middle luminosity range most relevant for low-income countries.&lt;/p&gt;
&lt;h3 id="q4-what-development-proxies-are-used-in-validation-and-across-what-samples"&gt;Q4. What development proxies are used in validation and across what samples?&lt;/h3&gt;
&lt;p&gt;Africa (DHS, 34 countries, 139 surveys, ~28km × 28km gridcells): mean years of schooling (respondents aged 15–39), DHS composite household wealth index, share of households with improved sanitation, share with electricity connection. All outcomes are standardized to mean zero, SD one. Mozambique (Census 1997, 2007, 2017, 1,126 admin-4 localities): mean years of schooling (aged 15–39) and non-agricultural employment (aged 15–24 or 19–24). Indonesia (PODES village census waves 1996–2018, 60,000+ villages): binary measures for garbage disposal, toilet use, drinking water access, gas/electricity for cooking, paved roads, and counts of kindergartens, primary, middle, and secondary schools — aggregated into a first principal component (eigenvalue ~3.5, capturing ~1/3 of variance). India (SHRUG dataset, 550,000+ towns and villages, Population Censuses 1991/2001/2011, Economic Censuses 1990/1998/2005/2013): population count, total non-farm employment, manufacturing employment, services employment.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented"&gt;Q5. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Spatial resolution: adjusted series outperforms unadjusted most at fine resolutions (2×2 gridcell blocks, ~56km × 56km at the equator); at coarse levels (12×12 blocks, ~336km × 336km), both series yield similar coefficients, as spatial aggregation attenuates noise in the unadjusted series. Urban vs. rural: cross-sectional estimates are similarly significant in urban and rural DHS samples. Panel estimates are statistically significant only with the adjusted series; urban panel coefficients are consistently larger than rural ones, echoing Asher et al. (2021)&amp;rsquo;s India finding. The adjustment matters more in rural areas than in urban areas in cross-section. Local variation (spatial RDD / fine fixed effects): with unadjusted series, panel wealth-index coefficients are statistically indistinguishable from zero until spatial fixed effects cover areas at least 7×7 gridcells (~200km × 200km at equator); with the adjusted series, coefficients remain significantly positive at all fixed-effect sizes including the finest 2×2 blocks. Top-coding vs. blooming: most of the improvement in Africa derives from blooming correction; top-coding correction has minor impact because fewer than 2% of lit African DMSP pixels approach the DN ceiling. Country-ethnic homelands (large areas, avg. 25,547 km²): adjustments matter little because spatial averaging already reduces noise. Applications replication: the precolonial institutions result (Michalopoulos and Papaioannou 2013) is robust and essentially unchanged because the units are very large. The national-institutions-at-border result (Michalopoulos and Papaioannou 2014) is strengthened in within-ethnicity specifications (coefficient marginally significant at 90% with adjusted series vs. p ≈ 0.15 with unadjusted); capital-proximity heterogeneity is sharpened. The regional-favoritism result (Hodler and Raschky 2014) strengthens: the log-lights lagged-leader coefficient rises from 0.038 to 0.058, and the lit-probability coefficient rises from ~3 to ~7 percentage points.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-and-specification-variations-are-run"&gt;Q6. What robustness checks and specification variations are run?&lt;/h3&gt;
&lt;p&gt;The paper compares four luminosity series (unadjusted Li et al.; blooming only; top-coding only; both combined + VIIRS fusion) to isolate each correction&amp;rsquo;s contribution. It checks the luminosity-GDP nexus at annual, five-year, and long-difference frequencies. It examines seven African countries&amp;rsquo; co-evolution of the harmonized series with electrification share (Kenya, DRC, Ghana, Tanzania, Nigeria, Mozambique, and one other) and finds no discontinuity at the 2012/2013 DMSP-VIIRS transition year. Spatial aggregation robustness: coefficients are computed across aggregation blocks ranging from 2×2 to 12×12 gridcells, showing stability in cross-section (~0.18) and mild size dependence in panel (~0.075, slightly rising with coarser units). Local variation robustness: fixed effects of increasing spatial coverage (2×2 to 12×12 cells) are added while the outcome remains at the gridcell level. Results replicated for schooling and electricity access (Appendix Section B.2) beyond the primary wealth-index outcome. Confounding by latitude in the ML model is assessed via regressions of prediction errors on latitude and longitude with and without country fixed effects. Median regressions confirm the OLS elasticity estimates at the cross-country level. The India analysis is replicated for both towns (urban) and villages (rural) separately.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Henderson et al. (2012): pioneer the use of luminosity as a cross-country GDP proxy and estimate a long-difference elasticity of 0.30–0.33 across 188 countries; this paper estimates 0.25–0.24 over a comparable specification, consistent but slightly lower. Gibson et al. (2021): show that VIIRS is superior to DMSP but find weak GDP-lights correlations outside cities for the early DMSP period in China, Indonesia, and South Africa; this paper addresses the concern by adjusting DMSP and merging it with VIIRS. Asher et al. (2021): validate luminosity as a strong proxy in India and find stronger urban-luminosity links; this paper replicates and extends those findings to Africa, Mozambique, and Indonesia and shows the adjusted series strengthens the Asher et al. patterns. Chen et al. (2024): find strong cross-sectional but weak panel associations; this paper&amp;rsquo;s adjusted series substantially strengthens panel associations. Bluhm and Krause (2022): provide the top-coding correction method adopted here. Cao et al. (2019): provide the blooming correction method. Nechaev et al. (2021): propose a CNN-based DMSP-VIIRS fusion but apply it to the unadjusted DMSP; this paper outperforms their RMSE slightly (1.50 vs. 1.57) and improves on their F1 score (0.72 vs. 0.71), with greater advantage in low-light regions. Li et al. (2020): propose a sigmoid-based fusion calibrated for high-light pixels; this paper substantially outperforms it (RMSE 1.50 vs. 3.27) particularly in low-luminosity areas. The paper thus synthesizes and extends multiple strands: it unifies the corrections of Bluhm-Krause and Cao et al., pairs them with state-of-the-art ensemble ML fusion, and provides by far the most comprehensive multi-country, multi-context validation of the resulting series.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary policy implication is methodological: researchers studying development in low-income countries should use the adjusted and harmonized nighttime lights series rather than raw DMSP data, and should be especially careful at fine spatial scales (e.g., spatial regression discontinuity designs, granular village-level analyses) and in panel specifications. The gains from adjustment are largest precisely where applied development research is moving — toward local identification strategies and over-time variation. For practitioners and statistical agencies, the series provides a low-cost annual proxy for local economic conditions in environments with weak administrative data, particularly across sub-Saharan Africa, South Asia, and Southeast Asia. Scope conditions: (a) Correlations are far from perfect — binary lit/unlit classification misses much variation in the many-zeros low-income context. (b) At large aggregate units (admin-1, country-ethnic homelands), the adjustments yield minimal additional improvement since noise averages out. (c) The series does not resolve the fundamental limitation that most of sub-Saharan Africa remains unlit (98.4% of DMSP pixels in Africa in 1992), so it captures variation among already-lit areas better than the development gradient at the zero-light frontier. (d) Future research blending nighttime lights with daytime imagery (traffic, built structures) is flagged as a promising extension, though daytime data are often proprietary.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-main-findings-from-the-three-replication-exercises"&gt;Q9. What are the main findings from the three replication exercises?&lt;/h3&gt;
&lt;p&gt;Michalopoulos and Papaioannou (2013) — precolonial ethnic institutions and contemporary development: Replication across 682 country-ethnic homelands confirms that areas with higher precolonial political centralization (as measured by a 0–4 jurisdictional hierarchy index) have significantly higher contemporary luminosity, conditional on country constants and geographic controls. With the adjusted series, the unlit share among homelands rises from 24% to 29% (because blooming correction removes spurious light), but the coefficients on political centralization are still highly significant, somewhat smaller in magnitude, and similar qualitatively. The main conclusion is robust because the units are large and spatial averaging already reduces noise in the raw series. Michalopoulos and Papaioannou (2014) — national institutions and split-border ethnic development: Replication across 38,427 gridcells of 220 systematically partitioned ethnic homelands. Cross-sectional results show a one-point increase in the rule-of-law index (range −2.5 to 2.5) is associated with a ~10 pp higher probability of a gridcell being lit. The within-ethnicity coefficient drops by more than half (~0.025). With the adjusted series, this within-ethnicity coefficient is marginally significant at 90% versus a p-value of ~0.15 with unadjusted. Spatial RDD coefficients remain small and insignificant regardless of adjustment. Capital-proximity heterogeneity: the positive association between rule of law and luminosity is significant only for ethnically split groups where both portions are close to their respective capitals, and this finding is more precisely estimated with the adjusted series; the effect is nil far from capitals in both series. Hodler and Raschky (2014) — regional favoritism: Panel replication across 38,427 subnational regions in 126 countries, 1992–2009. The lagged-leader dummy coefficient (log lights specification) rises from 0.038 to 0.058 with the adjusted series. The linear-probability-model lit indicator rises from ~3 to ~7 percentage points. All specifications with the adjusted series are at least two standard errors above zero, matching or exceeding the precision of the original.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-limitations-and-caveats-acknowledged-by-the-authors"&gt;Q10. What are the limitations and caveats acknowledged by the authors?&lt;/h3&gt;
&lt;p&gt;First, the correlations between luminosity and development are &amp;lsquo;far from perfect&amp;rsquo; — the binary lit/unlit transformation in particular fails to capture the significant continuous variation in assets, education, and public goods across regions that are all formally &amp;rsquo;lit.&amp;rsquo; Second, bottom-coding (under-recording of low-light areas) is acknowledged but not corrected; no existing method addresses it, though the authors note that their corrections nonetheless improve elasticities even in rural African regions with very low light. Third, downgrading VIIRS to DMSP by construction sacrifices some of the VIIRS data quality; the long-difference VIIRS elasticity for Africa (0.4) shrinks to 0.35 in the downgraded series. Fourth, daytime satellite imagery and combinations with nighttime lights (Jean et al. 2016; Yeh et al. 2020; Rossi-Hansberg and Zhang 2025) can better capture local wealth but are often proprietary and not replicable in standard economic research. Fifth, the top-coding correction in Africa is minor because very few pixels approach the DN=63 ceiling (0.98–1.7% of lit pixels in 1992–2012), so the main African improvement comes from blooming; other regions with denser urban cores may benefit more from top-coding correction. Sixth, the cross-sensor inter-calibration step is taken &amp;lsquo;off-the-shelf&amp;rsquo; from Li et al. (2020) and further investigation of sensor calibration is left to future work.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Top coding (DMSP)&lt;/strong&gt;: The truncation of Digital Number values at the 8-bit ceiling of 63 in DMSP-OLS data, caused by sensor calibration for cloud detection. Pixels with DN ≥ 55 also suffer &amp;lsquo;implicit&amp;rsquo; top coding because they represent averages of multiple potentially top-coded sub-readings. The paper corrects this by replacing top-coded pixels with structural values drawn from a truncated Pareto distribution, using the radiance-calibrated DMSP vintage to rank pixels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Blooming (spatial spillover of light)&lt;/strong&gt;: A measurement artifact in DMSP data whereby light from bright pixels spills into neighboring dark areas due to the sensor&amp;rsquo;s imprecise spatial accuracy and possible displacement of up to 3 km. The paper identifies pseudo-light pixels (lit pixels adjacent to at least one dark pixel), models the spillover as an inverse-squared-distance weighted function of neighboring lights, and subtracts the predicted blooming from each lit pixel. This correction raises the global unlit pixel share from 92% to 95% in 1992.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extremely randomized trees (ERT)&lt;/strong&gt;: An ensemble machine-learning method used to downgrade VIIRS luminosity data to the DMSP scale. Unlike standard random forests that find the best split thresholds within a random feature subset, ERT selects split thresholds randomly, reducing variance and improving computational efficiency. The authors train it on pixel statistics (mean, median, min, max) and neighborhood statistics within windows of varying sizes to predict DMSP-like values for 2014 onward from VIIRS readings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Harmonized (adjusted + fused) luminosity series&lt;/strong&gt;: The authors&amp;rsquo; main output: an annual global panel of nighttime lights from 1992 to 2023 that applies inter-sensor calibration, top-coding correction, and blooming correction to DMSP data (1992–2013), then uses the ERT ensemble model to convert post-2013 VIIRS data into DMSP-comparable units, yielding four variants (unadjusted, blooming only, top-coding only, both corrections) merged into a continuous time series at 30-arc-second (~1 km²) resolution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pseudo-light pixels (PLPs)&lt;/strong&gt;: In the blooming correction procedure, PLPs are defined as lit pixels (DN &amp;gt; 0) that have at least one dark neighbor (DN = 0). They are the pixels most likely to contain spurious light from neighboring bright areas. PLP light values are regressed on the inverse-squared-distance weighted sum of surrounding pixels to estimate the blooming decay function.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DHS composite wealth index&lt;/strong&gt;: Used in the validation analysis as a local development proxy: a principal-component aggregation of household characteristics including roof quality and ownership of consumer assets, constructed by the Demographic and Health Surveys program across African countries. The paper standardizes this and other outcomes to mean zero and standard deviation one for cross-outcome coefficient comparisons.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Spatial RDD (regression discontinuity design) using nighttime lights&lt;/strong&gt;: As applied in Michalopoulos and Papaioannou (2014) and referenced throughout, a design that restricts estimation to gridcells within a narrow band (e.g., 50 km) of a political or administrative border to compare otherwise similar areas on opposite sides, using luminosity as the outcome. The paper notes that such fine-resolution, localized comparisons are exactly the setting where measurement error in the unadjusted DMSP series is most consequential and where the adjusted series yields the largest improvement.&lt;/p&gt;</description></item><item><title>Import Liberalization as Export Destruction? Evidence from the United States</title><link>https://macropaperwarehouse.com/papers/import-liberalization-as-export-destruction-evidence-from-the-united-states/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/import-liberalization-as-export-destruction-evidence-from-the-united-states/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; How does import liberalization affect a country&amp;rsquo;s &lt;em&gt;export&lt;/em&gt; performance and welfare? Economic theory (Graham 1923, Ethier 1982, Krugman 1984) shows the answer hinges on whether production exhibits increasing returns to scale at the sector level. Krugman (1984) argued that with scale economies, import protection can be export-promoting because a protected industry expands, exploits scale economies, becomes more productive, and exports more — so conversely import liberalization is &amp;ldquo;export destroying.&amp;rdquo; The paper turns this logic into an empirical test: the sign of the import-liberalization-to-export relationship discriminates between constant-returns and increasing-returns trade models. Researchers otherwise lack tools to choose between these model classes, yet the choice matters greatly for multi-sector trade policy analysis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model and data.&lt;/strong&gt; The authors build a multi-sector general-equilibrium gravity model generalizing Krugman (1980) to many countries/sectors with input-output linkages (as in Caliendo-Parro 2015). The model nests constant returns (Armington, σ→∞) and increasing returns. The &amp;ldquo;scale elasticity&amp;rdquo; is 1/(σ−1); the &amp;ldquo;output elasticity&amp;rdquo; of exports equals the trade elasticity (ε−1) times the scale elasticity, and is positive iff there are increasing returns. The empirical application exploits US Permanent Normal Trade Relations with China (PNTR), passed Oct 2000, which removed tariff-revocation uncertainty. Exposure is measured by Pierce-Schott&amp;rsquo;s NTR gap (log gap between non-NTR and NTR tariffs; mean 0.23, SD 0.13, range 0–0.59). Trade data are from CEPII BACI; the baseline sample covers exports from 23 OECD countries (including the US) to 141 importers across 444 NAICS goods industries, in long differences (1995–2000 pre-period vs 2000–07 post-period).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings.&lt;/strong&gt; Reduced-form: US export growth fell in higher-NTR-gap industries after PNTR. The raw Figure 1 slope is −0.51 (SE 0.057); a 10-log-point NTR-gap increase is associated with 5.0 log points lower annual export growth, and the NTR gap explains 18% of cross-industry variation. This is inconsistent with constant returns and implies increasing returns in US goods production. An offsetting &lt;em&gt;input cost effect&lt;/em&gt; (lower imported-input costs) raises exports: PNTR reduced 2007 exports by 13% more for a 75th- vs 25th-percentile NTR-gap industry, but raised them 20% more for a 75th- vs 25th-percentile input-cost-shock industry; net effects range from −18% (Cigarettes) to +56% (Automobiles). A structural IV (NTR gap instrumenting output growth) yields an output elasticity of 0.74 (SE 0.41, preferred column).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantitative GE results.&lt;/strong&gt; Calibrating the output elasticity to 0.821 (matching the −0.10 conditional NTR-gap effect; trade elasticity set to 5), PNTR raised aggregate US exports/GDP by 3.2%, decomposed into −1.8% real market potential (export destruction), +2.4% input cost, and +2.7% foreign demand. Aggregate export growth is 28% larger with scale economies than without, because scale economies make the input-cost effect almost five times stronger (2.4% vs 0.5%). Exports nevertheless declined in the most exposed sectors (Textiles &amp;amp; Leather, Other Manufacturing), shifting US comparative advantage away from high-NTR-gap sectors. Welfare: PNTR raised US real income 0.068% (real expenditure 0.087%); gains are ~30% smaller than under constant returns because a negative specialization effect (−0.15%) offsets a larger ACR openness gain (0.22%). Chinese gains exceed US gains tenfold.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-theoretical-test-and-why-does-the-sign-of-the-import-liberalization-to-export-relationship-identify-returns-to-scale"&gt;Q1. What is the core theoretical test and why does the sign of the import-liberalization-to-export relationship identify returns to scale?&lt;/h3&gt;
&lt;p&gt;From the bilateral trade equation, the elasticity of exports to output equals the output elasticity (ε−1)/(σ−1), which is strictly positive iff there are increasing sector-level returns. Under constant returns (Proposition 1), conditional on foreign demand and domestic input costs, import liberalization does not affect exports (α1=0). Under increasing returns (Proposition 2), import liberalization shrinks domestic real market potential, lowers output, and — because productivity falls with output under scale economies — reduces exports to ALL destinations (α1&amp;lt;0), with the effect&amp;rsquo;s magnitude strictly increasing in the output elasticity. So estimating whether export growth falls in more-liberalized industries distinguishes the two model classes.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identification-strategy-and-its-main-threats"&gt;Q2. What is the identification strategy and its main threats?&lt;/h3&gt;
&lt;p&gt;A triple-difference: changes in US bilateral export growth by sector after PNTR relative to changes in other OECD exporters&amp;rsquo; growth, identified from the NTR gap interacted with Post and a US-exporter dummy. The estimating equation (12) uses importer-exporter-industry, importer-exporter-period, and importer-industry-period fixed effects to absorb importer demand, common-across-exporter technology shocks, and industry trends in supply capacity and trade costs. The NTR gap is plausibly exogenous because variation stems mostly from Smoot-Hawley (1930) non-NTR tariffs, unlikely related to economic conditions 70 years later; any endogeneity from NTR tariffs being higher in weak-growth industries would bias against finding a negative effect. Threat 1: unobserved US-specific technology shocks negatively correlated with the NTR gap not captured by input/skill/capital intensity controls. Addressed by re-estimating at HS 6-digit level with NAICS-industry-exporter-period fixed effects (Table 3), still finding negative effects. Threat 2: US-China competition in third markets — if PNTR shifted China&amp;rsquo;s export basket toward US-type products in high-NTR-gap industries. Tested by interacting with China&amp;rsquo;s market share (Table 4); the quadruple interaction is positive and insignificant, ruling this out.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-three-mechanisms-and-how-are-they-distinguished-empirically-and-quantitatively"&gt;Q3. What are the three mechanisms and how are they distinguished empirically and quantitatively?&lt;/h3&gt;
&lt;p&gt;(1) Real market potential / export destruction: import liberalization lowers the US price index, makes the domestic market more competitive, shrinks real market potential and output, and (under scale economies) cuts productivity and exports — identified by the negative α1 on the NTR gap. (2) Input cost effect: lower imported-input costs cut production costs and raise exports — identified by α2 on the input-output-weighted upstream NTR gap (CostShock), found negative and significant (lower input costs → higher exports). (3) Foreign demand effect: GE expansion of global demand and the trade-balance link between imports and exports — absorbed by fixed effects in the regression but recovered in the calibrated model&amp;rsquo;s decomposition (equation 16). In GE: −1.8% (market potential), +2.4% (input cost), +2.7% (foreign demand), netting +3.2%.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented"&gt;Q4. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Sector-level: the real market potential effect is negative in all goods sectors and stronger where the NTR gap is higher; the input cost effect is positively correlated with the NTR gap (due to heavy diagonal weight in the I-O table); the foreign demand effect is positive everywhere but uncorrelated with the NTR gap. Net exports/GDP rise in 12 of 15 goods sectors but fall in the highest-NTR-gap sectors — Textiles &amp;amp; Leather falls 22% (−32% market potential, +8.5% input cost, +4.6% foreign demand) and exports decline in 3 of the 4 highest-NTR-gap sectors. Under constant returns, by contrast, export growth is positive in all sectors and weakly POSITIVELY correlated with the NTR gap — qualitatively opposite. The correlation between sector-level export growth with vs without scale economies is insignificant (excluding Textiles &amp;amp; Leather) or significantly negative (including it).&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Appendix C checks robustness to: starting the post-period in 2001 instead of 2000; alternative NTR-gap definitions; aggregating exports across destinations; varying the exporter/importer/industry samples; allowing PNTR to affect domestic expenditure; and controlling for China import growth driven by non-PNTR shocks. An event study (equation 13, Figure 2) shows no NTR-gap/export relationship before 2000 and a negative one from 2001 until the 2007–08 financial crisis, ruling out pre-trends. The first-stage (Table 5) confirms higher-NTR-gap industries had lower OUTPUT growth (paralleling Pierce-Schott&amp;rsquo;s employment result). Alternative calibrations (Appendix D.5): without I-O linkages the market potential effect weakens but total export growth is roughly unchanged; allowing services scale economies raises US gains; combining Textiles &amp;amp; Leather with Other Manufacturing preserves results; using Bartelme et al. (2019) sector-varying elasticities still yields a negative specialization effect.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-output-elasticity-calibrated-and-how-does-it-compare-to-the-structural-estimate"&gt;Q6. How is the output elasticity calibrated and how does it compare to the structural estimate?&lt;/h3&gt;
&lt;p&gt;The output elasticity for goods is calibrated to 0.821 by matching the simulated NTR-gap effect to the −0.10 conditional reduced-form estimate (Table 2, column i), with services output elasticity set to zero and trade elasticity (ε−1) set to 5 (Head-Mayer 2014). This is below the value of 1 implied by Krugman (1980) or the Pareto-Melitz model but close to the Bartelme et al. (2019) mean of 0.83. It is reassuringly close to the independent structural IV estimate of 0.74 (SE 0.41). The simulated effect is decreasing in the output elasticity (consistent with Proposition 2 part ii) and rises sharply as the elasticity approaches one; the model has a unique solution for output elasticities below 0.95.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-welfare-decomposition-work-and-why-are-gains-smaller-with-scale-economies"&gt;Q7. How does the welfare decomposition work and why are gains smaller with scale economies?&lt;/h3&gt;
&lt;p&gt;Following Costinot-Rodríguez-Clare (2014), real-income gains decompose into an ACR term (changes in domestic expenditure share / trade openness) and a specialization term that exists only with scale economies (welfare from sectoral reallocation of employment, weighted by adjusted Leontief forward-linkage coefficients). With scale economies the ACR effect is +0.22% (vs +0.10% without), but it is more than offset by a −0.15% specialization effect, netting +0.068% real income — about 30% below the constant-returns gain. The specialization effect is negative because PNTR shifted resources toward services (weaker scale economies; goods output −0.55%, services +0.11%) and, more importantly per Appendix D.5, toward sectors with weaker FORWARD input-output linkages; cross-sectoral heterogeneity in scale economies alone contributes negligibly.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-relate-to-and-differ-from-closely-related-prior-work"&gt;Q8. How does this relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It extends Krugman (1984)&amp;rsquo;s partial-equilibrium oligopoly mechanism to a class of quantitative GE trade models (love-of-variety, external economies, Melitz-Pareto, or endogenous innovation — shown equivalent in Appendix A.3). Unlike prior scale-economy estimates (Antweiler-Trefler 2002, Lashkaripour-Lugovskyy 2018, Bartelme et al. 2019) and home-market-effect tests (Davis-Weinstein 2003, Costinot et al. 2019), it uses TRADE POLICY variation (not factor content, market size, or exchange rates) for identification and performs an ex-post policy analysis (echoing Goldberg-Pavcnik 2016). Relative to the PNTR/China-shock literature (Pierce-Schott 2016, Handley-Limão 2017, Autor-Dorn-Hanson 2013), it adds a new outcome — US EXPORTS and comparative advantage — and argues the &amp;lsquo;surprisingly swift&amp;rsquo; manufacturing decline would have been smaller absent scale economies. It complements Juhász (2018)&amp;rsquo;s infant-industry evidence (Napoleonic France) by quantifying the export-destruction cost while showing PNTR&amp;rsquo;s net effect on exports and welfare is positive. Dick (1994) tested the same hypothesis cross-sectionally for 1970 US data but found little support.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The findings support the existence of the scale-economies channel traditionally invoked to justify protection: pre-PNTR import protection shifted US comparative advantage toward the most-protected industries, and in the calibrated model targeted import protection CAN promote sector-level exports — but not under constant returns. However, the export-destruction effect is dominated, for most sectors and in aggregate, by export-promoting channels (input cost, foreign demand); total export growth is even greater WITH scale economies; and the negative specialization effect is more than offset by traditional gains from trade, so US gains from PNTR remain positive (+0.068% real income). Scope conditions: results rest on the calibrated output elasticity (0.821) and trade elasticity (5); the model assumes constant markups and full employment, so welfare excludes pro-competitive effects (Jaravel-Sager 2020, Amiti et al. 2020) and employment effects (Autor-Dorn-Hanson 2013); it studies a single liberalization episode; and the analysis cannot distinguish among alternative SOURCES of increasing returns. The authors stress accounting for scale economies (or their absence) is a prerequisite for correctly evaluating sector-level trade flows and welfare.&lt;/p&gt;
&lt;h3 id="q10-what-other-notable-findings-or-caveats-appear"&gt;Q10. What other notable findings or caveats appear?&lt;/h3&gt;
&lt;p&gt;PNTR is calibrated as a reduced-form openness shock (α5=0.43; equation 15), equivalent to a 13% average trade-cost reduction on US imports from China (SD 6.6% across industries) given trade elasticity 5 — matching Handley-Limão&amp;rsquo;s 13-percentage-point estimate. The calibrated economy has 12 economies and 24 sectors (15 goods). Chinese gains exceed US gains more than tenfold (because the US was much larger in 2000, so PNTR was a bigger shock to China), and China&amp;rsquo;s nominal wage rose 6.0% relative to the US, contributing to factor-price convergence. For comparison, Caliendo-Parro (2015) find NAFTA raised US welfare 0.08% and Fajgelbaum et al. (2020) find the Trump trade war cut US real income 0.04%. The model in changes is solved via exact hat algebra, holding each country&amp;rsquo;s trade deficit as a constant share of global value-added (which induces the positive import-export link in the foreign-demand term).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Information and the Formation of Inflation Expectations by Firms: Evidence from a Survey of Israeli Firms</title><link>https://macropaperwarehouse.com/papers/information-and-the-formation-of-inflation-expectations-by-firms-evidence-from-a-survey-of-israeli-firms/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/information-and-the-formation-of-inflation-expectations-by-firms-evidence-from-a-survey-of-israeli-firms/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; How do firms form and update inflation expectations during a monetary-policy regime change and a transition from high/volatile inflation to a low, stable, inflation-targeting environment? This matters because tracking and managing expectations is central to modern monetary policy (especially under forward guidance), yet high-quality firm-level expectations data—particularly across regime changes—are scarce (Bernanke 2007). A central tension in the literature is that firms and households in long-stable advanced economies are largely inattentive to inflation and monetary policy, plausibly because successful stabilization removes the incentive to monitor them. Israel offers a natural experiment: its recent history of high inflation and dollarization, followed by disinflation, de-dollarization, and the anchoring of expectations at the ~2% target midpoint around 2003.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and design.&lt;/strong&gt; The authors use the Bank of Israel Firms&amp;rsquo; Survey, a quarterly survey (quantitative inflation-expectation questions added in 1997), covering six industries (post-2009 shares: manufacturing 36%, services 36%, commerce 14%, transportation/communications 5%, hotels 5%, construction 4%). The main analysis sample is 2001Q3–2018Q3. The survey is voluntary, unbalanced, not nationally representative; late-sample participation fell to ~250–300 firms with a response rate around 30%. Identification exploits within-quarter variation in response timing: because Israel&amp;rsquo;s CPI is published monthly on the 15th and policy-rate decisions are scheduled, firms responding after a release (&amp;ldquo;treatment&amp;rdquo;) had information that firms responding earlier (&amp;ldquo;control&amp;rdquo;) did not. Surprises are defined relative to professional forecasters&amp;rsquo; mean expectations: an inflation (CPI) surprise and a monetary (policy-rate) surprise. Identification assumes response timing is random; the authors show firm characteristics generally do not predict either response period (Table 4) or the cross-section of expectations (Table 3). Estimation uses two-way (firm and quarter) fixed-effects panel regressions interacting treatment dummies with surprise size, plus a lagged dependent variable; local projections (Jordà 2005) first show output/employment respond to the shocks, motivating that beliefs should too.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main quantitative findings (Table 9, full sample 2001Q3–2018Q3).&lt;/strong&gt; A positive inflation surprise of one percentage point raises 1-year inflation expectations by about 0.5 pp from the second-monthly-CPI surprise (coefficient 0.467) and about 0.7 pp from the third-monthly-CPI surprise (0.700). The effect on 1-quarter expectations is weaker (≈0.12 and ≈0.29). Because the annual response exceeds the quarterly response, firms on average treat CPI surprises as persistent, not transitory. A surprise one-percentage-point hike in the policy rate lowers 1-year inflation expectations by about 0.3 pp (coefficient 0.343, negative sign) and 1-quarter expectations by roughly 0.15 pp. The mean second-month-CPI treatment dummy itself is small (-0.07 pp), so the interaction terms carry the economic content.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanisms and scope conditions.&lt;/strong&gt; The inflation-surprise result is robust across sub-periods, before/after 2010, firm sizes, and industries. The monetary-surprise result is NOT robust: dropping the large 2001–2002 policy shocks (sample 2002Q3–2018Q3) renders it insignificant and sign-flipped, consistent with policy shocks having little effect on beliefs in stable environments (Coibion et al. 2020; Ilek 2021 for Israeli forecasters). Implication: even after de-dollarization and prolonged low/stable inflation, Israeli firms keep monitoring macro news; (re)anchoring expectations—making them insensitive to news—may take a long time, an insight relevant for countries now facing high inflation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The strategy exploits variation in survey response timing within each quarter. Because Israel publishes CPI on the 15th of each month and policy-rate decisions are on scheduled dates, firms that respond after a release (treatment) have seen information that firms responding earlier (control) have not. Responses are grouped into Periods 1, 2, 3 (and Period 0 for missing/late dates), generating two CPI surprises (second- and third-monthly index) and one interest-rate surprise per quarter. The key identifying assumption is that response timing is as-good-as random. The main threat is selection—if attentive or expectation-distinctive firms systematically respond later, treatment status would be endogenous. The authors address this by regressing exposure-period indicators on observable firm characteristics (Table 4) and finding characteristics generally do not predict response period; they also confirm firm characteristics do not explain cross-sectional expectation levels (Table 3). A placebo test replacing the dependent variable with the prior quarter&amp;rsquo;s expectation (t-1) finds no effect (Appendix Table B5), supporting the timing identification. A residual threat is unobservable correlates of timing not captured by observables.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;Two mechanisms: (1) firms update inflation expectations to new CPI information, and (2) firms update to monetary-policy information. They are distinguished by using separate, independently timed surprises (CPI releases vs. policy-rate decisions) and separate interaction terms. Persistence vs. transitory perception is inferred from the horizon pattern: because the 1-year response to a CPI surprise (~0.5–0.7 pp) exceeds the 1-quarter response (~0.12–0.29 pp), firms must expect the price increase to continue over subsequent quarters, i.e., they perceive CPI shocks as persistent. For monetary policy, the smaller 1-quarter than 1-year effect is read as consistent with monetary policy operating with a lag. The output/employment local projections (Table 8) show a non-monotonic response to rate surprises (rises in quarters 0–1, declines in quarters 2–3), which the authors note could mix conventional contractionary effects with an information effect (a higher rate signaling a stronger economy).&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;By firm size (Table 11): all three size groups (small, medium, large) respond to CPI surprises on 1-year expectations and the differences across groups are generally not statistically significant; the interest-rate-surprise effect resembles the pooled estimate for medium and large firms but is not statistically significant for small firms. By industry (Table 12): the CPI-surprise effect on 1-year expectations is positive and statistically significant in nearly every industry, whereas the interest-rate-surprise effect on 1-year expectations (full sample) is negative and significant only in manufacturing. Over time (Table 10): the 1-year CPI-surprise effect is almost identical before and after 2010 (the year the monetary committee was established), and the 1-quarter effect is similar or if anything stronger in the later period. Cross-sectionally, firm size, industry, and region are mostly statistically and economically insignificant predictors of expectation levels (Table 3).&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;(1) Shorter sample 2002Q3–2018Q3 excluding the large 2001–2002 policy shocks—CPI-surprise results essentially unchanged, monetary-surprise results become insignificant and change sign. (2) Split before/after 2010 allowing time-varying effects (Table 10). (3) Heterogeneity by size (Table 11) and industry (Table 12) as consistency checks. (4) A placebo test regressing the previous quarter&amp;rsquo;s (t-1) expectation on current-quarter news, finding no effect (Appendix Table B5). (5) Checks that firm characteristics predict neither response timing (Table 4) nor expectation levels (Table 3), supporting the random-timing assumption. (6) Local projections on output and employment (Table 8) establishing that firms&amp;rsquo; real-side behavior responds to the shocks, motivating belief responses. Standard errors are White and clustered at the firm level throughout.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It builds on the firm-expectations literature (Coibion, Gorodnichenko, Kumar 2018; Candia, Coibion, Gorodnichenko 2023) showing firms&amp;rsquo; expectations lie between professional forecasters&amp;rsquo; and households&amp;rsquo;—confirmed here by intermediate disagreement among firms. It connects to expectation-formation work (D&amp;rsquo;Acunto et al. 2021 on shopping experience; Coibion-Gorodnichenko 2015 on exchange-rate sensitivity in Ukraine; Kumar et al. 2015 on New Zealand managers) and to studies of news effects on expectations (Beechey, Johannsen, Levin 2011). It is closest in spirit to Lamla and Vinogradov (2019), who compare household expectations before/after monetary announcements; the contribution is to study firms in an economy with a recent history of high inflation and dollarization undergoing disinflation. It also relates to regime-change classics (Sargent 1982 on ending hyperinflations; Mankiw, Reis, Wolfers 2003 on Volcker disinflation), filling the gap that little is known about firms&amp;rsquo; expectations across a policy-regime change. Its Israeli monetary-surprise null in the stable period echoes Coibion et al. (2020) and Ilek (2021).&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Central implication: even after successful de-dollarization and a prolonged low-and-stable inflation environment, Israeli firms continued to monitor and react to inflation news—so de-dollarization (firms&amp;rsquo; renewed trust in local currency) does not necessarily translate into inattention, and (re)anchoring expectations in the sense of making them insensitive to news may take a long time. For countries currently experiencing high inflation, the Israeli experience suggests firm expectations can remain news-sensitive for an extended period. Scope conditions: the firm sample is not nationally representative; results are specific to Israel&amp;rsquo;s institutional setting (monthly CPI on the 15th, scheduled rate decisions); the monetary-policy result is fragile—it is driven mainly by the unusually large 2001–2002 shocks and disappears in calmer periods, so the conclusion that monetary surprises move firm expectations holds chiefly when shocks are large.&lt;/p&gt;
&lt;h3 id="q7-are-there-other-significant-findings-or-caveats"&gt;Q7. Are there other significant findings or caveats?&lt;/h3&gt;
&lt;p&gt;Descriptive facts: firms&amp;rsquo; average annual inflation expectations (2001Q3–2018Q3) averaged 2.34% (vs. 1.81% for professional forecasters, 1.57% for the capital market); in the 2011Q1–2018Q3 panel households averaged 3.02% while firms averaged 1.83%, banks 1.07%. Firms&amp;rsquo; expectations are about one percentage point below households&amp;rsquo; but 0.5–1 pp above other (forecaster/market) sources, and disagreement among firms lies between that of households and professional forecasters—consistent with prior literature. Expectations co-move strongly across sources and across industries. Raw cross-period descriptive evidence (Table 5) shows average and median expectations decline as more information becomes available (Period 1 mean 2.52 → Period 3 mean 2.26), and disagreement weakly declines. The largest interest-rate surprises (1.5–2 pp) occurred at the sample start: in December 2001 the Bank cut the rate by 2 pp to 3.8%, triggering capital outflow, depreciation, and price increases, then reversed to 9.1%. A caveat is that the survey was discontinued at end-2020 (replaced by a CBS survey), and the unbalanced, voluntary panel limits representativeness.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Long-Distance Trade and Long-Term Persistence</title><link>https://macropaperwarehouse.com/papers/long-distance-trade-and-long-term-persistence/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/long-distance-trade-and-long-term-persistence/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether the location of economic activity adapts to changes in the location of trading opportunities, or whether historical patterns of trade permanently fix where cities emerge and grow. The question is fundamental to economic geography: many large cities owe their origins to access to long-distance trade that has since moved on, yet the cities persist. The empirical context is the staggered liberalization of direct transatlantic trade across the Spanish Empire in the second half of the 18th century. Before the reform, a mercantilist system confined legal trade to four American ports (Cartagena de Indias, Callao, Portobello/Nombre de Dios, and Veracruz) and a single European port (Seville, then Cadiz). Following Spain&amp;rsquo;s defeat in the Seven Years&amp;rsquo; War, a sequence of decrees opened direct trade to an additional 40-plus ports between 1765 and the early 19th century. The reform was driven by European interstate competition and implemented from above, creating staggered, quasi-exogenous variation in transportation times to Europe across American cities.&lt;/p&gt;
&lt;p&gt;The empirical strategy is a difference-in-differences design. The author constructs a novel panel of 62 cities in Spanish America observed every 50 years from 1600 to 1850 (372 observations), plus a settlement-level panel of 53,581 grid-cell-decade observations for 1710-1810. The key treatment variable is the time-varying transportation time to Europe, computed via a directed network using maritime logbooks (282,322 daily entries from the CLIWOC 2.1 database, 1750-1855) to estimate wind-conditional sailing speeds, and land travel models based on slope, elevation, landcover, and postal routes. The reduction in transportation time ranged from 0 to 38.3 days across locations, with an average pre-reform time of 93.5 days and an average reduction of 7.7 days - economically significant, representing 0 to 40 percent of the baseline average.&lt;/p&gt;
&lt;p&gt;The paper documents four main empirical patterns, all within a city-and-time fixed-effects framework that absorbs time-invariant location fundamentals. First, the reform improved market integration: non-bullion Spanish imports from the Americas rose nearly fourfold after the 1778 decree (following no secular trend before 1765), and a commodity price ratio between Spain and Spanish America converged beginning in the second half of the 18th century, consistent with lower transportation costs facilitating arbitrage. Second, lower transportation times raised urban population. In the preferred specification, a one-day reduction in transportation time to Europe increases city population by approximately 2 percent over a 50-year period (baseline coefficient -0.023, significant at conventional levels, with the sign and approximate magnitude stable across specifications adding viceroyalty-by-year or country-by-year fixed effects and controls interacted with year indicators). Third, the effects are concentrated among smaller cities and in the fringe regions of the empire (Argentina, Chile, Venezuela, the Caribbean, etc.): for the fringe region the point estimate is -0.016, while for the colonial core (Mexico, Peru, Bolivia) the effect is statistically indistinguishable from zero. A ten-day reduction in transportation time raises the probability that a grid cell contains a settlement by approximately one percentage point (against a sample mean of 11 percent), suggesting the primary margin was growth of existing cities rather than expansion to new frontier areas. Fourth, the cross-sectional elasticity of contemporary (year 2000) population density to pre-reform (1750) population size is 0.592 overall, but falls to 0.369 for cities that experienced large reductions in transportation times, and rises to 0.866 for cities that experienced little change - consistent with the reform attenuating the persistence of pre-reform settlement patterns specifically where the trade shock was large.&lt;/p&gt;
&lt;p&gt;To interpret mechanisms and simulate long-term implications, the author calibrates a dynamic spatial general equilibrium model built on Allen and Donaldson (2022). The model features cities that differ in productivity, land endowments, and trade/migration costs, with agents living two periods, static and dynamic agglomeration economies (parameters a1 = 0.055 and a2 = 0.063 from the data), and Frechet-distributed migration preferences. Counterfactual exercises simulate the model forward 300 years. In the benchmark counterfactual, the average reduction in transportation costs increases urban population by 1.27 percent (25th/75th percentile: -0.06 to 1.34 percent), with a maximum city-level gain of 11.77 percent and a minimum of -0.2 percent. Effects in the fringe region average 1.9 percent population gain versus 0.26 percent in the core. Decomposition exercises show that: differences in location fundamentals (productivity and land endowments, A and H) account for part of the core-fringe differential (the gap falls from 1.64 to 1.11 percentage points when fundamentals are equalized); equalizing the pre-reform population size across cities leaves the gap nearly unchanged (1.64 to 1.65), suggesting dynamic agglomeration from historical size plays little role in driving the core-fringe difference; by contrast, equalizing the spatial incidence of the shock (the amount by which transportation times fell) closes the differential almost entirely (gap falls to 0.16 percentage points), indicating that the fringe was simply more restricted before the reform and thus received a larger shock. Trans-Atlantic migration is also an important channel: when trans-Atlantic migration is made prohibitively costly in the model, the average population effect falls to roughly 14.86 percent of the benchmark, indicating that migration from Europe amplified the effect of lower trade costs on city populations.&lt;/p&gt;
&lt;p&gt;The overarching conclusion is that economic geography is not fully path-dependent: where trading opportunities move, economic activity can follow - but this adaptation is conditional. Cities that had already accumulated large populations before the change in trading locations are insulated from reallocation, because their internal market size reduces their reliance on long-distance external trade. In less-developed fringes with smaller internal markets, however, the spatial distribution of activity is more malleable and adjusts substantially to the change in trading opportunities.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The identification strategy is a two-way fixed-effects difference-in-differences, exploiting cross-city variation in the change in transportation time to Europe induced by the staggered port-opening reform. City fixed effects absorb all time-invariant unobserved location fundamentals (agroclimatic characteristics, disease environment, natural harbors, etc.). Year fixed effects absorb common time-varying shocks. The key identifying assumption is parallel trends: absent the reform, population growth would have evolved similarly across cities with different exposure to the transportation-time reduction. Three main threats are addressed. First, selective port targeting: if policymakers chose to open ports in anticipation of their commercial potential, the reform would not be exogenous to growth trajectories. The author argues against this: historical accounts indicate reluctance to open the wealthiest colony (New Spain/Mexico) precisely because its prosperity might divert trade from other regions, and the reform was driven by European interstate competition (the Seven Years&amp;rsquo; War) rather than by American economic conditions. Second, confounding from contemporaneous administrative reforms: Bourbon-era reorganizations, new viceroyalties (Rio de la Plata and Nueva Granada), and ecclesiastical changes could coincide with the trade reform. The author addresses this by dropping cities in the two new viceroyalties (coefficients remain similar) and by including viceroyalty-by-year fixed effects. Third, the transportation network itself might endogenously reflect urban growth (roads built to connect growing cities). The author notes the transportation times are constructed from predetermined geographic characteristics (wind patterns, slope, elevation, landcover) and pre-existing postal routes, not from contemporaneous road-building. The dynamic pre-trends test (interacting the reform-induced change in transportation time with year indicators) shows no significant difference in population growth across differentially exposed cities before 1750, supporting the parallel trends assumption. An alternative synthetic control design yields qualitatively similar results (treatment effect of approximately 19 percent over one century). The author also estimates the model on the sub-sample of cities far from ports (distance above median) and finds similar coefficients, addressing concerns that the reform directly targeted port cities for reasons correlated with their growth.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically-and-in-the-model"&gt;Q2. What are the main mechanisms and how are they distinguished empirically and in the model?&lt;/h3&gt;
&lt;p&gt;Two principal mechanisms are proposed. The first is a trade-cost channel: lower transportation times reduce the iceberg cost of exporting to European markets, lowering the price index for traded goods in affected cities and raising real income, which attracts labor migration. This channel operates even if migration costs are unchanged. The second is a migration-facilitation channel: lower transportation times reduce information frictions and direct travel costs for migrants from Europe, lowering migration frictions as well as trade costs. The quantitative model distinguishes these by running counterfactuals with and without changes in migration frictions (keeping migration costs fixed at 1760 levels). In the benchmark, allowing migration frictions to fall alongside trade costs yields an average 1.27 percent population increase; fixing migration frictions yields 0.66 percent. This comparison indicates that lower migration frictions approximately double the population effect relative to the pure trade-cost channel. The model also distinguishes between trans-Atlantic migration (Spain to Americas) and intracolonial migration. When trans-Atlantic migration is shut off (migration costs set prohibitively high for Europe-America pairs), the average population effect falls to approximately 15 percent of the benchmark value, implying that trans-Atlantic migration is the dominant driver of the population response. A third dimension of heterogeneity concerns internal market size: in the partial-equilibrium analytics, the marginal impact of a reduction in the trade cost to Europe is attenuated in larger cities because a larger local market reduces the share of consumption sourced externally, making the price index less sensitive to external trade costs. This mechanism is consistent with the finding that the reform had statistically significant effects only in smaller cities and fringe regions.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented-and-what-explains-it"&gt;Q3. What heterogeneity is documented and what explains it?&lt;/h3&gt;
&lt;p&gt;The paper documents three main dimensions of heterogeneity. First, the colonial core (Mexico, Peru, Bolivia) versus the fringe (Argentina, Chile, Venezuela, Caribbean, Central America): the average effect in the fringe region is -0.016 per day of transportation time in the baseline city regressions (statistically significant), while the core coefficient is indistinguishable from zero. In the counterfactual model, the fringe shows a 1.9 percent average population gain against 0.26 percent in the core. Second, large versus small cities: the effects are larger and more precisely estimated for cities with below-median pre-reform population. Third, within the fringe, there is wide dispersion: the 25th/75th percentile population change in the model is -0.06 to 2.47 percent, with individual city gains up to 11.77 percent (most sizable in Buenos Aires and Caribbean ports) and losses up to -0.2 percent (most negative in cities whose relative economic centrality declined, such as Veracruz and Cartagena). The decomposition of the core-fringe differential shows: (a) location fundamentals (A and H, i.e. productivity and land endowments) explain part of the differential - equalizing fundamentals reduces the gap from 1.64 to 1.11 percentage points; (b) initial population size contributes little - equalizing pre-reform population shares barely moves the gap (from 1.64 to 1.65); (c) the spatial incidence of the shock explains most of the differential - equalizing the size of the transportation-cost reduction across all cities virtually eliminates the gap (to 0.16 percentage points), because the fringe was more trade-restricted before the reform and thus received a larger absolute reduction in transportation times.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper runs extensive robustness exercises. On the reduced-form side: (1) Dynamic event-study specifications show no significant pre-trends before 1750 for the full sample, sub-samples by city size and macro-region, and for settlements. (2) Synthetic control method: treating cities as a group, the synthetic control closely matches pre-reform population trends, with a divergence beginning in 1800 and an implied treatment effect of approximately 19 percent (0.338 log points) over a century, and the true treatment group has the highest post/pre-RMSE ratio relative to all placebo assignments. (3) Dropping outliers (cities outside the 5th-95th percentile of pre-reform population growth rates): coefficients remain around -0.018. (4) Weighting by population size: coefficients similar at around -0.024. (5) Spatial standard errors following Conley (1999): results hold. (6) Robustness value analysis following Cinelli and Hazlett (2020): a confounder would need to explain 16.9 percent of both outcome and treatment variation to fully account for the effect, and 7 percent to render it statistically insignificant - both larger than the combined R2 of observable fundamentals. (7) Interior cities only (distance to port above median): coefficient around -0.016, similar to baseline. (8) Dropping the viceroyalties of Nueva Granada and Rio de la Plata (formed in the 18th century): similar coefficients. (9) Estimating only through 1800 to exclude independence-era effects: point estimates similar for smaller cities. (10) Alternative transportation cost measures including a simple distance measure, showing qualitative robustness. On the model and counterfactual side: (1) Alternative values of the elasticity of substitution (sigma 3-7), Frechet shape parameter (theta 2-4), land expenditure share (1-mu: 0.4-0.6), and agglomeration parameters (a1 in [0.04, 0.07], a2 in [0.02, 0.07]) all yield qualitatively similar results. (2) Alternative trade-cost elasticity from Baum-Snow et al. (2018): similar. (3) Incorporation of national borders after independence (15 percent additional trade cost for cross-border flows): similar, somewhat larger effects. (4) Secular productivity improvements and secular declines in transportation costs (0.88 percent per year starting 1800 per Harley 1988): average effects similar or larger than baseline.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The paper connects to four main strands. First, the literature on history dependence in economic geography. Davis and Weinstein (2002) use WWII bombing shocks to show that Japanese cities return to their pre-shock size, highlighting persistence from locational advantages. Bleakley and Lin (2012) find that US portage sites retain elevated population density long after canals made them obsolete, a classic multiple-equilibria story. Redding, Sturm and Wolf (2010) exploit German division and reunification to show airports exhibit path dependence. Michaels and Rauch (2018) compare Roman and non-Roman cities in France, finding Roman legacy persists. This paper contributes by using a large-scale historical policy reform that changed the location of trading opportunities itself - controlling for time-invariant location fundamentals by construction - and showing that adaptation occurs but is contingent on initial urbanization levels. Henderson et al. (2018) use cross-country data to show locational advantages governing trade matter less in early developers (countries that developed under high transportation costs). This paper supports that cross-sectional finding and gives it a causal interpretation within a single institutional setting. Second, the literature on transportation costs and income. Frankel and Romer (1999) and Feyrer (2019) find large reduced-form effects. Pascali (2017) uses steamship diffusion and finds little aggregate effect except in countries with inclusive institutions - this paper focuses within countries (single institutional environment) and finds robust effects on the spatial distribution rather than aggregate national income. Third, the historical institutions literature. Acemoglu, Johnson and Robinson (2002) establish that pre-industrial population density negatively predicts current income (reversal of fortune). This paper reframes that as partly attributable to trade institutions, showing that Bourbon-era reforms interacted with pre-existing geography to shape the reversal. Fourth, the literature on 18th-century Spanish empire reforms. Valencia (2019), Alvarez-Villa and Guardado (2020), Arteaga (2022), and Chiovelli et al. (2024) examine Bourbon administrative and ecclesiastical reforms. This paper is distinct in focusing on commercial policy and in constructing time-varying bilateral transportation time matrices rather than relying on cross-sectional variation.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central policy implication is that trade liberalization that reduces access costs to long-distance markets can reshape the spatial distribution of economic activity within a country, particularly benefiting peripheral regions that were previously excluded from international trade networks. However, this finding comes with important scope conditions. First, the magnitude of the effect is larger in locations with small pre-existing internal markets. Regions with larger pre-existing urban agglomerations are relatively insulated from reallocation because their size makes them less dependent on external trading opportunities. Policy interventions that reduce international trade costs may therefore have limited spatial rebalancing effects in already-urbanized contexts. Second, the adaptation is not a rapid reallocation: the estimated 2 percent population gain per day-reduction in transportation time reflects cumulative adjustment over 50-year periods. Third, the historical context involves extractive colonial institutions. The paper notes that lower transportation costs influenced spatial development within countries even under extractive institutions, suggesting the result does not require inclusive institutions - but the magnitude and form of adjustment may differ under different institutional regimes. Fourth, migration plays an important amplifying role: in the model, restricting trans-Atlantic migration halves to nearly eliminates the population effect. In modern contexts where immigration is restricted, the spatial reallocation effect of trade liberalization may be substantially smaller. Fifth, the author cautions that the reform involved abrupt, large changes in trade costs, which may produce different adjustment dynamics than gradual reductions.&lt;/p&gt;
&lt;h3 id="q7-how-is-the-transportation-network-constructed-and-validated"&gt;Q7. How is the transportation network constructed and validated?&lt;/h3&gt;
&lt;p&gt;Maritime transportation times are estimated by regressing daily sailing speed (in knots) from 188,687 logbook entries (after removing implausibly fast observations above 10 knots, anchored ships, steamships, and coastal entries) on wind speed and the cosine of the angle between direction of travel and wind direction. The model is estimated on a training sample (179,255 entries) and validated on a holdout sample (9,432 entries), yielding a mean squared error of 2.16. Fitted sailing speeds are then extrapolated to a 0.16 x 0.16 degree global grid using modern wind data from NOAA&amp;rsquo;s Global Forecasting System (2011-2017), assuming wind patterns are sufficiently stable (the correlation between historical logbook wind speed and modern wind speed is 0.24; for wind direction, 0.33). The Dijkstra algorithm finds time-minimizing routes through this grid. Land transportation is modeled using a Tobler-style hiking function adjusted for slope, elevation, and landcover, based on the Weiss et al. (2018) parameterization, applied to a 0.16-degree land grid with postal route locations from Stangl (2019b) treated as roads. Validation compares maritime times to seadistances.org sailing times across 21 ports (strong positive correlation), and land times to the Human Mobility Index and Google Maps driving times (again strongly correlated). The transportation time to Europe from city i in period t is defined as the minimum over the set of ports open to direct trade at time t of the sum of the inland travel time to the nearest open port plus the maritime travel time from that port to Cadiz.&lt;/p&gt;
&lt;h3 id="q8-how-are-the-spatial-models-parameters-identified-and-what-are-the-key-parameter-values"&gt;Q8. How are the spatial model&amp;rsquo;s parameters identified and what are the key parameter values?&lt;/h3&gt;
&lt;p&gt;The model has six parameters (sigma = elasticity of substitution, theta = Frechet shape parameter for migration, mu = expenditure share on traded goods, b = preference shifter for transatlantic goods, a1 = static agglomeration externality, a2 = dynamic/historical agglomeration externality), two vectors of location fundamentals (A and H), and time-varying trade and migration cost matrices (T and M). Sigma is set to 5 following Simonovska and Waugh (2014). Theta is set to 3.18 following Bryan and Morten (2019). Mu = 0.5 is the midrange estimate of the land income share for colonial Mexico and Peru from Arroyo Abad and van Zanden (2016). a1 = 0.055 is taken from the mid-range of estimates in Combes and Gobillon (2015). The trade cost elasticity with respect to transportation time (kappa) is estimated from a port-level gravity model of Spanish imports from Spanish America (1797-1820) using PPML with viceroyalty fixed effects, yielding a transportation time elasticity of trade flows of -2.23, which gives kappa = 0.56. The preference shifter b = 0.45 is chosen to match the observed Spanish import share from the Americas in 1750 (approximately 25 percent per Prados de la Escosura and Casares 1983). The migration cost elasticity lambda is estimated similarly from migration gravity, yielding -lambda*theta = -1.16, so lambda = 0.363. The dynamic agglomeration parameter a2 = 0.063 is identified by estimating the structural version of the reduced-form city-size equation (regressing log population on log price index, log real income, lagged log population, and location controls), where the coefficient on lagged population identifies a2 via the model&amp;rsquo;s equilibrium conditions. Location fundamentals A and H are recovered by inverting the model to exactly match the observed population distribution and nominal wages in 1750.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-paper-contribute-to-understanding-of-the-reversal-of-fortune-in-the-americas"&gt;Q9. What does the paper contribute to understanding of the &amp;lsquo;reversal of fortune&amp;rsquo; in the Americas?&lt;/h3&gt;
&lt;p&gt;Acemoglu, Johnson and Robinson (2002) established that areas with higher pre-industrial (circa 1500) population density tend to have lower income today, interpreting this as evidence that Spanish colonization was most extractive in densely populated areas (which later fell behind) and that sparser-populated frontier areas had better institutions (property rights) that supported later development. This paper complements that institutional story by showing that trade institutions also matter for explaining the reversal. The Bourbon reform - driven by dynastic change from Habsburg to Bourbon rule and by European interstate competition - specifically opened direct trade access to peripheral areas that had been systematically excluded under the Habsburg mercantilist system. The paper&amp;rsquo;s persistence results (lower elasticity of contemporary to pre-colonial population density in areas more exposed to the reform) suggest that the trade reform contributed to the subsequent relative rise of peripheral regions. The finding thus supports the view that the reversal of fortune is partly rooted in institutional change (trade liberalization) interacting with pre-existing geography, rather than in population-density-determined institutions alone. The scope condition is important: the core-versus-fringe heterogeneity shows the reform&amp;rsquo;s spatial effects were largest precisely in the sparsely populated periphery - consistent with the Acemoglu et al. mechanism but augmenting it with a trade-access channel.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-limitations-of-the-analysis-acknowledged-by-the-author"&gt;Q10. What are the limitations of the analysis acknowledged by the author?&lt;/h3&gt;
&lt;p&gt;The author acknowledges four main limitations. First, the reform involved sizeable and abrupt changes in trade costs. More gradual liberalizations might produce different adjustment dynamics, potentially slower convergence or different spatial sorting. Second, the absence of individual-level migration data prevents a more direct examination of whether the city-population effects operate primarily through trans-Atlantic immigration, intracolonial migration, or natural population growth. The model-based inference that trans-Atlantic migration matters substantially is indirect. Third, path dependence likely plays a more important role in industrialized contexts with stronger agglomeration economies (larger a2 than estimated here). The pre-industrial colonial setting, with relatively modest agglomeration forces and thin labor markets, may not generalize to modern industrialized spatial economies. Fourth, the study&amp;rsquo;s focus on within-country (within-empire) variation means it cannot directly address the effect of trade liberalization on aggregate national income, only on the spatial distribution of activity within the empire.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-significance-of-the-finding-that-the-reform-primarily-affected-city-size-rather-than-frontier-settlement"&gt;Q11. What is the significance of the finding that the reform primarily affected city size rather than frontier settlement?&lt;/h3&gt;
&lt;p&gt;The settlement-level analysis uses a balanced panel of 53,581 grid-cell-decade observations for 1710-1810, with an indicator for whether a cell contains any settlement. The baseline result is that a ten-day increase in transportation time to Europe reduces the probability of a cell containing a settlement by one percentage point, against a sample mean of 11 percent, and this effect is small relative to the urban population effects. Event-study plots for settlement formation show no significant pre-trends and only modest post-reform effects. This implies that the reform&amp;rsquo;s primary spatial impact was to concentrate more people in existing urban centers rather than to push economic activity into entirely new locations. This is consistent with the model, in which cities have pre-existing productivity advantages (embedded in A and H) that make them focal points for agglomeration. It also implies the reform did not create entirely new urban systems in frontier areas but rather amplified existing ones, which is important for interpreting the persistence results: even in the fringe, the settlements that grew were already established before 1765.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Transportation time to Europe&lt;/strong&gt;: The time-minimizing route from a given location in Spanish America to Cadiz (the dominant European trading port), computed by combining maritime sailing speed estimates (from logbooks, conditional on wind speed and direction) and land travel speed estimates (based on slope, elevation, landcover, and road location) via the Dijkstra algorithm; time-varying because the set of ports permitted to trade directly with Europe changes as the reform proceeds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Comercio Libre (free trade reform)&lt;/strong&gt;: The staggered series of Spanish royal decrees between 1765 and the early 19th century that progressively lifted the mercantilist restriction confining direct transatlantic trade to four American ports and a single Spanish port, ultimately opening more than 45 American ports to direct trade with Europe; motivated by European interstate competition rather than by the commercial potential of specific American locations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic agglomeration externality (a2)&lt;/strong&gt;: In the Allen-Donaldson (2022) framework as applied here, the component of city-level total factor productivity that depends on the city&amp;rsquo;s own population in the previous period rather than the current period; it encodes the idea that historically larger cities are persistently more productive through channels such as durable local infrastructure, accumulated local knowledge, or input-sharing networks. Estimated at a2 = 0.063 in this setting, smaller than values found in Allen and Donaldson (2022).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;First-nature fundamentals&lt;/strong&gt;: Time-invariant geographic endowments that determine a location&amp;rsquo;s intrinsic productivity and land availability independent of the scale of economic activity, captured in the model by the vectors A (productivity) and H (arable land); these are recovered by inverting the spatial model to match observed 1750 population and wages and are correlated with caloric potential, elevation, terrain ruggedness, and proximity to rivers and coasts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second-nature fundamentals&lt;/strong&gt;: The agglomeration forces that arise from the scale of economic activity already present at a location, including static (current population) and dynamic (lagged population) agglomeration economies; in this paper, the term is used to explain why larger pre-reform cities in the core are insulated from the trade reform&amp;rsquo;s spatial reallocation effects - their scale generates internal-market advantages that reduce reliance on long-distance external trade.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Internal market size (market insulation)&lt;/strong&gt;: The degree to which a city&amp;rsquo;s price index for traded varieties is determined by local production rather than external trade costs; in the model, cities with larger local productivity (higher Ait) have a less sensitive price index to changes in the trade cost with Europe because local goods compete with imported varieties, dampening the welfare and migration effects of trade liberalization; this is the central mechanism explaining the core-fringe heterogeneity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Persistence elasticity&lt;/strong&gt;: The coefficient relating contemporary (year 2000) population density or size to pre-reform (1500 or 1750) population density or size in a cross-sectional regression, interpreted as a measure of how much historical settlement patterns predict current ones; found to be 0.866 for cities with below-median changes in transportation time (little treated) and 0.369 for cities with above-median changes (strongly treated), documenting that the reform attenuated the persistence of pre-reform settlement patterns.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Migration-facilitation channel&lt;/strong&gt;: The mechanism by which lower transportation times to Europe reduce not only trade costs but also migration frictions - through lowering the direct cost of travel and through improving information flows about opportunities in American cities - thereby amplifying city population growth beyond the pure trade-cost effect; quantified in the model by comparing counterfactuals that allow migration frictions to decline with those that hold them fixed at 1760 levels.&lt;/p&gt;</description></item><item><title>Manipulation of information in times of crisis: evidence from Covid excess mortality</title><link>https://macropaperwarehouse.com/papers/manipulation-of-information-in-times-of-crisis-evidence-from-covid-excess-mortality/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/manipulation-of-information-in-times-of-crisis-evidence-from-covid-excess-mortality/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Karlinsky and Shayo ask which governments manipulate public information, in which direction, and by how much — questions that are normally intractable because the ground truth is unobservable. The Covid-19 pandemic supplies an unusual opportunity: all countries faced a broadly similar crisis simultaneously, and all-cause mortality — collected by national statistical offices as a routine bureaucratic function independently of Covid — provides a manipulation-resistant benchmark against which officially-reported Covid deaths can be evaluated.&lt;/p&gt;
&lt;p&gt;The authors hand-collect all-cause mortality data for 134 countries and territories from national statistical offices, population registries, health ministries, and, in some cases, right-to-information requests facilitated by local journalists. Data span 2015–2021 at weekly, monthly, or annual frequency. Their sample covers 93 percent of countries with at least 75 percent Death Registration Completeness. They compute, for each country, a Misreporting Rate (MRR) defined as estimated Covid deaths minus officially reported Covid deaths, normalised by expected total deaths derived from pre-pandemic trends. Estimated Covid deaths equal excess mortality — itself estimated from a country-specific model with weekly/monthly fixed effects and an annual trend (R² = 0.997 in pre-pandemic prediction) — minus adjustments for excess deaths attributable to conflicts, natural disasters, and other identifiable non-Covid causes. Those adjustments are small: the mean total adjustment across the sample is 0.04 percent of expected deaths.&lt;/p&gt;
&lt;p&gt;Six main findings emerge. First, between 45 and 55 percent of the 134 countries misreported Covid deaths. Second, the direction of manipulation is overwhelmingly one-sided: of 131 countries with sufficient data to estimate confidence intervals, 59 reported accurately, 62 significantly underreported, and only 10 overreported. The theoretical prediction that governments might exaggerate a crisis — to rally populations, legitimise repressive measures, or attract foreign aid — finds no empirical support. Third, the magnitude of underreporting is large: the sample reported 5.08 million Covid deaths in 2020–2021 while estimated actual Covid deaths were 12.47 million, nearly 2.5 times the official figure; the implied global MRR is 12.8 percent. Among the 62 underreporting countries, the average MRR is 14.5 percent of expected total deaths and the median is 12 percent. Individual-country MRRs range from above 37 percent (Bolivia, Nicaragua) downward, with Russia at 24 percent. Fourth, state capacity in counting and registering deaths explains some but far from most cross-country variation; the R² of the best capacity-only regression is 0.115. Chile and Russia have virtually identical Death Registration Completeness and Percent Well-Certified Death Registrations, yet Chile accurately reported while Russia&amp;rsquo;s MRR is 24 percent. Fifth, the extent of underreporting is strongly associated with constraints on governmental power. In individual regressions conditioning on capacity, each of three institutional constraint measures — Clean Elections, Executive Constraints, and Freedom of the Press — is associated with a 0.4–0.5 standard deviation lower MRR per one standard deviation stronger constraint. In a joint model including all 12 factors from four domains (macroeconomic incentives, culture, audience sophistication, institutions), institutional constraints are the strongest predictor (partial R² ≈ 0.11), followed by audience sophistication (partial R² ≈ 0.04–0.06). Macroeconomic incentives — tourism reliance, unemployment, foreign direct investment — are not jointly significant. Cultural factors (trust, individualism, religiosity) lose significance once other factors are controlled. The full model explains more than 50 percent of MRR variation. Sixth, countries with a communist legacy (defined as having had a communist or socialist regime for at least 10 years, covering 34 countries) show significantly higher misreporting even holding current institutional and cultural conditions constant. Countries that held elections during 2020–2021 also show significantly higher misreporting.&lt;/p&gt;
&lt;p&gt;The results are robust to alternative expected-mortality models, alternative MRR normalisations, the inclusion of Bangladesh, China, and Indonesia (treated separately due to data quality concerns), year-by-year (2020 vs. 2021) splits, controls for age structure and GDP per capita, and alternative manipulation measures (underdispersion, Benford&amp;rsquo;s law deviations). The evidence that manipulation cannot be attributed to varying standards for false-positive attribution of cause of death is direct: four pre-pandemic measures of a country&amp;rsquo;s tendency to use unspecified cause-of-death categories are uncorrelated with MRR and individually account for less than 1 percent of its variation.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s contribution to the economics of information manipulation is methodological as well as empirical: it provides a comparable, country-level measure of governmental misinformation based on actual observable actions regarding a policy issue of central importance, covering a large and diverse cross-section of countries.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The strategy compares officially-reported Covid deaths (the variable that attracted political attention and over which governments had strong incentives and ability to intervene) with estimated Covid deaths derived from excess all-cause mortality (a statistic collected routinely by national bureaucracies under very different incentive structures, harder to manipulate, and less visible publicly during the pandemic). The identifying assumption is that all-cause mortality data are not themselves systematically manipulated in response to Covid. The authors defend this on four grounds: (1) all-cause mortality has long been collected independently of Covid; (2) ascertaining that someone died is far easier than attributing a cause of death; (3) Covid figures attracted vastly more public attention, making their manipulation more urgent; (4) when governments appear to have discovered the evidential value of excess mortality, their response has been to delay publication of all-cause data rather than to alter it (Belarus is cited as an example). The main remaining threat is that the adjustment for non-Covid excess deaths (conflicts, disasters, traffic accidents, suicides, homicides) is imperfect in countries with poor data on those causes. The authors note this caveat but show mean adjustments are tiny (0.04% of expected deaths) and the largest individual adjustments (Armenia 6.1%, Azerbaijan 3.2%) are driven by the Nagorno-Karabakh war and are handled explicitly.&lt;/p&gt;
&lt;h3 id="q2-how-is-excess-mortality-estimated-and-how-sensitive-are-the-results-to-modelling-choices"&gt;Q2. How is excess mortality estimated, and how sensitive are the results to modelling choices?&lt;/h3&gt;
&lt;p&gt;Country-specific models are estimated using 2015–2019 all-cause mortality data, including country-specific weekly or monthly fixed effects and a country-specific annual trend to capture seasonality and long-run factors (population ageing, improvements in health care, etc.). The model achieves R² = 0.997 in predicting pre-pandemic mortality. The authors report in Supplementary Material B that alternative expected-mortality approaches from the literature yield very similar results, as do alternative normalisations of the MRR. Sensitivity to model choice is low because the discrepancies between excess and reported deaths in weak-institution countries are so large that they persist across methodological variants.&lt;/p&gt;
&lt;h3 id="q3-how-do-the-authors-distinguish-intentional-manipulation-from-limited-state-capacity"&gt;Q3. How do the authors distinguish intentional manipulation from limited state capacity?&lt;/h3&gt;
&lt;p&gt;They use two pre-pandemic, capacity-specific measures: (1) Death Registration Completeness (DRC) — the share of deaths captured by the vital registration system — and (2) Percent of Well-Certified Death Registrations (PWC) — the share with proper cause-of-death attribution. Both are computed before the pandemic so they are not contaminated by Covid-era behaviour. Regressions confirm that capacity predicts MRR negatively (R² up to 0.115), but the residual variation remains large. The clearest illustration is Chile vs. Russia: both have complete DRC and near-identical high PWC, yet Chile reports accurately and Russia has an MRR of 24 percent. All subsequent analysis of correlates conditions on these capacity measures.&lt;/p&gt;
&lt;h3 id="q4-how-do-the-authors-rule-out-the-possibility-that-differences-in-false-positive-aversion-rather-than-manipulation-explain-mrr-variation"&gt;Q4. How do the authors rule out the possibility that differences in false-positive aversion (rather than manipulation) explain MRR variation?&lt;/h3&gt;
&lt;p&gt;They construct four pre-pandemic measures from WHO Mortality Database ICD-10 cause-of-death data: (1) number of ICD codes reported; (2) share of specific-viral deaths among all viral deaths; (3) share of specific-infection deaths among all infection deaths; (4) share of specific-respiratory deaths among all respiratory deaths. A country more averse to false positives would report less specific causes. None of the four measures is significantly associated with MRR, and none accounts for more than 1 percent of its variation. This rules out differences in diagnostic/reporting standards as a driver of the observed discrepancies.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-direction-of-manipulation-and-what-does-this-imply-for-theories-of-governmental-information-behaviour"&gt;Q5. What is the direction of manipulation and what does this imply for theories of governmental information behaviour?&lt;/h3&gt;
&lt;p&gt;Of 131 countries with estimable confidence intervals, 62 significantly underreported and only 10 overreported. The four main theoretical channels for overreporting — rally-around-the-flag effects, legitimising repression, attracting foreign aid, and inducing flight-to-safety compliance — find no empirical support. The authors argue that the rally-around-the-flag mechanism requires an outgroup-related threat (Covid, unlike a foreign military attack, was not easily framed this way), that Covid mortality does not signal repressive capacity, and that international economic actors appear sufficiently sophisticated to be sceptical of inflated figures. The pattern is consistent instead with governments downplaying to project competence, reduce accountability, and justify inadequate responses.&lt;/p&gt;
&lt;h3 id="q6-what-factors-are-most-strongly-associated-with-misreporting-and-how-are-they-ranked"&gt;Q6. What factors are most strongly associated with misreporting, and how are they ranked?&lt;/h3&gt;
&lt;p&gt;In joint regressions with all 12 factors from four domains, after conditioning on capacity: (1) Institutional constraints (Clean Elections, Executive Constraints, Freedom of the Press) have the highest partial R² (approximately 0.11 for Executive Constraints alone) and are jointly significant at p &amp;lt; 0.001; each standard deviation of stronger institutional constraint is associated with roughly 0.4–0.5 standard deviations lower MRR. (2) Audience Sophistication (tertiary education, HDI Education Index, internet access) is the second strongest domain (partial R² in the range of 0.04–0.06 per variable; jointly significant at p &amp;lt; 0.05). (3) Cultural factors (trust, individualism, religiosity) are individually significant in bivariate regressions but lose significance when institutional and other factors are controlled. (4) Macroeconomic incentives (tourism, unemployment, net FDI) are not jointly significant in any specification. Specification-curve analysis across all combinations of controls confirms that Executive Constraints is the single most robust predictor, retaining sign, magnitude, and significance across all models. The full model (Table 4, column 1) has R² exceeding 0.50.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-communist-legacy-finding-and-how-is-it-interpreted"&gt;Q7. What is the communist legacy finding and how is it interpreted?&lt;/h3&gt;
&lt;p&gt;Countries defined as having had a communist or socialist regime for at least 10 years (34 countries) show significantly higher MRRs even after conditioning on contemporary institutional constraints, audience sophistication, culture, and capacity. The coefficient is statistically significant at p &amp;lt; 0.05 or better in the main and most robustness specifications. The authors point to Harrison (2017) on the pervasiveness of information manipulation in communist states as a historical precedent, and interpret the finding as a persistent legacy operating through channels not fully captured by current measures. This suggests that historical exposure to a political culture of systematic information manipulation may have durable effects on bureaucratic behaviour or political norms that current V-Dem indices do not fully absorb.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-elections-finding"&gt;Q8. What is the elections finding?&lt;/h3&gt;
&lt;p&gt;Countries holding national parliamentary or presidential elections during 2020–2021 (76 of 134 countries) show significantly higher misreporting, consistent with electoral incentive theories of information manipulation. This finding is robust to including controls for GDP per capita, population age structure, and other domains, and is stable across the 2020-only and 2021-only sub-samples.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-performed"&gt;Q9. What robustness checks are performed?&lt;/h3&gt;
&lt;p&gt;The authors conduct: (1) specification-curve analysis across all combinations of covariates; (2) a joint model with all 12 individual factors; (3) principal component analysis within each domain to recover common variation and reduce dependence on specific measurement choices; (4) alternative expected-mortality models (Supplementary Material B.1); (5) alternative MRR normalisations (Supplementary Material B.2); (6) separate year-by-year analysis for 2020 and 2021; (7) inclusion of Bangladesh, China, and Indonesia as robustness cases despite lower data reliability; (8) addition of GDP per capita to check whether the institution-misreporting link is proxying for development; (9) analysis using underdispersion (Kobak 2022) and Benford&amp;rsquo;s law deviations as alternative manipulation measures; (10) exploration of colonial legacy as an additional historical variable (no significant effect found). The primacy of institutional constraints is robust across all of these.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-authors-treat-china-bangladesh-and-indonesia"&gt;Q10. How do the authors treat China, Bangladesh, and Indonesia?&lt;/h3&gt;
&lt;p&gt;These three large countries are excluded from the main analysis because their all-cause mortality data come from surveys (Bangladesh, China) rather than vital registration systems, or are very incomplete (Indonesia), making excess mortality estimation unreliable. They are included in a robustness regression (Table 4, column 6) and results are described as qualitatively similar. The authors flag that China&amp;rsquo;s data may itself be informative as a potential indicator of data suppression.&lt;/p&gt;
&lt;h3 id="q11-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q11. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;The paper is closest in spirit to Olken (2007), who uses the gap between reported and actual infrastructure spending to measure corruption, and Martinez (2022), who compares GDP growth to night-time-light-implied growth and finds autocracies overstate growth by more than a third. The authors extend this approach to a different domain (health/mortality) with broader country coverage. Prior Covid-specific work documented anomalies — underdispersion (Kobak 2022) and Benford&amp;rsquo;s law deviations (Kapoor et al. 2020; Kilani 2021) — and noted that autocratic regimes reported lower-than-expected deaths (Annaka 2021; Cassan and Van Steenvoort 2021), but these studies relied on regime type as the sole or primary explanatory variable and did not systematically rank competing factors. Neumayer and Plümper (2022) and Wigley (2024) used the authors&amp;rsquo; own World Mortality Dataset to test data manipulation. This paper is distinctive in that it: (a) provides what the authors describe as the most systematic estimates to date of Covid mortality and misreporting; (b) examines a broad range of factors across four domains without a priori privileging any; (c) directly tests and rejects capacity and false-positive aversion as alternative explanations; and (d) identifies communist legacy and elections as additional significant correlates.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Three implications are highlighted. First, unconstrained regimes appear to manipulate not only economic statistics but also health information during the most salient public policy event of the era; travel restrictions and multilateral actions during the pandemic relied on reported Covid figures, so manipulation had direct international externalities. This raises broader questions about the credibility of official data from such governments across domains — foreign aid targeting, climate action, vaccination campaigns. Second, the MRR provides a comparable cross-country measure of institutional quality grounded in actual governmental behaviour, potentially useful as an input to studies of institutions, conflict, electoral outcomes, and economic performance. Third, some countries that score respectably on conventional executive constraint indices — Albania, El Salvador, India, Serbia — show high MRRs, suggesting these rates may be leading indicators of democratic erosion not yet captured by standard measures. The scope condition the authors flag is external validity: if pandemic mortality is an extreme case with unique incentive structures (tourism, investment, aid eligibility), then findings about determinants of manipulation may not generalise beyond crisis settings. The authors argue against this interpretation on the grounds that macroeconomic factors — which would be pandemic-specific — are not significant, while institutional constraints — which reflect general governmental behaviour — are.&lt;/p&gt;
&lt;h3 id="q13-what-limitations-do-the-authors-acknowledge"&gt;Q13. What limitations do the authors acknowledge?&lt;/h3&gt;
&lt;p&gt;First, the analysis is explicitly descriptive rather than causal; factors are correlates, not proven determinants. Second, the MRR may understate true manipulation if all-cause mortality data are themselves selectively withheld or manipulated; the authors argue this is probably modest but acknowledge it cannot be fully ruled out. Third, important large countries — Pakistan, Nigeria, Ethiopia, Venezuela — cannot be scored because sufficient all-cause mortality data are not publicly available; the authors note this absence may itself be informative but cannot be quantified. Fourth, data on other causes of excess deaths (traffic accidents, suicides, homicides) are patchy in many countries, though the scale of these adjustments is very small. Fifth, some capacity controls (PWC) use data from as early as 2003, introducing measurement error. The paper does not claim to fully separate the channels through which institutions reduce manipulation (electoral accountability, press scrutiny, judicial oversight, professional agency independence), treating them as joint constraints rather than separately identified mechanisms.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Misreporting Rate (MRR)&lt;/strong&gt;: The paper&amp;rsquo;s central measure, defined as (estimated Covid deaths minus officially reported Covid deaths) divided by expected total deaths for the country in the same period based on pre-pandemic trends. A positive MRR indicates underreporting; a negative MRR indicates overreporting. Normalising by expected total deaths rather than by reported Covid deaths accounts for differences in population size, age structure, and baseline mortality across countries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Excess mortality&lt;/strong&gt;: The number of deaths above and beyond what would have been expected in the absence of the pandemic, estimated from country-specific models with weekly or monthly fixed effects and an annual trend fitted to 2015–2019 data. Used as the primary building block for estimated Covid deaths after subtracting excess deaths due to identified non-Covid causes (conflict, natural disasters, traffic accidents, homicides, suicides).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Death Registration Completeness (DRC)&lt;/strong&gt;: In this paper&amp;rsquo;s usage, the share of all deaths in a country captured by its vital registration system each year, measured using pre-pandemic data. Treated as the most basic indicator of a country&amp;rsquo;s capacity to count deaths. Used as a control to separate capacity constraints from intentional manipulation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Percent of Well-Certified Death Registrations (PWC)&lt;/strong&gt;: The share of death certificates in a country that carry a properly specified cause of death, measured using pre-pandemic data. Used alongside DRC as a second capacity control capturing not just whether deaths are registered but whether causes are correctly attributed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Informational Autocrat&lt;/strong&gt;: Following Guriev and Treisman (2022), the paper uses this concept to describe executives in countries where formal and informal checks and balances are weak, who systematically manipulate public information to project competence and reduce accountability. The paper&amp;rsquo;s empirical results are interpreted as evidence that such executives behave as informational autocrats not only in economic statistics but also in health data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;False-positive aversion&lt;/strong&gt;: The tendency of some countries to apply a higher evidentiary bar before attributing a death to a specific cause — such as Covid — rather than leaving the cause unspecified, independently of capacity or intention to deceive. The paper operationalises this using pre-pandemic ICD-10 data on specificity of reported causes of death and shows it is uncorrelated with MRR, ruling it out as a driver of observed discrepancies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Communist legacy&lt;/strong&gt;: The paper&amp;rsquo;s binary indicator for countries that had a communist or socialist regime for at least 10 consecutive years (34 countries). The variable captures historical exposure to a political culture of systematic information manipulation and is found to be a significant positive predictor of MRR even after conditioning on current institutional constraints, consistent with persistent norms or bureaucratic practices.&lt;/p&gt;</description></item><item><title>Market Opacity and Fragility: Why Liquidity Evaporates When It Is Most Needed</title><link>https://macropaperwarehouse.com/papers/market-opacity-and-fragility-why-liquidity-evaporates-when-it-is-most-needed/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/market-opacity-and-fragility-why-liquidity-evaporates-when-it-is-most-needed/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: The paper asks why market liquidity sometimes behaves in a stabilizing way (an illiquidity hike curbs liquidity demand and attracts liquidity supply) but on other occasions &amp;ldquo;evaporates when it is most needed,&amp;rdquo; degenerating into a disorderly run for the exit and a flash crash, often with no fundamentals news. Motivated by flash events (the May 6, 2010 US flash crash where the Dow Jones fell about 9% intraday; the October 15, 2014 Treasury crash; the August 24/25, 2015 ETF freeze; the 1987 crash; and the COVID-19 Treasury market dislocation), Cespa and Vives argue that lack of transparency about order flow is a key ingredient that can jam the &amp;ldquo;rationing&amp;rdquo; function of the cost of trading.&lt;/p&gt;
&lt;p&gt;Model setup: It is a stylized, two-period (trading rounds) rational-expectations model with no noise traders and no asymmetric information about payoffs — only about order flow. A single risky asset (liquidation value v ~ N(0, 1/tau_v)) is traded by competitive CARA agents. There are risk-averse dealers with risk tolerance gamma: a mass mu in [0,1] of &amp;ldquo;full&amp;rdquo; D-dealers present in both periods and 1-mu &amp;ldquo;restricted&amp;rdquo; RD-dealers present only in period 1; both post price-contingent (limit) orders. Overlapping unit-mass cohorts of risk-averse hedgers (risk tolerance gamma_H) receive independent endowment shocks u_t ~ N(0, 1/tau_u) in a non-tradable, perfectly correlated security and submit MARKET orders. Second-period hedgers observe a noisy signal s_u1 = u1 + eta of the first-period order imbalance, with eta ~ N(0, 1/tau_eta); tau_eta indexes transparency (infinity = full transparency, 0 = full opacity). The authors solve for linear equilibria and introduce a novel total-illiquidity measure, the Weighted Average Price Impact (WAPI), which volume-weights the heterogeneous price impacts of u1, u2, and eta.&lt;/p&gt;
&lt;p&gt;Main findings and mechanism: Under full transparency, second-period hedgers can perfectly infer u1, face no price (execution) risk, and supply liquidity via contrarian marketable orders (speculative aggressiveness b &amp;gt; 0); the price impacts of the two cohorts&amp;rsquo; shocks (Lambda_2 and Lambda_21) are independent, liquidity demand slopes DOWN in trading cost, and the equilibrium is unique. Under opacity the signal is noisy (b = 0 under full opacity), Lambda_2 and Lambda_21 become strategic SUBSTITUTES, generating strategic complementarity in illiquidity that can produce MULTIPLE equilibria and make liquidity demand slope UP in trading cost. Multiplicity arises when 0 &amp;lt; tau_u&lt;em&gt;tau_v &amp;lt; gamma/(4&lt;/em&gt;(gamma+gamma_H)^3): three equilibria (two stable extremal, one unstable intermediate). Example with tau_u = 0.1, tau_v = 0.1, gamma = 1, gamma_H = 0.1: Lambda_2 in {8.96, 1.98, 0.12}, Lambda_21 in {0.12, 1.98, 8.96}, Lambda_1 in {0.0001-ish (10^-2), 0.43, 8.84}; with tau_u = 2 a unique equilibrium with Lambda_21 = Lambda_2 = 4.61, Lambda_1 = 2.34. Traders facing the LARGEST trading cost trade most intensely at equilibrium.&lt;/p&gt;
&lt;p&gt;Quantitative comparative statics: An unanticipated, perceived-permanent rise in endowment-shock dispersion produces a flash crash raising WAPI by 44% (from 4.62 to 6.67) and price volatility by 70% (from 4.62 to 7.87); recovery restores the original equilibrium. Halving tau_v raises WAPI by 89% and price volatility by 138%; an 11% decline in gamma raises WAPI by 20% and volatility by 14% (the latter preserving a unique equilibrium — fragility without multiplicity). With restricted dealers, an 11% cut in mu (0.9 to 0.8) when transparency is low can plunge the market to the opposite equilibrium: Lambda_2 from 1.47 to 9.6 (a 653% jump) and WAPI from 5.7 to 10.3 (+80%); a 10% cut (mu 1 to 0.9) raises WAPI from 4.55 to 6.19 (+36%) without multiplicity.&lt;/p&gt;
&lt;p&gt;Implications: When the equilibrium is unique, total welfare is increasing in transparency (tau_eta) and in the mass of always-present dealers (mu), with gains accruing to hedgers and a transfer away from dealers. This supports policies for cheaper, consolidated order-flow information (EU/UK consolidated tape; US Treasury post-trade transparency; the SEC February 2024 dealer rule), while flagging a trade-off: more transparency can erode dealer participation, particularly for riskier securities.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-mechanism-that-turns-a-benign-illiquidity-hike-into-a-liquidity-rout"&gt;Q1. What is the core mechanism that turns a benign illiquidity hike into a liquidity rout?&lt;/h3&gt;
&lt;p&gt;Order-flow opacity. When second-period hedgers cannot observe the first-period endowment shock u1, the price impacts of the first- and second-period shocks (Lambda_21 and Lambda_2) become strategic substitutes: a higher Lambda_2 makes the price more driven by u2, raising cohort-1 hedgers&amp;rsquo; execution risk and shrinking their liquidity demand (|a21| down), which lowers Lambda_21, which in turn lowers cohort-2 execution risk and boosts their demand (|a2| up), further raising Lambda_2. This self-reinforcing loop (formalized by an aggregate best-response Phi(Lambda_2) that is strictly increasing in Lambda_2) is the strategic complementarity that can yield multiple equilibria and fragility. Under transparency the loop is killed because Lambda_2 and Lambda_21 are independent.&lt;/p&gt;
&lt;h3 id="q2-how-is-this-an-identificationequilibrium-selection-question-rather-than-an-empirical-one"&gt;Q2. How is this an &amp;lsquo;identification&amp;rsquo;/equilibrium-selection question rather than an empirical one?&lt;/h3&gt;
&lt;p&gt;This is a theory paper with no econometric identification. The analogue of &amp;lsquo;identification&amp;rsquo; is equilibrium selection and the formal conditions for multiplicity. The sufficient conditions for fragility are: overlapping cohorts of risk-averse hedgers suffering endowment shocks and submitting market orders; enough opacity about period-1 order flow; and risk-averse dealers. The necessary condition for multiplicity is sufficiently strong strategic complementarity, which is increasing in opacity. The closed-form multiplicity region is 0 &amp;lt; tau_u&lt;em&gt;tau_v &amp;lt; gamma/(4&lt;/em&gt;(gamma+gamma_H)^3).&lt;/p&gt;
&lt;h3 id="q3-how-does-the-model-distinguish-a-liquidity-dry-up-from-a-flash-crash"&gt;Q3. How does the model distinguish a &amp;rsquo;liquidity dry-up&amp;rsquo; from a &amp;lsquo;flash crash&amp;rsquo;?&lt;/h3&gt;
&lt;p&gt;Both arise when an unexpected shock (a jump in endowment-shock dispersion, i.e. a fall in tau_u, or a rise in dealer risk aversion / fall in gamma, or a fall in tau_v) pushes a market from a unique high-liquidity equilibrium into the multiplicity region and best-response dynamics attract it to a low-liquidity equilibrium. A dry-up is the transition to low liquidity; a flash crash is the same plus rapid recovery once the shock dissipates, all over a short interval. A shock to dispersion gravitates the market to the high-Lambda_2/low-Lambda_21 equilibrium; a shock to dealer risk aversion gravitates it to the low-Lambda_2/high-Lambda_21 equilibrium; in both, WAPI and price volatility rise.&lt;/p&gt;
&lt;h3 id="q4-what-does-the-wapi-measure-add-and-why-is-it-needed"&gt;Q4. What does the WAPI measure add and why is it needed?&lt;/h3&gt;
&lt;p&gt;Because period-2 price reacts with DIFFERENT impacts to u1, u2, and the signal noise eta (coefficients Lambda_21, Lambda_2, Lambda_22), no single price coefficient captures total illiquidity. WAPI is a volume-weighted average of these price impacts, with weights given by the expected absolute volumes from equilibrium responses (using E|z| = sqrt(2/pi)*sigma_z for normals). It is analogous to a volume-weighted spread for an order that walks the book. WAPI is shown to be U-shaped in transparency tau_eta, even though total welfare is monotonically increasing in tau_eta.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-role-of-the-contrarian-marketable-order-by-second-period-hedgers"&gt;Q5. What is the role of the contrarian marketable order by second-period hedgers?&lt;/h3&gt;
&lt;p&gt;With good information on u1, second-period hedgers post a contrarian market(able) order (b &amp;gt; 0) that offsets the first cohort&amp;rsquo;s selling/buying pressure, providing additional risk-sharing, enhancing the market&amp;rsquo;s risk-bearing capacity, and rationalizing first-period hedgers&amp;rsquo; decision to split their order across rounds. b is increasing in signal precision tau_eta. Under full opacity b = 0 because hedgers cannot predict the direction of the period-1 imbalance, so only dealers absorb the imbalance and risk-bearing capacity collapses.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-across-equilibria-and-cohorts-is-documented"&gt;Q6. What heterogeneity across equilibria and cohorts is documented?&lt;/h3&gt;
&lt;p&gt;At fragile (multiple) equilibria, trading costs are heterogeneous across cohorts: Lambda_2 and Lambda_21 are negatively correlated (one high, the other low). The cohort facing the HIGHEST market impact demands MORE liquidity (hedging intensity is increasing in the cost of trading it induces). Dealers speculate (consume liquidity) more aggressively in the most illiquid equilibrium — consistent with HFTs stepping up liquidity demand during extreme moves (Brogaard et al. 2018; Bellia et al. 2022). The persistence parameter beta = Lambda_21/Lambda_2 equals 1 at unique/intermediate equilibria (random walk noise), and beta&amp;gt;1 is an indicator of multiple equilibria and fragility.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-welfare-results-and-their-scope-conditions"&gt;Q7. What are the welfare results and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Restricted to the UNIQUE-equilibrium case (because with multiplicity hedger payoffs are complex-valued and cannot be ranked), and computed numerically with gamma = gamma_H = 1, tau_v = 1, tau_u = 2: total welfare TW(mu; tau_eta) is increasing in both transparency tau_eta and dealer mass mu. The gain is driven by higher hedger certainty equivalents (CEH_1, CEH_2); restricted dealers&amp;rsquo; CE falls with tau_eta, and D-dealers&amp;rsquo; CE falls with mu and (when tau_eta is not too small) with tau_eta. So transparency/dealer-presence policies raise welfare via a transfer from liquidity providers to consumers. A well-defined-payoffs condition is gamma_H^2&lt;em&gt;tau_u&lt;/em&gt;tau_v &amp;gt; 1 (which, when tau_eta=0 and mu=1, also implies a unique equilibrium).&lt;/p&gt;
&lt;h3 id="q8-what-is-the-transparency-versus-dealer-participation-trade-off"&gt;Q8. What is the transparency-versus-dealer-participation trade-off?&lt;/h3&gt;
&lt;p&gt;More transparency spurs second-period hedgers&amp;rsquo; speculation, eroding dealers&amp;rsquo; profits, which in a free-entry sense raises effective entry costs and induces some dealer exit (lower mu). Keeping total welfare constant against rising tau_eta requires a smaller mu cut for riskier securities (tau_v = 1) than for safer ones (tau_v = 3). Hence moderate transparency increases can reduce always-present dealer mass and may hurt welfare, especially for risky securities. With low transparency, raising mu has a NON-MONOTONIC effect on fragility (can move from multiple to unique and back), so enhancing transparency — not just dealer presence — is the key tool to eliminate fragility.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-relate-to-and-differ-from-prior-fragility-literature"&gt;Q9. How does the paper relate to and differ from prior fragility literature?&lt;/h3&gt;
&lt;p&gt;It departs on three dimensions: (i) the disruptive strategic complementarity is on the liquidity DEMAND side, not the supply side (unlike Brunnermeier-Pedersen 2009, Gromb-Vayanos 2002 funding constraints, Cespa-Foucault 2014, Cespa-Vives 2015); (ii) fragility relies on NO irrationality, noise trading, or exogenous demand/supply (unlike crash models of Gennotte-Leland 1990, Jacklin et al. 1992, Madrigal-Scheinkman 1997); (iii) asymmetric information is about the order flow, not payoffs. It also endogenizes an AR(1) noise-trading process whose persistence beta is determined in equilibrium. It supersedes the authors&amp;rsquo; earlier working paper Cespa-Vives (2019).&lt;/p&gt;
&lt;h3 id="q10-how-does-the-model-map-to-fragmentation-and-otc-markets"&gt;Q10. How does the model map to fragmentation and OTC markets?&lt;/h3&gt;
&lt;p&gt;Trading rounds 1 and 2 can be reinterpreted as separate venues; opacity then captures the limited flow of order information across venues, and mu (always-present dealers) is a reduced-form proxy for fragmentation-related dealer presence. Results should hold a fortiori in fragmented OTC markets, which are more opaque than centralized ones. Unlike Chen-Duffie (2021), Malamud-Rostek (2017), and Manzano-Vives (2021) — where fragmentation can raise welfare via traders&amp;rsquo; price impact — here traders are competitive, so those advantages do not arise.&lt;/p&gt;
&lt;h3 id="q11-what-robustness-and-extension-checks-are-reported"&gt;Q11. What robustness and extension checks are reported?&lt;/h3&gt;
&lt;p&gt;The partially-opaque case (finite tau_eta) is studied numerically: one or three equilibria can arise, with multiplicity when transparency is low; b&amp;gt;0 and increasing in tau_eta dampens complementarity. The general model with restricted dealers and partial opacity is simulated (Figure 9 partitions (mu, tau_eta) into unique vs. multiple-equilibria regions). Remark 1 allows period-specific endowment variances (tau_u1, tau_u2) and confirms the substitutes logic; as tau_u1 to infinity the transparent solution is recovered. Internet Appendices cover a partially informative signal, comparative statics for tau_v and gamma_H, the AR(1) noise process, the case where first-period hedgers observe u2, and a ranking of hedging aggressiveness across regimes (Corollary 11).&lt;/p&gt;
&lt;h3 id="q12-what-real-world-episodes-does-the-model-claim-to-rationalize-and-how-is-the-empirical-case-made"&gt;Q12. What real-world episodes does the model claim to rationalize, and how is the empirical case made?&lt;/h3&gt;
&lt;p&gt;It is consistent with the May 6, 2010 flash crash, the 2015 ETF freeze (where uncertainty over ETF constituents sidelined arbitrageurs and the SPY-RSP spread reached 21 dollars at one point), and the COVID-19 US Treasury dislocation around March 12, 2020 (spreads up roughly tenfold and depth virtually disappearing, per Duffie 2023). Empirical support for non-standard liquidity provision via contrarian marketable orders is drawn from Brogaard et al., Biais et al. (2017), Anand et al. (2013, 2021). The paper itself runs calibrated simulations (normal-volatility tau_v=1,tau_u=2 giving ~30% return volatility per Yuan 2005; and a liquidity-crisis tau_v=tau_u=0.1 case) rather than original econometric estimation.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;placeholder&lt;/strong&gt;: placeholder&lt;/p&gt;</description></item><item><title>Medical innovation and health disparities</title><link>https://macropaperwarehouse.com/papers/medical-innovation-and-health-disparities/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/medical-innovation-and-health-disparities/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks why medical innovation can widen health disparities even when it unambiguously improves health for everyone who takes it. The authors argue that the standard access-versus-preferences dichotomy is a false one: disadvantaged patients can rationally forgo effective medications because treatment side effects interfere with work, and the income cost of not working is particularly severe for low-education workers who hold physically demanding, inflexible jobs. Health-maximizing and welfare-maximizing behavior are therefore not the same thing, and the gap between the two is systematically larger for lower-education individuals.&lt;/p&gt;
&lt;p&gt;The empirical setting is the introduction of Highly Active Antiretroviral Therapy (HAART) for HIV in the mid-1990s. HAART was substantially more effective than prior mono- and combo-therapy at preventing AIDS progression and death, but it produced harsh physical side effects (fatigue, diarrhea, headache, fever). Data come from the Multi-Center AIDS Cohort Study (MACS), a semi-annual panel of men who have sex with men in Baltimore, Chicago, Pittsburgh, and Los Angeles, covering 1991–2003. After sample restrictions, the analysis uses 11,290 person-visit observations for 1,201 HIV-positive individuals aged 30–64, approximately 63% of whom hold a college degree or more. The study dichotomizes education into less-than-college versus college-or-more and tracks treatment choices, labor supply, immune-system health (CD4 count, with AIDS threshold at 250), physical ailments, income, insurance, and out-of-pocket medical expenditures.&lt;/p&gt;
&lt;p&gt;The structural model is a lifecycle discrete-choice dynamic programming framework in which forward-looking individuals simultaneously choose treatment (no treatment, monotherapy, combotherapy, and post-1995 HAART) and full-time work or non-work each half-year period to maximize expected lifetime utility. Health and survival evolve stochastically as functions of prior health, treatment, and age. Utility is a function of consumption (income minus out-of-pocket expenses), ailments, and labor supply, with utility parameters allowed to differ by education. The model is estimated via maximum likelihood using nested backwards induction; the quasi-experimental introduction of HAART as an unanticipated shock helps identify utility parameters.&lt;/p&gt;
&lt;p&gt;Key quantitative results: (1) HAART drastically reduced mortality for both groups—six-month mortality fell from 9% to 2% for less-educated men and from 6% to 1% for college graduates—and raised the probability of maintaining a high CD4 count from 62% to 78% (less-educated) and 68% to 83% (college+). (2) Despite equivalent access (both groups face roughly 91-95% insurance coverage and similarly low out-of-pocket costs), lower-educated men adopted HAART at a lower rate (58% of post-HAART visits versus 66% for college graduates) and approximately five months later. (3) The structural utility parameters confirm that while the direct disutility of ailments is not significantly different across education groups, the disutility of working while experiencing ailments is substantially larger in magnitude for less-educated men (estimated parameter -2.73) than for college graduates (-1.97). (4) Measured as expected lifetime utility, HAART&amp;rsquo;s introduction increased value for low-CD4 men by 236.1% (less-educated) versus 176.6% (college+), but in absolute utility units the gains were larger for college graduates—establishing that HAART increased welfare inequality. (5) Decompositions show the largest single driver of the education gap in HAART value is the differential survival process; income differences also matter but financial access variables (insurance, out-of-pocket costs) explain little. (6) A simulated six-month HAART mandate improves health—by 1.7 percentage points more for less-educated men—but reduces expected lifetime value by 2.8% for the less-educated versus 1.4% for college graduates, and reduces employment by 4.1% versus 1.6%, as mandated HAART forces men into ailment-producing treatment whose side effects they cannot manage alongside work. (7) A counterfactual $10,000-per-six-months non-labor income subsidy (similar to COVID-19 transfer policies) reduces work by 31–49% for less-educated men and by 25–39% for college graduates, while inducing an 81.2% increase in HAART take-up among less-educated men in good health who were not previously on treatment (from 5% to 9% baseline probability), and a 44.5% increase for similar college graduates (8% to 11%). For men with AIDS-level CD4 counts not on treatment, the policy raises the probability of being healthy next period by 12.6% for less-educated men and 5.3% for college graduates.&lt;/p&gt;
&lt;p&gt;The central mechanism is a wedge between health and welfare that is steeper for disadvantaged workers: occupational conditions make it harder to work while experiencing side effects, so the opportunity cost of HAART compliance is higher. This means effective medical innovation—precisely by creating more severe side effects than older regimens—can widen welfare inequality even as it compresses mortality gaps. Clinical trials that randomize assignment to treatment and measure health outcomes will register the innovation as a success while masking the distributional welfare costs. Policy interventions that reduce the cost of not working (income transfers, labor market restructuring) can simultaneously increase HAART take-up and improve health, with effects concentrated among the disadvantaged.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-main-identification-strategy-and-what-are-the-key-threats-to-identification"&gt;Q1. What is the main identification strategy and what are the key threats to identification?&lt;/h3&gt;
&lt;p&gt;The model is estimated by maximum likelihood using nested backwards induction over observable state variables. A key identifying variation is the quasi-experimental, unanticipated introduction of HAART in 1995, which shifts the choice set mid-panel and allows the authors to trace behavioral responses to an exogenous change in treatment efficacy and side-effect profiles. Disutility of ailments and work parameters are identified by conditional choice probabilities given state variables (health, ailment status, prior treatment) and by comparing behavior before and after HAART availability. The authors follow Magnac and Thesmar (2002) to establish that under the distributional assumptions (Type I EV shocks, fixed discount factor β=0.95) and the normalization imposed, the likelihood has a unique maximum. The main threats are: (a) the assumption that individuals were surprised by HAART (no forward-looking anticipation), which simplifies the model but is explicitly noted—Hamilton et al. (2021) show that incorporating individual expectations substantially complicates the framework; (b) the exclusion of unobserved heterogeneity in the utility function, though specifications including it produce very small probabilities of a second type (below 5%); (c) the absence of borrowing and saving, which could allow more educated individuals to smooth consumption across treatment cycles—the authors note this would bias downward the disutility of working with ailments for higher-educated individuals, meaning the estimated cross-education difference in that parameter is a lower bound; (d) the sample is restricted to white men in four cities, limiting external validity; and (e) the education dichotomy collapses heterogeneity within education groups.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-through-which-education-moderates-the-health-welfare-tradeoff-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms through which education moderates the health-welfare tradeoff, and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The paper identifies two nested channels. First, the estimated structural utility parameter for working while experiencing ailments is larger in magnitude for less-educated men (θ = -2.73) than for college graduates (θ = -1.97), indicating greater disutility from combining work and side effects. The paper argues this reflects occupational sorting: lower-education men are significantly more likely to hold manual occupations (occupation score 5.12 versus 4.49 for college graduates, where higher scores indicate more manual tasks per Autor et al. 2003), making physical side effects especially incompatible with job performance. Second, lower-educated men have lower incomes ($15,373 versus $22,290 per half-year for less-educated versus college-educated, pre-HAART), so the income cost of not working is larger in relative terms, creating stronger incentives to maintain employment even at the cost of forgoing treatment. The authors decompose the relative contribution of these mechanisms in the non-labor income subsidy simulation: when they give lower-educated men the income process of higher-educated men (Appendix Figure A1), the gap in behavioral response narrows but does not close; when they give lower-educated men the disutility parameters of higher-educated men (Figure A2), similarly the gap narrows but remains. Both mechanisms are jointly operative.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-in-haart-take-up-and-welfare-value-is-documented"&gt;Q3. What heterogeneity in HAART take-up and welfare value is documented?&lt;/h3&gt;
&lt;p&gt;Education is the primary heterogeneity dimension examined. Post-HAART, lower-educated men used HAART in 58% of observations versus 66% for college graduates, were slower to start (5 months later on average), and less likely to ever use it (67% versus 81%). Health status interacts with education: low-CD4 men gain more in percentage terms from HAART because they are more in need of its health-improving effects (236.1% gain for less-educated low-CD4 versus 176.6% for college-educated low-CD4; 85.7% versus 76.3% for high-CD4 men, with college graduates gaining more in absolute utility units throughout). The welfare cost of a treatment mandate is higher for less-educated men (2.8% lifetime value decline versus 1.4%), and the employment reduction induced by the mandate is also larger for them (4.1% versus 1.6%). In the income subsidy simulation, low-CD4 men not on any medication show the largest health response. The paper does not examine race/ethnicity heterogeneity, having excluded non-white individuals from the analysis due to sampling methodology concerns.&lt;/p&gt;
&lt;h3 id="q4-what-does-the-value-decomposition-reveal-about-why-haart-benefited-more-educated-men-more"&gt;Q4. What does the value decomposition reveal about why HAART benefited more-educated men more?&lt;/h3&gt;
&lt;p&gt;Table A17 sequentially replaces the processes and parameters of lower-educated agents with those of higher-educated agents. Giving lower-educated men the income process of college graduates narrows but does not close the gap—income is not the primary driver. Replacing the insurance and medical expenditure processes slightly reduces value for less-educated men relative to giving them only the income process, because more-educated individuals actually have somewhat higher out-of-pocket costs. Changing the health and ailments processes has modest positive effects. The largest single contributor to closing the education gap is the survival process: less-educated men face much higher baseline mortality, which depresses the expected present value of all future flows including the gains from HAART. This suggests that policies targeting survival differentials (e.g., access to other health services) could partially close the HAART welfare gap. Finally, replacing the utility parameters mechanically closes the remaining gap, but preferences are less amenable to direct policy intervention than the survival process.&lt;/p&gt;
&lt;h3 id="q5-what-do-the-treatment-mandate-simulations-show-and-why-do-they-matter-for-evaluating-clinical-trials"&gt;Q5. What do the treatment mandate simulations show, and why do they matter for evaluating clinical trials?&lt;/h3&gt;
&lt;p&gt;A six-month HAART mandate mimics randomized assignment to treatment in a clinical trial. It improves health—the probability of high CD4 rises by 1.7 percentage points more for less-educated men than baseline (reflecting a larger baseline gap in HAART use)—which would appear a policy success from a health-only perspective. However, expected lifetime utility falls by 2.8% for less-educated men and 1.4% for college graduates, because mandated HAART forces individuals into ailment-inducing treatment they would not have chosen, inhibiting labor supply. Employment falls by 4.1% for less-educated men versus 1.6% for college graduates. Appendix analyses removing the ailment-producing properties of treatment largely eliminate both the welfare cost and the employment effect, confirming that ailments are the mediating channel. This shows that clinical trials—which typically report health endpoints and do not measure welfare or distributional consequences—can mask the costs that effective but side-effect-heavy treatments impose, and that those costs fall disproportionately on less-advantaged patients.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-non-labor-income-subsidy-simulation-show-and-which-groups-respond-most"&gt;Q6. What does the non-labor income subsidy simulation show, and which groups respond most?&lt;/h3&gt;
&lt;p&gt;A permanent $10,000-per-six-months increase in non-employment income (approximately 50% of median income, calibrated to COVID-era transfer policies) induces labor force exit across all groups but concentrates its health-promoting effects among disadvantaged men who were not already on HAART. Among relatively healthy (high-CD4) less-educated men not using any medication, HAART take-up rises by 81.2% (from 5% to 9%); the corresponding figure for college graduates is 44.5% (from 8% to 11%). Among men with AIDS-level (low) CD4 not on treatment, the probability of being healthy next period increases by 12.6% for less-educated men and 5.3% for college graduates. Men already on HAART—who are unlikely to change treatment regardless—show little response. The policy has small but positive health externalities beyond the immediate recipients, since people on antiretrovirals have lower viral loads and lower transmission risk. Decomposition simulations (Appendix Figures A1–A2) show that both the income-level channel and the disutility-of-work-with-ailments channel independently contribute to the larger lower-education response, with neither alone sufficient to fully explain the differential.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The paper is most closely related to Papageorge (2016, Quantitative Economics), which uses the same MACS data and setting to link non-uptake of HAART to labor supply and side effects. The key difference is scope: Papageorge (2016) focuses on individual-level mechanisms; the present paper&amp;rsquo;s goal is to characterize distributional differences in the health-welfare tradeoff across education groups and to show that innovation can exacerbate existing inequality. Chan, Hamilton, and Papageorge (2016, Review of Economic Studies) also use the MACS setting to study the value of medical innovation, and Hamilton, Hincapié, Miller, and Papageorge (2021, International Economic Review) examine the diffusion of HAART. Relative to the sociological fundamental cause theory literature (Link and Phelan 1995; Phelan et al. 2010), which documents that medical innovations tend to widen health disparities, the present paper provides a structural quantification of the specific mechanisms and their relative magnitude. Relative to papers attributing health disparities primarily to access barriers (insurance, cost), the paper provides evidence that for this sample—where insurance coverage exceeds 91% even for less-educated men and HIV drugs are inexpensive—access explains little of the educational disparity in HAART use or health outcomes.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The core implication is that policies reducing the cost of not working—income transfers, disability benefits, worker protections—can raise HAART adoption and improve health among disadvantaged patients, precisely the group for whom standard health-access policies have limited traction. The non-labor income subsidy simulation suggests that the health improvements are modest in absolute magnitude (a 0.2% rise in probability of being healthy next period for the best-responding group among high-CD4 non-HAART users, and 13% for low-CD4 non-HAART users), but there are unmodeled positive externalities through reduced transmission risk that would multiply the social return. Scope conditions: (1) The sample is white men who have sex with men in four U.S. cities during 1991–2003, enrolled in a prospective cohort study; generalizability to other populations (women, racial minorities, other diseases) is uncertain. (2) The income subsidy that triggers HAART take-up must be large enough to induce labor force exit; a $10,000 per-six-months transfer is needed to generate the simulated behavioral response, larger for higher-income workers. (3) The paper explicitly notes that drug costs and insurance are not binding constraints in this sample, and the policy conclusions may differ in settings with weaker drug coverage. (4) Mental health is excluded from the model; the paper shows depression variables have smaller effects on treatment choice than the physical mechanisms included, but mental health could independently affect some populations&amp;rsquo; response. The paper&amp;rsquo;s conclusions extend to other conditions where effective treatment has disabling side effects and disadvantaged patients hold inflexible physical jobs—the authors invoke COVID-19 as a contemporary analog.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-conducted"&gt;Q9. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The authors report several robustness exercises. Treatment transition results are shown to be robust to defining the HAART introduction period as survey visit 23 or 25 rather than 24. Ailment specifications are noted to be robust to varying the type or frequency of ailments counted (citing Papageorge 2016 for this). Specifications including unobserved heterogeneity in the utility function produce very small second-type probabilities (below 5%), arguing against its inclusion. The treatment mandate simulations are run under three alternative shock-assignment methods (2 draws, 8 draws, and the preferred 2-draw approach), with results consistent across methods on the main welfare-versus-health asymmetry. Appendix Tables A19 and A20 remove ailments from all medications and from HAART only, respectively, confirming that the welfare cost of mandates is driven by treatment-induced ailments. Appendix Figures A1 and A2 mechanically decompose the education-differential response to the income subsidy by replacing income processes and disutility parameters separately, confirming that both channels are active. The model fit (Table A9) shows overall employment (66% model, 66% data) and HAART use (33% model, 36% data) closely matching, though the model slightly over-predicts medication use among low-CD4 individuals.&lt;/p&gt;
&lt;h3 id="q10-why-does-the-paper-focus-on-white-men-only-and-what-does-this-imply-for-interpretation"&gt;Q10. Why does the paper focus on white men only, and what does this imply for interpretation?&lt;/h3&gt;
&lt;p&gt;The authors drop 1,098 observations from 390 non-white individuals because of concerns about the sampling methodology used to recruit the refresher sample for those individuals—specifically, non-white participants entered the panel via a different selection process that could confound estimates. The paper does not investigate racial disparities in HAART take-up, which are also well-documented in the literature. This is a significant limitation because HIV/AIDS has disproportionately affected Black men in the United States, and the mechanisms the paper identifies—occupational sorting, income constraints, disutility of working with ailments—may operate differently or more intensely along racial lines. The authors acknowledge this limitation and note that the structural framework could in principle be applied to other groups if appropriate data were available.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Health-welfare tradeoff&lt;/strong&gt;: In this paper, the wedge between the action that maximizes health (taking effective medication despite side effects) and the action that maximizes lifetime utility (avoiding medication to remain employed and maintain income). The tradeoff is not a bias or error but a rational response to economic constraints, and it is wider for less-educated individuals whose occupational conditions make working with side effects especially costly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HAART (Highly Active Antiretroviral Therapy)&lt;/strong&gt;: A combination antiretroviral HIV treatment introduced in the mid-1990s, far more effective than prior mono- or combo-therapy at improving CD4 count and preventing AIDS-level immune decline and death. In this paper&amp;rsquo;s model, HAART serves as the innovation whose adoption the authors study: it is more efficacious but produces harsher side effects than earlier treatments, and its introduction is treated as an unanticipated aggregate shock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Disutility of working with ailments&lt;/strong&gt;: A structural utility parameter (θ_2,f=0) capturing how much worse-off an agent feels from working while experiencing physical ailments (fatigue, diarrhea, headache, fever). Estimated at -2.73 for less-educated men and -1.97 for college graduates, this parameter is the primary driver of the differential health-welfare tradeoff across education groups and explains why side-effect-bearing treatments like HAART are disproportionately avoided by lower-education workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Treatment mandate simulation&lt;/strong&gt;: A counterfactual in which all agents are assigned to HAART for six months (eliminating choice among other treatment options), used to mimic randomized assignment in a clinical trial. The simulation is designed specifically to illustrate that health improvements observable in a clinical trial coexist with welfare reductions and employment disruptions that would not be captured in standard trial endpoints.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fundamental cause theory&lt;/strong&gt;: A sociological framework (Link and Phelan 1995) arguing that socioeconomic status is a &amp;lsquo;fundamental cause&amp;rsquo; of health disparities that persists despite or is even amplified by medical innovation, because more advantaged individuals are better positioned to adopt and benefit from new treatments. The paper provides structural economic microfoundations for this theory by quantifying the mechanisms through which HAART&amp;rsquo;s introduction widened the welfare gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-labor income subsidy&lt;/strong&gt;: A counterfactual policy simulation in which non-employment income is raised by $10,000 per six months (approximately 50% of the median person&amp;rsquo;s income), modeled after COVID-19 transfer policies. In the paper&amp;rsquo;s model this policy reduces employment but increases HAART take-up and health improvements particularly for less-educated HIV-positive men who were previously forgoing treatment to maintain income from work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Source text origin&lt;/strong&gt;: Not a paper-specific concept but denoted here: the full working paper text was obtained from the NBER Working Paper (No. 28864), not from abstract-only, satisfying the GUARD requirement.&lt;/p&gt;</description></item><item><title>Nonlinear Monetary Policy Tradeoffs</title><link>https://macropaperwarehouse.com/papers/nonlinear-monetary-policy-tradeoffs/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/nonlinear-monetary-policy-tradeoffs/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper measures how the inflation-unemployment tradeoff associated with monetary policy varies with both the sign of the monetary intervention (easing versus tightening) and the state of the business cycle (booms versus recessions) for the US economy over 1973:M1 to 2019:M6. The motivation is that standard linear Phillips-curve estimates implicitly impose a constant tradeoff, yet a flat Phillips curve would simultaneously predict that (i) stimulating activity during a recession costs nothing in terms of inflation and (ii) reducing inflation costs very large amounts of unemployment — both empirically extreme predictions that have very different policy implications. The paper challenges both extremes.&lt;/p&gt;
&lt;p&gt;The empirical strategy extends the Proxy-SVAR approach of Mertens-Ravn (2013) and Stock-Watson (2018) to a nonlinear setting. The economy is described by a Vector Moving Average augmented with nonlinear functions of the monetary policy shock — specifically its absolute value (capturing sign dependence) and its interaction with a recession indicator (capturing state dependence). Under a finite-order VARX representation assumption and a linear monetary policy rule assumption, the paper proves (Proposition 1) that even though the underlying VARX is nonlinear, the monetary shock can be recovered as the projection of an external instrument onto residuals of a misspecified linear VAR. Once the shock is recovered, it and its nonlinear functions are used as regressors in a VARX to estimate nonlinear impulse responses. The instrument is the Degasperi-Ricco (2022) extension of Miranda-Agrippino and Ricco (2021), with a baseline span of 1991:M1-2015:M12 extrapolated to the full sample. The VAR contains five variables: the 1-year Treasury bond rate, industrial production growth, the Gilchrist-Zakrajsek excess bond premium, the unemployment rate, and CPI inflation, estimated with 7 lags. The recession indicator equals 1 when average GDP growth over the previous 12 months is negative.&lt;/p&gt;
&lt;p&gt;The monetary policy tradeoff is defined analogously to the fiscal multiplier: the ratio of the cumulative average impulse response of inflation (unemployment) to the cumulative average impulse response of unemployment (inflation) over horizons H. In a nonlinear setting the easing tradeoff and tightening tradeoff are no longer inverses of one another and must be treated separately.&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. For monetary easing during recessions, the inflation cost of reducing unemployment is small and statistically insignificant: point estimates of T+ range from -0.03 to -0.17 (in absolute value) across horizons H = 12 to H = 48 months, with 68% confidence intervals spanning from approximately -5.3 to +2.8 at H = 12 and -3.4 to +2.7 at H = 48. For monetary tightening during booms, the unemployment cost of reducing inflation is moderate and statistically significant: T- estimates range from -0.51 to -0.61 across H = 12 to H = 48, with 68% confidence intervals entirely below zero (e.g., -1.10 to -0.26 at H = 12 and -1.23 to -0.24 at H = 48). In other words, reducing inflation by 1 percentage point during a boom requires raising unemployment by roughly 0.5 to 0.6 percentage points. These results are qualitatively robust to excluding the post-2008 zero-lower-bound period (pre-2009 subsample) and to alternative specifications. By contrast, monetary tightening during recessions implies a very large and unfavorable tradeoff. Easing during booms is extremely inflationary with virtually no real effect.&lt;/p&gt;
&lt;p&gt;A Likelihood Ratio test for the null hypothesis that all nonlinear terms are zero is rejected at the 1% level, confirming the statistical importance of nonlinearities. The null hypothesis of shock invertibility (Assumption A4) is not rejected at the 5% level across all combinations of VAR lags and residual leads tested.&lt;/p&gt;
&lt;p&gt;A simple model with downward nominal wage rigidities — in which the wage floor introduces a kink in the aggregate supply curve — provides a theoretical rationale for the sign- and state-dependent tradeoff: an expansionary shock in a full-employment economy raises inflation with no output effect (the economy sits on the vertical AS segment), while a contractionary shock makes the wage rigidity binding and reduces output with no price effect (the horizontal AS segment). Monte Carlo validation using artificial data generated by the calibrated DSGE model shows that the proposed empirical procedure recovers the theoretical nonlinear impulse responses very accurately.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-assumptions-required"&gt;Q1. What is the identification strategy and what are the main assumptions required?&lt;/h3&gt;
&lt;p&gt;Identification proceeds in two steps. First, the monetary shock is recovered by projecting an external instrument (Degasperi-Ricco 2022) onto the residuals of a standard linear VAR — this is justified by Proposition 1, which shows that even though the VAR is misspecified (it omits the nonlinear terms), the shock can still be recovered as a linear combination of VAR residuals under four assumptions: (A0) a structural VMA representation in which the shock is orthogonal to past observables and to the remaining structural shocks at all leads and lags; (A1) a finite-order VARX representation; (A2) invertibility of the Wold representation; (A3) a valid instrument (relevance and exogeneity); and (A4) informational sufficiency, meaning the monetary shock can be expressed as a linear combination of current and past observables — a condition implied by a linear monetary policy rule. Second, once the estimated shock and its nonlinear functions (absolute value and interaction with the state dummy) are in hand, they are used as exogenous regressors in a VARX to estimate nonlinear impulse response functions.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-threats-to-identification"&gt;Q2. What are the main threats to identification?&lt;/h3&gt;
&lt;p&gt;Three main threats are acknowledged. (1) Instrument validity: if the instrument (Degasperi-Ricco 2022) is weak or contaminated by information shocks, the first-stage projection may recover a mislabeled shock. The authors note the first-stage F-statistic is adequate per Miranda-Agrippino and Ricco (2021) but acknowledge that the weak-instrument problem in the nonlinear context is non-trivial and left for future research. (2) Assumption A4 (informational sufficiency): if the central bank follows a nonlinear rule or the VAR variables are not sufficient to recover the shock, identification fails. The authors test this using the Forni-Gambetti-Ricco (2023) invertibility test — regressing the instrument on current and future VAR residuals and checking whether future residuals matter — and fail to reject invertibility at 5% across all lag/lead combinations. (3) Model misspecification in the nonlinear VARX: the VARX approximation may not capture all relevant nonlinearities generated by the true DSGE. The Monte Carlo validation on artificial DSGE data provides reassurance that the approach recovers the true nonlinear responses accurately.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-distinguish-sign-dependence-from-state-dependence"&gt;Q3. How does the paper distinguish sign dependence from state dependence?&lt;/h3&gt;
&lt;p&gt;The paper includes two nonlinear terms as regressors in the VARX: the absolute value of the shock |u_t^r|, which captures sign-dependent effects (i.e., whether a tightening and an easing of equal magnitude have asymmetric effects), and the product s_{t-1} * u_t^r, which captures state-dependent effects (i.e., whether the same-sign shock has different effects depending on whether the economy was in a recession before the shock arrived). The two components are estimated simultaneously, allowing their separate contributions to be read off impulse responses in Figure 3. Robustness checks in the Online Appendix report models estimated with only sign dependence and only state dependence in isolation, with results described as qualitatively similar to Barnichon-Matthes (2018) and Tenreyro-Thwaites (2016), respectively.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-key-quantitative-results-on-impulse-responses"&gt;Q4. What are the key quantitative results on impulse responses?&lt;/h3&gt;
&lt;p&gt;In the full nonlinear model, monetary tightening generates large and significant effects on real variables (unemployment, industrial production) regardless of the state, while monetary easing has more muted real effects. For prices, sign and state components operate in opposite directions: the largest inflation responses are associated with tightening during expansions. Numerically, the tradeoff estimates from Table 2 show: (a) easing during recessions — T+ point estimates of -0.03 at H=12, -0.12 at H=24, -0.17 at H=36, -0.17 at H=48 months (all statistically insignificant at 68%); (b) tightening during booms — T- point estimates of -0.51 at H=12, -0.61 at H=24, -0.59 at H=36, -0.53 at H=48 months (all statistically significant at 68%). For the pre-2009 subsample (excluding the ZLB period), tightening-in-booms estimates are somewhat larger in absolute value (-0.63 to -0.70) but confidence intervals widen to include zero at longer horizons.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-key-implication-for-pushing-on-a-string-results-in-the-prior-literature"&gt;Q5. What is the key implication for &amp;lsquo;pushing on a string&amp;rsquo; results in the prior literature?&lt;/h3&gt;
&lt;p&gt;Tenreyro-Thwaites (2016) and Barnichon-Matthes (2018) document that monetary easing is less effective at stimulating real activity, especially during recessions — an apparent &amp;lsquo;pushing on a string&amp;rsquo; result. The current paper accepts that the real effect of easing in recessions is muted, but adds a crucial dimension: price responses are also muted in the same circumstances, so the inflation-unemployment tradeoff is actually favorable even when the absolute size of real effects is small. The policy implication is that central banks can still usefully deploy monetary easing during recessions as long as interventions are sufficiently aggressive to achieve the desired stimulus, since the inflationary cost of doing so is low.&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-measure-the-tradeoff-differently-from-phillips-curve-regressions"&gt;Q6. How does this paper measure the tradeoff differently from Phillips-curve regressions?&lt;/h3&gt;
&lt;p&gt;The tradeoff is defined as the ratio of the cumulative average impulse response of inflation to the cumulative average impulse response of unemployment (or vice versa) in response to an identified monetary shock, analogous to a fiscal multiplier. This approach avoids three problems that plague standard Phillips-curve estimates: (i) it does not require specifying a structural Phillips-curve equation, reducing misspecification risk; (ii) it does not require data on inflation expectations or the natural rate of unemployment, which are unobserved and introduce measurement error; (iii) identification comes from exogenous monetary shocks rather than OLS variation in unemployment, so the endogeneity problem is avoided.&lt;/p&gt;
&lt;h3 id="q7-what-theoretical-mechanism-rationalizes-the-nonlinear-tradeoffs"&gt;Q7. What theoretical mechanism rationalizes the nonlinear tradeoffs?&lt;/h3&gt;
&lt;p&gt;A simple New-Keynesian-style model with downward nominal wage rigidities (Wt &amp;gt;= theta * W_{t-1}) generates a kink in the aggregate supply curve. When the economy operates at full employment and inflation is non-negative, an expansionary monetary shock stimulates demand but the wage rigidity is non-binding, so the economy sits on the vertical segment of the AS curve: output cannot exceed its natural level, and the only effect is higher inflation. By contrast, a contractionary shock makes the wage rigidity binding, pushing the economy onto the flat segment of the AS curve: firms cut employment rather than nominal wages, so output falls but prices are unaffected. More generally, averaging over periods of full employment and periods of involuntary unemployment, tightening has larger real effects and weaker price effects than easing — matching the empirical pattern — because a contractionary shock keeps the economy below full employment for a longer time.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-are-conducted"&gt;Q8. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;Three main robustness checks are reported in the main text, each presented with impulse-response figures (Figures 6, 7, 8): (1) replacing the authors&amp;rsquo; state dummy (based on 12-month average GDP growth) with NBER recession dates; (2) replacing the 1-year Treasury bond rate with the Federal Funds rate and with the 6-month Treasury Bill rate; (3) replacing the baseline Degasperi-Ricco instrument with the Jarocinski-Karadi (2020) instrument both raw and cleaned (regressed on six lags of VAR variables). In all cases, the qualitative result — tightening in booms produces larger real effects than easing in recessions, while price responses are more muted in recessions — is preserved, and the tradeoff pattern remains favourable for easing in recessions and tightening in booms. The Online Appendix additionally reports results using: the unemployment rate as the state variable (instead of industrial production); the VAR extended with the 10-year Treasury Bill rate and M2 monetary aggregate; models with only sign dependence; models with only state dependence; and an alternative estimation using the instrument directly in place of the estimated shock (which yields implausible results, validating the two-stage procedure).&lt;/p&gt;
&lt;h3 id="q9-what-does-the-monte-carlo-validation-using-the-dsge-model-establish"&gt;Q9. What does the Monte Carlo validation using the DSGE model establish?&lt;/h3&gt;
&lt;p&gt;The paper generates 1000 artificial realizations from a calibrated downward-nominal-wage-rigidity DSGE model (beta=0.99, sigma=1, theta=1, phi_pi=1.5, rho_m=0.5, sigma_r=0.25%, sigma_a=0.45%, solved by nonlinear global projection using Chebyshev polynomials). It then applies the nonlinear Proxy-SVAR procedure to each artificial dataset and compares average estimated impulse responses with average true (model-generated) generalized impulse responses. The two are described as &amp;lsquo;very similar&amp;rsquo; (Figure 10), demonstrating that the empirical nonlinear VARX representation accurately approximates the nonlinearities of the DSGE even though the VARX is in principle misspecified relative to the true model. This validates both the econometric procedure and the interpretive link between the empirical findings and the theoretical mechanism.&lt;/p&gt;
&lt;h3 id="q10-why-does-the-paper-estimate-the-shock-from-a-misspecified-linear-var-rather-than-the-varx-directly"&gt;Q10. Why does the paper estimate the shock from a misspecified linear VAR rather than the VARX directly?&lt;/h3&gt;
&lt;p&gt;The monetary shock is latent. Proposition 1 shows that, under the stated assumptions, the monetary shock equals (up to a scaling constant) the projection of the external instrument onto the VAR residuals of the linear VAR, even though the VAR omits the nonlinear terms. This is because the linear monetary policy rule implies the shock is a linear combination of current observables, and the VAR residuals span the same space. Using the instrument directly in the VARX instead of going through steps I and II introduces a non-proportional bias in the nonlinear case (unlike the linear case where the attenuation bias from measurement error in the instrument is proportional across units and corrects under normalization). The Online Appendix shows that bypassing the two-stage shock-estimation procedure yields implausible impulse response estimates.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-scope-of-the-empirical-findings-and-what-caveats-apply"&gt;Q11. What is the scope of the empirical findings and what caveats apply?&lt;/h3&gt;
&lt;p&gt;Three scope conditions are explicitly stated. (1) State uncertainty: the tradeoff varies significantly with the state of the economy, so if the central bank is uncertain about current economic conditions, interventions carry considerable risk — a disinflation during what turns out to be a weaker-than-anticipated economy could incur very large unemployment costs. (2) Historical average: estimates reflect the effects of average monetary interventions over 1973-2019 and may not generalize to unusually large, persistent, or unconventional policy actions. (3) Accompanying fiscal policy: the tradeoff could be influenced by fiscal policy measures that accompanied monetary interventions during the sample period. The sample also excludes the post-2019 inflation surge, so inference about that episode is not direct. The identification requires a valid external instrument, whose strength in the nonlinear context is an open question.&lt;/p&gt;
&lt;h3 id="q12-how-does-this-paper-relate-to-barnichon-mesters-2020-2021-and-gali-gambetti-2020"&gt;Q12. How does this paper relate to Barnichon-Mesters (2020, 2021) and Gali-Gambetti (2020)?&lt;/h3&gt;
&lt;p&gt;Barnichon-Mesters (2020, 2021) and Gali-Gambetti (2020) also exploit identified monetary shocks to estimate the conditional inflation-unemployment relationship (the &amp;lsquo;Phillips multiplier&amp;rsquo;) and to investigate whether the Phillips curve slope has changed over time. The main additional contribution of the present paper is to show that the relationship is not only time-varying but specifically sign- and state-dependent, driven by the direction of monetary intervention and the current phase of the business cycle. The sign- and state-dependent tradeoff framework provides a richer characterization that can explain why a flat aggregate Phillips curve is compatible with moderate costs of disinflation and low inflationary costs of stimulus — something a time-varying-slope model alone does not deliver.&lt;/p&gt;
&lt;h3 id="q13-what-does-the-paper-say-about-the-implications-for-disinflation-episodes-like-2022-23"&gt;Q13. What does the paper say about the implications for disinflation episodes like 2022-23?&lt;/h3&gt;
&lt;p&gt;The paper does not directly analyze the 2022-23 episode (the sample ends at 2019:M6 and the paper was written with November 2025 dating for the online appendix). However, the results imply that if the economy is in a boom when disinflation begins — as was broadly the case in 2022 — the unemployment cost of reducing inflation is moderate (roughly 0.5-0.6 percentage points of unemployment per percentage point of inflation at a 24-36 month horizon), substantially less than would be implied by a flat Phillips curve. The authors explicitly note that their results suggest central banks can pursue disinflation without necessarily incurring very large unemployment costs, subject to the caveats about state uncertainty and scale of the intervention.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Monetary policy tradeoff&lt;/strong&gt;: In this paper&amp;rsquo;s usage: the ratio of the cumulative average impulse response of inflation to the cumulative average impulse response of unemployment (for easing) or vice versa (for tightening), in response to an identified monetary shock, averaged over a horizon H. In a linear model easing and tightening tradeoffs are inverses; in the nonlinear model they must be estimated separately. The concept is deliberately defined without assuming a Phillips curve and without requiring inflation expectations or the natural rate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sign dependence&lt;/strong&gt;: The property that a monetary easing and a monetary tightening of equal magnitude have asymmetric effects on inflation and unemployment, not just opposite-signed effects of the same absolute magnitude. Captured in the VARX by including the absolute value of the monetary shock as an exogenous regressor.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;State dependence&lt;/strong&gt;: The property that the effects of a monetary shock of given sign and magnitude differ depending on whether the economy was in a recession or a boom in the period before the shock arrived. Captured in the VARX by including the product of the recession indicator (s_{t-1}) and the monetary shock as an exogenous regressor.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nonlinear Proxy-SVAR&lt;/strong&gt;: The paper&amp;rsquo;s proposed econometric framework: a Vector Moving Average augmented with nonlinear functions of the monetary shock, which admits a VARX representation. Identification extends the standard Proxy-SVAR by showing — via Proposition 1 — that the latent monetary shock can be recovered from the residuals of a misspecified linear VAR, using an external instrument, under a linear monetary policy rule. The estimated shock and its nonlinear functions are then used as exogenous regressors to recover nonlinear impulse response functions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Downward nominal wage rigidity&lt;/strong&gt;: A labor market friction, modeled as the constraint W_t &amp;gt;= theta * W_{t-1}, that creates a kink in the aggregate supply curve. When the constraint binds (during downturns), firms respond to contractionary shocks by cutting employment rather than nominal wages, generating unemployment without deflation. When the constraint is non-binding (during expansions), expansionary shocks raise nominal wages and prices without affecting employment beyond full-employment output. In this paper the rigidity is the key mechanism generating a sign- and state-dependent monetary tradeoff.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Informational sufficiency (Assumption A4)&lt;/strong&gt;: The identifying assumption that the monetary policy shock can be expressed as a linear combination of current and past observable variables — equivalently, that the central bank follows a linear monetary policy rule. This allows the shock to be recovered from the residuals of a standard linear VAR even when the true model is nonlinear. Tested empirically via the Forni-Gambetti-Ricco (2023) invertibility test (checking whether the instrument Granger-causes future VAR residuals); not rejected at the 5% level in the authors&amp;rsquo; data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Generalized Impulse Response Function (GIRF)&lt;/strong&gt;: In this nonlinear context, defined as E(x_{t+h} | u_t^r = u-bar) - E(x_{t+h} | u_t^r = 0) for h = 0, 1, &amp;hellip;, where u-bar is a given shock size. Unlike linear IRFs, GIRFs depend on the sign and magnitude of the shock and on the state of the economy, and are computed by summing the linear response alpha(L)*u-bar and the nonlinear response Phi(L)*g(u_t^r, &amp;hellip;).&lt;/p&gt;</description></item><item><title>On the Effects of Monetary Policy Shocks on Income and Consumption Heterogeneity</title><link>https://macropaperwarehouse.com/papers/on-the-effects-of-monetary-policy-shocks-on-income-and-consumption-heterogeneity/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/on-the-effects-of-monetary-policy-shocks-on-income-and-consumption-heterogeneity/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks how conventional and informational monetary policy shocks affect the cross-sectional distributions of labor earnings, consumption, and financial income in the United States. The motivation is the growing concern, particularly in the aftermath of the global financial crisis, about distributional consequences of central bank actions. Existing studies either include scalar inequality statistics in standard VARs — losing information about the full distribution — or rely on indirect approaches that hold household portfolio compositions fixed. Chang and Schorfheide instead apply the functional VAR (fVAR) framework developed in Chang, Chen, and Schorfheide (2024, JPE forthcoming) that stacks macroeconomic aggregates alongside the full time-varying cross-sectional density, represented as a log probability density function approximated via a cubic-spline sieve. This allows simultaneous, internally-consistent IRFs for percentiles, Gini coefficients, 90-10 ratios, standard deviations, and other distributional statistics without the risk of quantile crossings.&lt;/p&gt;
&lt;p&gt;The earnings analysis uses monthly micro data from the Current Population Survey (CPS), sample period 1990:M2 to 2016:M12. The consumption and financial income analyses use quarterly Consumer Expenditure Survey (CEX) data from 1990:Q2 to 2016:Q4. Monetary policy shocks are identified via the Jarocinski-Karadi (2020) high-frequency instruments — surprises in the three-month fed funds futures and in S&amp;amp;P 500 index — used as internal instruments in the structural VAR. The instruments isolate (a) conventional monetary policy shocks (interest rate surprise, stock price opposite direction) and (b) informational shocks (interest rate and stock price surprise in the same direction). Sign restrictions set-identify the two shocks. Bayesian estimation uses a Chan (2022) Normal-Inverse Gamma prior suitable for high-dimensional VARs; model selection (sieve order K, lag length p, hyperparameters) is done by maximizing the marginal data density (MDD). The shock normalization corresponds to an unanticipated 25-basis-point cut in the three-month federal funds rate.&lt;/p&gt;
&lt;p&gt;Main quantitative findings:&lt;/p&gt;
&lt;p&gt;Earnings (conventional shock): An expansionary shock reduces earnings inequality, primarily through the employment (extensive) margin. At the posterior median, the 10th earnings percentile rises by up to 5% relative to steady state, the 20th percentile by up to 1%, while the 80th and 90th percentiles are essentially unaffected. The Gini coefficient for labor earnings falls from approximately 0.431 to 0.428 over a 36-month horizon. The 90-10 earnings ratio falls from approximately 12.27 to 11.76 after 36 months. These effects are driven almost entirely by individuals moving from unemployment into employment (the point mass at zero in the earnings distribution falls as the unemployment rate drops by approximately 0.3 percentage points at the posterior median after three years). When the unemployed point mass is excluded from the inequality computation, the inequality effect is small and short-lived, confirming that the employment channel dominates. The estimated Gini drop of 0.001–0.003 is broadly consistent with the HANK model of Ma (2021) with indivisible labor, which predicts a drop of approximately 0.001 for a comparable shock.&lt;/p&gt;
&lt;p&gt;Consumption (conventional shock): The expansionary shock generates a weakly positive (inequality-increasing) effect on consumption inequality at the posterior median, but with wide credible bands that span both positive and negative values. The cross-sectional standard deviation of consumption, the 90-10 ratio, and the Gini coefficient all peak upon impact and remain above steady state. The slight increase appears concentrated in durable goods expenditure; nondurable and service consumption inequality shows little response at the posterior median. The contrast with the earnings result reflects: (i) only labor income is captured in the earnings analysis, while wealthy households&amp;rsquo; capital income (rising with equity and bond prices) also rises; (ii) potentially higher interest-rate sensitivity of high-consumption households.&lt;/p&gt;
&lt;p&gt;Financial income (conventional shock): No statistically significant effect on financial income inequality. The cross-sectional standard deviation and Gini coefficient of financial income do not respond to the shock. An important caveat is that the CEX misses the top-10 percent of households by financial income (visible from CDF comparison with the Survey of Consumer Finances in 2012). The households most likely to benefit from equity and bond price appreciation — captured in other studies — are absent from the sample.&lt;/p&gt;
&lt;p&gt;Informational shock: A negative informational shock (unexpected simultaneous drop in interest rates and stock prices, signaling worse-than-expected output) increases earnings inequality, mainly via a rise in unemployment. The 10th earnings percentile drops by about 2% at the posterior median. Consumption inequality, by contrast, shows the opposite pattern: the 90-10 ratio and Gini coefficient for consumption decrease, and the posterior median responses are negative, though uncertainty is substantial.&lt;/p&gt;
&lt;p&gt;Policy implication: The authors conclude that earnings inequality effects of conventional monetary policy are well-proxied by the unemployment rate response, so standard macro indicators subsume the distributional information for earnings. The small and highly uncertain responses of consumption and financial income inequality provide, in their view, support for central banks continuing to focus primarily on macroeconomic aggregates.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-monetary-policy-shocks-and-what-are-the-main-threats-to-validity"&gt;Q1. What is the identification strategy for monetary policy shocks and what are the main threats to validity?&lt;/h3&gt;
&lt;p&gt;The paper uses the Jarocinski-Karadi (2020) high-frequency instruments as internal instruments in a structural VAR. The two instruments are surprises in the three-month federal funds futures (ff4_hf) and surprises in the S&amp;amp;P 500 index (sp500_hf), measured in narrow windows around FOMC announcements. Sign restrictions separate two shocks: a conventional shock is identified by an interest rate increase combined with a stock price fall; an informational shock by both increasing. The key assumptions are instrument relevance (the instruments are correlated with the policy shocks) and instrument validity (the instrument innovations are uncorrelated with non-policy structural shocks). As a robustness check the authors also use the Nakamura-Steinsson (2018) instruments and report very similar results. The main threat to validity is the standard one for external-instrument SVARs: the instruments may capture other economic news released simultaneously with FOMC decisions, violating the exclusion restriction. The informational shock identification partially addresses this by explicitly modeling the central bank&amp;rsquo;s information revelation.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-functional-var-approach-and-why-is-it-preferred-over-simpler-alternatives"&gt;Q2. What is the functional VAR approach and why is it preferred over simpler alternatives?&lt;/h3&gt;
&lt;p&gt;The functional VAR stacks macroeconomic aggregates Yt with the time-varying cross-sectional log-density of micro outcomes. The log-density is approximated by a finite-dimensional linear sieve (cubic spline basis of order K). Sieve coefficients are estimated period-by-period by maximum likelihood from the cross-section, then treated as observations in a standard VAR. The MDD selects K, lag order p, and Minnesota-type hyperparameters jointly. Compared to simply including a few inequality statistics in a VAR, the functional approach (a) derives a single coherent model from which arbitrarily many distributional statistics can be computed without quantile crossings; (b) achieves tighter credible intervals by efficiently compressing cross-sectional information through the sieve; (c) avoids the problem of internally inconsistent forward projections of stacked quantile VARs. Compared to indirect approaches (e.g., McKay-Wolf 2023), it does not require the assumption that household income or portfolio composition is fixed in response to the shock. Compared to panel approaches, it does not require high-frequency panel data, which are unavailable for the US at relevant horizons.&lt;/p&gt;
&lt;h3 id="q3-how-is-the-earnings-distribution-modeled-to-handle-unemployment"&gt;Q3. How is the earnings distribution modeled to handle unemployment?&lt;/h3&gt;
&lt;p&gt;The earnings distribution is treated as a mixture of a point mass at zero (representing unemployed individuals, whose weight equals the CPS-based unemployment rate) and a continuous part (the density of positive earnings of employed individuals, normalized to integrate to one minus the unemployment rate). The sieve density is estimated only from the positive-earnings observations, with a top-coding adjustment for right-censored values. The unemployment rate is included separately as an aggregate variable in the Yt vector. This mixture representation allows the analysis to separately identify the extensive-margin (employment) channel — changes in the probability mass at zero — from the intensive-margin channel (changes within the positive-earnings density). The key finding is that inequality effects are driven almost entirely by the extensive margin.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-in-earnings-responses-is-documented"&gt;Q4. What heterogeneity in earnings responses is documented?&lt;/h3&gt;
&lt;p&gt;In percentage terms, the expansionary monetary policy shock has the largest impact at the 10th earnings percentile (posterior median response of 0 to 5%), capturing workers moving out of unemployment. The 20th percentile rises by 0 to 1%. The 80th and 90th percentiles show essentially zero response. Earnings above 2 times GDP per capita (roughly twice the labor share of GDP per capita) are essentially unaffected. When the point mass at zero is excluded and only the continuous part of the earnings distribution is analyzed, the effect on inequality statistics (Gini, 90-10 ratio) is small and short-lived, confirming that the heterogeneous response across the full distribution is driven almost entirely by the employment transition at the bottom.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-in-consumption-responses-is-documented-and-why-might-consumption-inequality-rise-while-earnings-inequality-falls"&gt;Q5. What heterogeneity in consumption responses is documented, and why might consumption inequality rise while earnings inequality falls?&lt;/h3&gt;
&lt;p&gt;At the posterior median, both the 10th and 20th consumption percentiles initially rise above steady state (h=1), then fall 0.9% to 1.3% below baseline from h=5 onwards. The 80th and 90th percentile responses are quantitatively similar in shape but slightly larger in magnitude, leading to a weakly positive net inequality effect. The Gini coefficient and 90-10 ratio for consumption peak upon impact and stay above steady state. The authors offer two explanations for the inequality-increasing result despite earnings inequality falling: (i) wealthy households also earn substantial capital income (equities, bonds) that rises with the expansionary shock, boosting their total resources and hence consumption, a channel not captured by earnings alone; (ii) higher-consumption households may have more interest-rate-sensitive consumption decisions (larger direct Euler-equation effect), or may be wealthy hand-to-mouth consumers with high MPCs. The component analysis shows the increase is concentrated in durable goods, while nondurable and services Gini responses are near zero at the posterior median.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-financial-income-analysis-find-and-what-data-limitation-is-most-important"&gt;Q6. What does the financial income analysis find and what data limitation is most important?&lt;/h3&gt;
&lt;p&gt;The financial income distribution estimated from the CEX shows no statistically significant response to either the level or inequality of financial income following a conventional monetary policy shock. The cross-sectional standard deviation and Gini coefficient of financial income are essentially flat. The most important caveat is that the CEX substantially underrepresents high-financial-income households. A CDF comparison with the Survey of Consumer Finances for 2012 shows that the CEX misses the top-10 percent of households by financial income. These are precisely the households most likely to experience capital gains from equity and bond price appreciation following an interest rate cut. The fraction of households with essentially zero financial income (the point mass κt) fluctuates between 0.65 and 0.82 over the sample, so the analysis is largely capturing the lower 65–82 percent of the financial income distribution.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-informational-shock-and-how-do-its-distributional-effects-differ-from-the-conventional-shock"&gt;Q7. What is the informational shock and how do its distributional effects differ from the conventional shock?&lt;/h3&gt;
&lt;p&gt;An informational shock is defined as an unanticipated change in interest rates that conveys private central-bank information about the state of the economy — for example, a rate cut that signals the central bank expects worse output and prices than the public. It is identified by the simultaneous drop in interest rates and stock prices, the opposite pattern from the conventional shock. Aggregate effects: real GDP drops approximately 20 basis points and unemployment rises up to 0.15 percentage points after one year. Earnings distributional effects are roughly the mirror image of the conventional shock: the 10th earnings percentile drops about 2% at the posterior median, while other percentiles change little. The Gini coefficient and 90-10 ratio for earnings rise in the long run, driven by the increase in unemployment. Consumption distributional effects are different: relative consumption at the 10th and 20th percentiles rises, while the 90th percentile falls slightly, so consumption inequality (90-10 ratio, Gini) decreases. However, since aggregate consumption also falls, the rise in relative consumption at the bottom does not imply an absolute gain.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-coibion-gorodnichenko-kueng-and-silvia-2017"&gt;Q8. How does this paper relate to and differ from Coibion, Gorodnichenko, Kueng, and Silvia (2017)?&lt;/h3&gt;
&lt;p&gt;CGKS (2017) include inequality statistics directly in a VAR and use the Romer-Romer shock measure. For earnings, they find the Gini coefficient rises by about 0.0025 per 100bp contractionary shock (i.e., falls by 0.0025 for an expansionary shock); adjusting for shock size this is slightly smaller than the Chang-Schorfheide estimate of a 0.001–0.003 Gini drop per 25bp expansionary shock (which scales to 0.004–0.012 per 100bp). For consumption, CGKS find that inequality decreases in response to an expansionary shock, the opposite sign from Chang-Schorfheide&amp;rsquo;s posterior-median result (weakly increasing). The discrepancy may reflect: (i) the functional approach&amp;rsquo;s more flexible modeling of the full distribution versus using a single Gini; (ii) differences in shock identification (Romer-Romer vs. JK instruments); (iii) sample period differences. The wide credible bands in the consumption result mean the two findings are not statistically inconsistent.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-conducted"&gt;Q9. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The authors run the following robustness exercises: (i) Nakamura-Steinsson (2018) instruments instead of Jarocinski-Karadi (2020) for the earnings VAR — results are very similar. (ii) Model selection across sieve order K ∈ {4,6,8,10} and lag length p ∈ {1,2,3,4} via MDD maximization, confirming that results are robust to the choice of approximation order. (iii) For the earnings inequality analysis, the paper explicitly separates the contribution of the employment margin from the wage distribution within employment, by recomputing inequality statistics excluding the point mass at zero — confirming that the employment channel dominates. (iv) Comparison of aggregate IRFs across all four model specifications (aggregate VAR, earnings fVAR, consumption fVAR, financial income fVAR) showing that inclusion of cross-sectional data does not substantially alter inference about aggregate variables. (v) Comparison with time-aggregated monthly-to-quarterly rescaled IRFs to validate that monthly and quarterly specifications produce consistent results.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-scope-conditions-and-limitations-of-the-findings"&gt;Q10. What are the scope conditions and limitations of the findings?&lt;/h3&gt;
&lt;p&gt;Key scope conditions: (a) The sample runs through 2016:Q4/M12, so the post-2016 period and the 2020 pandemic episode are excluded. (b) The paper uses repeated cross-sections rather than a panel, so it directly estimates how the cross-sectional distribution evolves but cannot separately identify cohort effects, individual trajectories, or nonlinearities in unit-level histories. (c) The CEX substantially misses high-financial-income households, making the financial income results inapplicable to the top 10% of the financial income distribution. (d) The functional VAR models the unconditional distribution; it does not identify heterogeneous responses by subgroup in the sense of comparing specific groups (e.g., mortgagors vs. owners) as pseudo-panel approaches do. (e) The approach identifies the average linear response to a 25bp shock; nonlinear or asymmetric effects (large shocks, ZLB periods) are not modeled. (f) The simultaneous drop in earnings inequality and (weakly) rising consumption inequality cannot be fully reconciled without a complete model including capital income; the paper acknowledges this limitation explicitly.&lt;/p&gt;
&lt;h3 id="q11-how-do-the-quantitative-results-compare-to-the-ma-2021-hank-model-benchmark"&gt;Q11. How do the quantitative results compare to the Ma (2021) HANK model benchmark?&lt;/h3&gt;
&lt;p&gt;Ma (2021) incorporates an indivisible labor supply mechanism into a HANK model and shows that an expansionary monetary policy shock raises wages, inducing low-productivity workers to enter the labor market, raising earnings in the left tail. His calibration produces a Gini coefficient drop of approximately 0.001 for a comparable shock (scaled from his Figure 3: −0.4/(4×100) = −0.001 on a 0-to-1 scale for a 100bp shock). The Chang-Schorfheide empirical estimate is a drop of between 0.001 and 0.003 for a 25bp shock, which is broadly consistent with Ma&amp;rsquo;s model. The qualitative mechanism — earnings inequality reduction driven by low-productivity workers transitioning out of unemployment — is also consistent with the Chang-Kim (2006) heterogeneous-agent model with indivisible labor, which generates a negative correlation between idiosyncratic productivity and reservation wage.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-for-central-banks"&gt;Q12. What are the policy implications for central banks?&lt;/h3&gt;
&lt;p&gt;The paper provides semi-structural empirical evidence relevant for central banks concerned about distributional effects. The main conclusion is that for labor earnings inequality, the distributional effect of conventional monetary policy is well-summarized by the unemployment rate response: reducing unemployment compresses earnings inequality, and a central bank that targets unemployment de facto targets earnings inequality. The small, uncertain, and sometimes-positive effects on consumption and financial income inequality suggest that tracking these additional distributional statistics adds little actionable information beyond what standard macro aggregates already convey. The authors therefore conclude that there is an empirical case for central banks to continue focusing on macroeconomic aggregates. An important qualifier is that the financial income results are constrained by CEX top-coding, so the analysis cannot speak to very-high-income households&amp;rsquo; welfare.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Functional VAR (fVAR)&lt;/strong&gt;: A vector autoregression in which macroeconomic aggregates are stacked with the full cross-sectional log-probability density function of micro outcomes. The log-density is approximated by a finite-dimensional sieve (cubic spline basis), with sieve coefficients estimated period-by-period from cross-sectional data and then entered as observations in a linear VAR. This yields coherent IRFs for the entire distribution — percentiles, Gini, 90-10 ratio, etc. — from a single model, avoiding the quantile-crossing inconsistency of stacked-quantile approaches.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Employment channel (extensive margin)&lt;/strong&gt;: In this paper, the mechanism by which an expansionary monetary policy shock lowers earnings inequality: it reduces the unemployment rate, moving workers from a point mass of zero earnings into the positive-earnings distribution. The paper distinguishes this from the intensive margin (changes in wage rates conditional on employment), and finds empirically that the extensive margin dominates the inequality response of labor earnings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Informational shock (central bank information shock)&lt;/strong&gt;: As defined following Jarocinski-Karadi (2020): an unanticipated change in short-term interest rates that conveys the central bank&amp;rsquo;s private assessment of economic conditions. Identified by the simultaneous movement of interest rates and stock prices in the same direction, opposite to a conventional monetary policy shock. A negative informational shock (rates and equity prices both fall) signals that the central bank expects weaker output and prices than the public, and leads in this paper to rising earnings inequality via higher unemployment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Point mass at zero (earnings distribution)&lt;/strong&gt;: The concentration of probability mass at zero earnings, corresponding to the fraction of individuals in the labor force who are unemployed (the CPS-based unemployment rate). The total earnings density is modeled as a mixture of this point mass and a continuous density for positive earnings. The IRF for the point mass is the IRF for the unemployment rate; including it in inequality computations is necessary to capture the full distributional effect of employment transitions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Log probability density function (log-pdf) sieve representation&lt;/strong&gt;: The modeling device that represents each period&amp;rsquo;s cross-sectional distribution as the logarithm of a probability density, approximated by a finite linear combination of cubic spline basis functions (order K chosen by MDD). Working in log-pdf space avoids non-negativity and monotonicity constraints, enabling coherent linear propagation through the VAR law of motion; the density is recovered by exponential normalization in each period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marginal data density (MDD) model selection&lt;/strong&gt;: The Bayesian integrated likelihood used in this paper to jointly select the sieve approximation order K, lag length p, and Minnesota-type hyperparameters. The MDD balances in-sample fit (the log-spline likelihood) against a dimensionality penalty, thereby avoiding overfitting. A key result is that the preferred earnings fVAR uses K = 10 with a single lag, while the smoother consumption distribution is adequately captured with K = 6.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;κt (financial income point mass)&lt;/strong&gt;: The time-varying fraction of households in the CEX with financial income below a threshold x (set at the 10th percentile of pooled standardized financial income ≈ 0.0014 of the capital share of per-capita GDP). κt fluctuates between 0.65 and 0.82 over 1990–2016, meaning 65–82 percent of households have negligible financial income in a given quarter. The CEX data constraint — missing the top-10 percent of high-financial-income households — is the principal limitation on the financial income analysis.&lt;/p&gt;</description></item><item><title>On the elasticity of substitution between labor and ICT and IP capital and traditional capital</title><link>https://macropaperwarehouse.com/papers/on-the-elasticity-of-substitution-between-labor-and-ict-and-ip-capital-and-traditional-capital/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/on-the-elasticity-of-substitution-between-labor-and-ict-and-ip-capital-and-traditional-capital/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper estimates the elasticities of substitution between labor, information and communication technology (ICT) and intellectual property (IP) capital, and traditional capital using a nested constant elasticity of substitution (CES) production function. The motivation is twofold: standard macroeconomic models aggregate all capital into a single input and thus miss potentially distinct substitution relationships, and competing estimates of the labor-capital elasticity of substitution diverge sharply — with some finding gross substitutability (Karabarbounis and Neiman 2013) and others gross complementarity (Glover and Short 2020) — leaving unexplained the observed decline in labor income share across advanced economies.&lt;/p&gt;
&lt;p&gt;The data come from the 2023 release of the EU KLEMS database for nine Euro Area economies (Austria, Belgium, Finland, France, Germany, Italy, Netherlands, Portugal, and Spain) over 1996-2020 (with Germany ending in 2019 and Portugal starting in 2001). The nesting structure places an ICT-IP capital aggregate (itself a CES nest of ICT equipment and IP capital, which includes software, databases, patents, and R&amp;amp;D capital) together with labor in an inner nest, and that combined aggregate is then nested with traditional capital in an outer nest. The rationale for grouping ICT and IP capital is their joint and complementary use — computers and software — and the observation that roughly 25% of granted patents in the sample period are ICT-related. Estimation follows the normalized CES methodology of Grandville (1989), Klump, McAdam, and Willman (2007), and Leon-Ledesma, McAdam, and Willman (2010), which jointly estimates the logged and normalized production function together with its first-order conditions using feasible generalized nonlinear least squares, weighting by country-year employment shares and correcting for heteroscedasticity and serial correlation. This approach is preferred because normalization anchors the point elasticity at sample averages and Monte Carlo evidence shows it outperforms first-order-condition-only or translog alternatives, especially when identifying factor-augmenting technological change alongside substitution elasticities.&lt;/p&gt;
&lt;p&gt;The main results (Table 4, column 1) are as follows. The elasticity of substitution between labor and traditional capital (ε1) is estimated at 0.745 (standard error 0.009), statistically significantly below 1, implying gross complementarity. The elasticity between labor and the ICT-IP aggregate (ε2) is 1.187 (0.010), significantly above 1, implying gross substitutability. The elasticity between ICT and IP capital themselves (ε3) is 0.961 (0.003), significantly below 1, implying gross complementarity within the ICT-IP nest. The ICT capital-augmenting technological change parameter (γ_ICT) is estimated at 0.725, several orders of magnitude larger than the labor-augmenting parameter (γ_L = 0.003), consistent with rapid technological progress in ICT. The IP capital-augmenting parameter (γ_IP) is negative (−0.111), and the traditional capital-augmenting parameter (γ_TK) is negative but statistically insignificant (−0.002). For the US, ε2 is substantially larger at 1.712 (0.133), with ε1 = 0.724 (0.024) and ε3 = 0.922 (0.017).&lt;/p&gt;
&lt;p&gt;A counterfactual accounting exercise (fixing ICT and IP technological progress indexes and capital stocks at their 1996 levels) finds that absent these developments, labor income share would have slightly increased in European countries rather than declining, and would have declined by about 75% less in the US over the sample period. ICT accumulation and technological progress is the dominant driver of the fall: absent ICT changes alone, labor share would have risen significantly in Europe.&lt;/p&gt;
&lt;p&gt;The paper also derives the implied aggregate labor-capital elasticity (εL,K) using Hicks&amp;rsquo;s formula applied to the nested production function. The imputed εL,K for European countries ranges from approximately 1.36 to 1.43 over 1996-2020, rising through 1996-2008 and declining afterward. The US imputed values are substantially higher, ranging from approximately 2.14 to 2.37. By contrast, when the author directly estimates a two-input CES function combining labor with aggregate capital, the estimated elasticity is significantly below 1 (approximately 0.988 for European countries in the constant-CES specification), far below the imputed values. This divergence demonstrates that production function specification is consequential for identifying the labor-capital elasticity, and that models treating all capital as a single input can generate downward-biased estimates of this parameter.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The author jointly estimates a normalized CES production function and first-order conditions (capital return equations and the wage equation) using feasible generalized nonlinear least squares with multiple starting points, selecting results by log likelihood, AIC, BIC, and R-squared. Normalization anchors the elasticity as a point elasticity at geometric sample averages, which is theoretically motivated and improves finite-sample identification. Main threats include: (1) endogeneity of factor inputs — the system of equations is estimated jointly but without instrumental variables, relying on non-arbitrage conditions to close the model; (2) negative estimates for γ_IP and γ_TK, which the author acknowledges may capture markups or capital underutilization rather than true technical change (Jiang and Leon-Ledesma 2018 show that omitting markups can bias the sign of capital-augmenting technology); (3) the US results are sensitive to initial values for the estimation algorithm, possibly because of the small sample size (24 observations); and (4) the counterfactual exercise abstracts from equilibrium effects and free-factor supply adjustments.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-distinguishing-the-three-capital-types-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms distinguishing the three capital types, and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;ICT capital (computers, communication devices, peripherals) and IP capital (software, databases, patents, R&amp;amp;D capital) are grouped in an inner nest on the grounds of their complementary joint use. Traditional capital (machinery, transport, construction and structures) forms the outer nest. This nesting allows the elasticity of substitution between labor and the ICT-IP aggregate (ε2 &amp;gt; 1, gross substitute) to differ from the elasticity between labor and traditional capital (ε1 &amp;lt; 1, gross complement), which the paper argues is consistent with the automation literature&amp;rsquo;s emphasis on ICT displacing routine tasks. The elasticity of substitution within the ICT-IP nest (ε3 &amp;lt; 1) reflects gross complementarity between ICT equipment and IP assets (one needs software to use computers). The empirical distinction comes from the separate first-order conditions for each capital type, which link each capital&amp;rsquo;s income share to its stock and price, allowing the three elasticities to be separately identified.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented-across-countries-or-time"&gt;Q3. What heterogeneity is documented across countries or time?&lt;/h3&gt;
&lt;p&gt;The main estimates pool 9 European countries weighted by employment shares; the author does not report country-by-country elasticity estimates but does report country-level descriptive statistics (Table I in the Data Appendix). Time-series heterogeneity is addressed through the imputed aggregate elasticity εL,K, which rises from approximately 1.367 in 1996 to a peak around 1.388-1.426 near 2008 (varying across the sensitivity columns of Table 6) and then declines to approximately 1.369-1.411 by 2020. The US elasticities are systematically higher than the European ones (εL,K ranging approximately 2.14-2.37 for the US vs. 1.36-1.43 for Europe; ε2 = 1.712 for the US vs. 1.187 for Europe). The time-varying aggregate capital specification in Table 7 shows the estimated ε1 for European countries follows an inverted-U shape over the sample period, while the US estimate shows the contrary pattern (though the latter is imprecise due to the small sample).&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper estimates two alternative CES nesting structures (equations 20 and 21, reported in columns 2 and 3 of Table 4) to assess sensitivity to the nesting assumption. In specification (20), labor and traditional capital are nested first and then combined with the ICT-IP aggregate, so the elasticity between labor and ICT-IP equals that between traditional capital and ICT-IP. In specification (21), the different capital types are nested first and then combined with labor. Both alternatives confirm that ICT and IP capital are gross substitutes for labor. The paper also estimates a two-input labor-aggregate capital function in three variants: constant CES, elasticity as a linear function of compensation shares and relative prices, and elasticity as a quadratic polynomial of time (Table 7). Results using US data from the EU KLEMS database are reported separately (column 4 of Table 4 and columns 8-9 of Table 6). The imputed εL,K is further verified using data counterparts of the compensation shares rather than model-predicted shares (column 7 of Table 6), yielding essentially identical results with higher variability.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Relative to Karabarbounis and Neiman (2013), this paper agrees that labor and aggregate capital are gross substitutes (imputed εL,K &amp;gt; 1) and that capital deepening drives the labor share decline, but attributes the mechanism specifically to ICT and IP capital accumulation rather than the fall in all capital prices. It contrasts with Glover and Short (2020), whose below-1 estimates the paper reconciles by showing that treating all capital as a single input biases the aggregate elasticity downward. Relative to Eden and Gaggl (2018, 2019), who use US data and find ICT (including software) substitutes for labor in first-order-condition-only estimates, this paper adds normalization and biased technical change parameters and uses European panel data, and also separates ICT equipment from IP/software. Relative to Koh, Santaeulalia-Llopis, and Zheng (2020), who perform an accounting exercise attributing the labor share decline to IP capital capitalization, this paper provides structural estimates of substitution elasticities and corroborates the IP capital importance. Relative to Aum and Shin (2024), who use Korean firm-level data and find software substitutes for labor while ICT equipment complements it, this paper uses a different nesting (ICT and IP grouped together) and European aggregate data, and finds the combined ICT-IP aggregate is a gross substitute for labor — consistent with Aum and Shin&amp;rsquo;s software result driving the within-nest finding. The normalization approach distinguishes the paper from Antras (2004) and earlier aggregate studies that estimate only first-order conditions (which can produce upward-biased elasticity estimates when biased technical change is omitted).&lt;/p&gt;
&lt;h3 id="q6-what-does-the-paper-find-about-the-source-of-the-labor-share-decline-and-what-are-the-scope-conditions-on-this-result"&gt;Q6. What does the paper find about the source of the labor share decline, and what are the scope conditions on this result?&lt;/h3&gt;
&lt;p&gt;The counterfactual exercise (Section 4.2, Panel B of Table 3) finds that absent ICT and IP capital technological progress and accumulation, labor income share would have slightly increased in European countries over 1996-2020 rather than falling. Absent ICT changes alone, labor share would have risen significantly in Europe. The ICT-driven decline is the dominant contributor. By contrast, absent IP capital trends, labor share would have fallen substantially more (suggesting IP capital compensation growth, when attributed to capital rather than labor, partially offsets the ICT effect on labor&amp;rsquo;s share but its own share rise is the proximate driver of labor share decline). For the US, absent ICT and IP developments, labor share decline would have been about 75% smaller. Scope conditions: this is a static accounting exercise holding free factors at initial values and abstracting from general equilibrium effects. The results apply to total industrial value added (not individual sectors) and to the nine Euro Area countries in the sample. The exercise assumes the estimated production function parameters are the correct structural parameters, and thus inherits any limitations of the identification strategy.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-implication-for-the-measured-aggregate-labor-capital-elasticity-and-why-does-it-differ-from-standard-estimates"&gt;Q7. What is the implication for the measured aggregate labor-capital elasticity, and why does it differ from standard estimates?&lt;/h3&gt;
&lt;p&gt;When the paper estimates a two-input (labor, aggregate capital) CES function directly, the estimated aggregate elasticity is significantly below 1 and close to estimates from Herrendorf, Herrington, and Valentinyi (2015). When it instead imputes the aggregate elasticity from the nested-CES parameter estimates using Hicks&amp;rsquo;s formula, the imputed values exceed 1 and are much larger. The paper shows analytically that εL,K &amp;gt; ε2 when the relative capital cost of ICT compared to traditional capital (pKICT&lt;em&gt;KICT / pTK&lt;/em&gt;TK) takes sufficiently low values, which is the case in the data. This divergence arises because the single-input capital specification conflates the high substitutability of labor with ICT-IP capital and the low substitutability with traditional capital, yielding a biased estimate that depends on the capital composition. The paper concludes that production function specification is consequential for identifying the aggregate labor-capital substitution elasticity.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-key-data-features-that-drive-the-results"&gt;Q8. What are the key data features that drive the results?&lt;/h3&gt;
&lt;p&gt;ICT investment prices fell at an average annual rate of -4.6% relative to value added prices over the sample, while IP and traditional capital investment prices changed by -0.3% and +0.1% per year, respectively. Real ICT capital stocks grew at 4.9% per year, versus 3.4% for IP capital and 1.6% for traditional capital. ICT and IP capital depreciate rapidly (20.1% and 24.1% per year) compared to traditional capital (3.6%). These patterns imply computed rates of return on ICT capital that were very high at the start of the sample (131% in 1996, largely reflecting the fall in ICT prices that year) and fell sharply to 24% by 2020. The average share of labor and ICT-IP compensation in value added is approximately 71%, with labor making up about 92% of that combined share. The ICT share within the ICT-IP nest is about 21%, meaning IP capital compensation is substantially larger than ICT capital compensation.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Allen-Uzawa elasticity of substitution&lt;/strong&gt;: A point elasticity measuring the percentage change in the ratio of two inputs in response to a percentage change in their price ratio, holding output and other input prices constant. In this paper, it is estimated as a structural parameter of the nested CES production function, normalized at sample geometric averages; values above 1 imply gross substitutability and values below 1 imply gross complementarity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Normalized CES production function&lt;/strong&gt;: A CES specification that is indexed to sample averages of output and inputs so that the elasticity of substitution is defined as a point elasticity at those averages. This normalization, following Grandville (1989) and Leon-Ledesma et al. (2010), facilitates identification of both elasticity parameters and factor-augmenting technological change parameters, avoiding the conflation that arises in unnormalized specifications.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gross substitutes / gross complements&lt;/strong&gt;: Two inputs are gross substitutes (elasticity of substitution &amp;gt; 1) if a fall in the relative price of one leads to a rise in the share of cost devoted to it, reducing the other input&amp;rsquo;s cost share. They are gross complements (elasticity &amp;lt; 1) if a fall in relative price instead reduces cost share. In this paper, labor and ICT-IP capital are gross substitutes; labor and traditional capital and ICT with IP capital are gross complements.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Traditional capital (TK)&lt;/strong&gt;: In this paper&amp;rsquo;s taxonomy, all non-ICT, non-IP capital: machinery, transport equipment, construction, and structures. It is the residual capital category and is defined as a gross complement of labor in the estimated nested CES structure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intellectual property (IP) capital&lt;/strong&gt;: Capital comprising software, databases, patents (including R&amp;amp;D capital), and other forms of intellectual property as measured in the EU KLEMS database. IP capital is grouped with ICT equipment in an inner CES nest on the grounds of complementary use. Its compensation share rise is the proximate accounting factor in the labor share decline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Factor-augmenting technological change&lt;/strong&gt;: Hicks-neutral or biased technical progress that enters multiplicatively with a specific factor input in the production function (e.g., γ_ICT for ICT capital), scaling the effective quantity of that input. In this paper, the ICT-augmenting parameter is estimated to be very large and positive (0.725), reflecting rapid ICT productivity growth, while IP- and traditional-capital-augmenting parameters are negative, which the author suggests may partly reflect markups or underutilization rather than pure technology.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Imputed aggregate labor-capital elasticity&lt;/strong&gt;: The elasticity of substitution between labor and total capital derived analytically from the nested CES parameters using Hicks&amp;rsquo;s formula, rather than estimated directly from a two-input specification. In this paper, the imputed value exceeds 1 for Europe (~1.36-1.43) and is substantially higher for the US (~2.14-2.37), contrasting with directly estimated values that are below 1, illustrating the sensitivity of this parameter to production function specification.&lt;/p&gt;</description></item><item><title>Optimal Combination of Patent Instruments in a Cumulative-Innovation Growth Model</title><link>https://macropaperwarehouse.com/papers/optimal-combination-of-patent-instruments-in-a-cumulative-innovation-growth-model/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/optimal-combination-of-patent-instruments-in-a-cumulative-innovation-growth-model/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper develops a tractable general equilibrium model of endogenous growth driven by cumulative innovation, and uses it to characterize optimal patent policy — both for patent breadth (via a &amp;ldquo;non-infringing inventive step&amp;rdquo; requirement) and patent length — with a focus on their welfare implications and optimal combination.&lt;/p&gt;
&lt;p&gt;The central motivation is that cumulative innovation creates positive knowledge spillovers: each new idea strictly builds on the best existing technology, and the disclosure that patenting requires diffuses knowledge to future innovators. Because private firms do not internalize these spillovers, the decentralized equilibrium features strictly lower R&amp;amp;D investment than the social optimum. The key wedge is an intertemporal spillover effect: firms discount future profits at a rate that includes the hazard of being superseded (rho + lambda&lt;em&gt;v&lt;/em&gt;L), while the social planner uses only the pure time preference rate (rho). Appropriability and business-stealing externalities exactly offset each other, so the intertemporal spillover is the sole source of under-investment.&lt;/p&gt;
&lt;p&gt;The model has a continuum of differentiated varieties, a single labor input, a Poisson idea arrival process (rate lambda per R&amp;amp;D worker), and productivity improvements drawn i.i.d. from a standardized Pareto distribution with shape parameter theta &amp;gt; 1. The Pareto structure yields the key tractability: the log of the k-th best productivity level is Gamma-distributed with mean k/theta, which allows closed-form welfare expressions. In steady state, all outcomes depend on just three deep parameters: the discount rate rho, the Pareto shape theta, and the innovative capacity lambda*L.&lt;/p&gt;
&lt;p&gt;The patent breadth instrument is formalized as a &amp;ldquo;non-infringing inventive step&amp;rdquo; (NIS) requirement B &amp;gt;= 1: a new idea must deliver a productivity at least B times the current patent-holder&amp;rsquo;s productivity to qualify for a patent. Raising B creates two opposing forces. The &amp;ldquo;profit effect&amp;rdquo; extends incumbent monopoly duration by reducing the hazard rate of supersession (from lambda&lt;em&gt;v&lt;/em&gt;L to lambda&lt;em&gt;v&lt;/em&gt;L&lt;em&gt;B^{-theta}), raising innovation incentives. The &amp;ldquo;hurdle effect&amp;rdquo; raises the bar an idea must clear to be patentable, reducing the expected return to R&amp;amp;D. These forces generate a non-monotonic (inverted-U) relationship between R&amp;amp;D effort and B (Proposition 2): there is a unique B_v that maximizes the innovation rate, with dv/dB &amp;gt; 0 for B &amp;lt; B_v and dv/dB &amp;lt; 0 for B_v &amp;lt; B &amp;lt; B_0 (the upper bound beyond which no R&amp;amp;D occurs). Explicitly, B_v = [lambda&lt;/em&gt;L / (rho*(theta-1))]^{1/theta}. Proposition 3 further establishes that in economies whose innovative capacity falls just below the threshold for positive growth at B=1, a well-chosen NIS can shift the economy from a zero-growth to a positive-growth steady state.&lt;/p&gt;
&lt;p&gt;The welfare-maximizing breadth B_w is shown to be unique, binding (B_w &amp;gt; 1), and strictly below B_v (Proposition 4 and 5). The welfare optimum trades off the dynamic gain from greater innovation against the static consumer surplus loss from higher markup power. Because the dynamic gain is still positive when B &amp;lt; B_v (R&amp;amp;D is still rising) but the static loss grows continuously in B, the welfare maximum necessarily occurs in the region where research is still increasing — i.e., B_w &amp;lt; B_v.&lt;/p&gt;
&lt;p&gt;Numerically, at baseline parameters (rho = 0.07, theta = 4, lambda&lt;em&gt;L = 1), B_w = 1.14 and the equilibrium R&amp;amp;D share is v(B_w) = 0.22, implying an asymptotic maximum real wage growth rate of 4.8%. The optimal breadth is most sensitive to theta (Pareto tail thickness) and less sensitive to rho and lambda&lt;/em&gt;L.&lt;/p&gt;
&lt;p&gt;When patent length (Omega) is added as a second instrument, the model yields a sharp result: the welfare-maximizing policy sets Omega → infinity together with B = B_w (Proposition 6). Unlike patent breadth, patent length has no hurdle effect — a longer patent duration raises R&amp;amp;D monotonically (dv/dOmega &amp;gt; 0, Lemma 2). With no diminishing returns to innovation effort in this model (the Poisson arrival rate is proportional to vL), the marginal dynamic gain from extending Omega always strictly outweighs the marginal static loss, so infinite patent length is always superior to any finite length. With Omega = 20 years (the TRIPS standard), the baseline calibration implies B_w = 1.13 and v(B_w) = 0.21 — only slightly below the infinite-length benchmark — suggesting the qualitative infinite-length result has limited quantitative bite for realistic patent durations.&lt;/p&gt;
&lt;p&gt;Proposition 7 shows that patent breadth and patent length are policy complements: when patent length is exogenously constrained to a finite value, the welfare-maximizing breadth increases in Omega (dB_w/dOmega &amp;gt; 0). Intuitively, a shorter patent duration weakens innovation incentives, so the optimal NIS compensates by providing stronger breadth protection.&lt;/p&gt;
&lt;p&gt;The paper provides a unified rationalization of several empirical puzzles: the weak or negative relationship between patent strength and innovation rates (Sakakibara-Branstetter 2001 on Japan; Bessen-Maskin 2009 on US software) is consistent with B being set above B_v, where the hurdle effect dominates; the causal evidence in Galasso-Schankerman (2014) that patents impede cumulative knowledge accumulation is consistent with the hurdle effect operating at the margin.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-is-this-a-theoretical-or-empirical-paper"&gt;Q1. What is the identification strategy, and is this a theoretical or empirical paper?&lt;/h3&gt;
&lt;p&gt;This is a purely theoretical paper. There is no empirical identification strategy. The core contribution is an analytically tractable general equilibrium model in which the key results (Propositions 1–7) are derived from first-order conditions, comparative statics, and the application of the intermediate value theorem. The Pareto-improvement distribution is the key parametric assumption that enables closed-form expressions for welfare and the growth rate.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-key-model-departure-from-kortum-1997-and-eaton-kortum-2001"&gt;Q2. What is the key model departure from Kortum (1997) and Eaton-Kortum (2001)?&lt;/h3&gt;
&lt;p&gt;Kortum (1997) and Eaton-Kortum (2001) model ideas as drawn from a stationary distribution over productivity levels — new ideas may or may not surpass the existing frontier, and as ideas accumulate it becomes progressively less likely that a new draw beats the current best. This generates growth only if the workforce grows. Chor and Lai instead model productivity improvements (ratios Z_{k+1}/Z_k) as i.i.d. Pareto draws, so each new idea strictly improves on the frontier regardless of how many ideas have arrived. This cumulative structure generates endogenous growth with a constant workforce and introduces knowledge spillovers that are absent in Kortum (1997).&lt;/p&gt;
&lt;h3 id="q3-what-exactly-is-the-non-infringing-inventive-step-nis-and-how-does-it-differ-from-other-breadth-concepts-in-the-literature"&gt;Q3. What exactly is the &amp;rsquo;non-infringing inventive step&amp;rsquo; (NIS) and how does it differ from other breadth concepts in the literature?&lt;/h3&gt;
&lt;p&gt;The NIS requirement B stipulates that a new idea must achieve a productivity at least B times the productivity of the current best patent (i.e., Z_new &amp;gt;= B * Z_current) to be patentable and non-infringing (what the paper calls &amp;rsquo;leading breadth&amp;rsquo;). The paper notes this is distinct from — though related to — patentability requirements studied by O&amp;rsquo;Donoghue (1998), which focused on the minimum improvement to qualify for a new patent but not necessarily on infringement. It also differs from the Gilbert-Shapiro (1990) and Klemperer (1990) breadth concepts, which focus on horizontal product differentiation (consumer willingness to substitute away from a patent) rather than vertical quality improvements. In the paper&amp;rsquo;s model, both patentability and non-infringement requirements are captured by a single parameter B, with the simplifying assumption that meeting the B hurdle is both necessary and sufficient for non-infringement.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-three-externalities-in-the-model-and-which-one-drives-the-market-planner-wedge"&gt;Q4. What are the three externalities in the model, and which one drives the market-planner wedge?&lt;/h3&gt;
&lt;p&gt;Three externalities are present: (1) The intertemporal spillover effect — firms do not internalize that their innovation raises the knowledge base for future innovators. (2) The appropriability effect — firms capture only private profits, not the full consumer surplus gain from each innovation. (3) The business-stealing effect — each innovator imposes a negative externality on the incumbent patent-holder by eroding their profits. Effects (2) and (3) exactly offset each other in the Pareto specification, so only the intertemporal spillover effect remains. This is verified formally: the market equilibrium condition features a discount rate of rho + lambda&lt;em&gt;v&lt;/em&gt;L (including the creative destruction hazard), whereas the social planner&amp;rsquo;s problem involves only rho. The wedge between v_eqm and v_SP stems entirely from this higher effective discount rate in decentralized equilibrium.&lt;/p&gt;
&lt;h3 id="q5-why-is-the-welfare-maximizing-patent-breadth-strictly-less-than-the-innovation-rate-maximizing-breadth"&gt;Q5. Why is the welfare-maximizing patent breadth strictly less than the innovation-rate-maximizing breadth?&lt;/h3&gt;
&lt;p&gt;At B_v, research effort is at its maximum, but this is achieved by granting patent-holders maximum protection, imposing the largest static consumer surplus loss. For B between B_w and B_v, increasing B further raises the static loss but no longer raises the innovation rate significantly enough to compensate; in fact for B &amp;gt; B_v, research effort falls while the static loss remains. The welfare optimum trades off the dynamic benefit (higher innovation) against the static cost (monopoly pricing). Because welfare must also account for the static loss at each period, and this loss is already large at B_v, the welfare optimum is achieved at a lower level of protection. Formally, dU_0/dB &amp;lt; 0 for all B in [B_v, B_0), and the unique welfare maximum lies strictly in [1, B_v).&lt;/p&gt;
&lt;h3 id="q6-why-is-the-optimal-patent-length-infinite"&gt;Q6. Why is the optimal patent length infinite?&lt;/h3&gt;
&lt;p&gt;Unlike patent breadth, patent length has only a profit effect and no hurdle effect — a longer patent strictly raises R&amp;amp;D effort (Lemma 2). Moreover, the model has no diminishing returns to innovation effort: the Poisson arrival rate of ideas is simply proportional to the total number of R&amp;amp;D workers at each date (lambda&lt;em&gt;v&lt;/em&gt;L), so each additional unit of research labor generates the same expected innovation flow regardless of how much research has already been done. This means the marginal dynamic gain from raising Omega (via increased innovation) is approximately constant, while the marginal static loss (additional consumer surplus ceded per period) is also roughly constant. The dynamic gain always strictly exceeds the static loss as long as the economy can sustain positive R&amp;amp;D (Lemma 1 condition holds), so Omega → infinity is always welfare-improving. This result breaks down if one introduces diminishing returns to R&amp;amp;D (e.g., a fishing-out effect or a congestion externality in research).&lt;/p&gt;
&lt;h3 id="q7-are-patent-breadth-and-patent-length-policy-substitutes-or-complements"&gt;Q7. Are patent breadth and patent length policy substitutes or complements?&lt;/h3&gt;
&lt;p&gt;They are policy complements (Proposition 7): when patent length is shorter (e.g., exogenously constrained by TRIPS or ethical considerations), the welfare-maximizing breadth B_w is lower; conversely, a longer patent length calls for a higher optimal breadth. This is because a longer patent length increases the dynamic gain from research, which raises the marginal value of also increasing breadth (since breadth further amplifies the monopoly profit effect). Formally, d^2U^l_0/(dB d Omega) &amp;gt; 0 at B_w, implying dB_w/d Omega &amp;gt; 0 by the implicit function theorem.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-quantitative-calibration-and-what-are-the-key-numerical-results"&gt;Q8. What is the quantitative calibration, and what are the key numerical results?&lt;/h3&gt;
&lt;p&gt;The calibration is illustrative rather than structural. Baseline: rho = 0.07 (matching real stock market returns as in Kortum 1997), theta = 4 (implying expected profits = 25% of per-variety expenditure, since 1/(1+theta) = 0.20 &amp;hellip; actually 1/(1+4) = 0.20, with the text stating 1/(1+theta) = 0.25 implying theta=3; the paper states theta=4 gives 1/(1+theta) = 0.20 — there is a slight inconsistency in the text&amp;rsquo;s wording, but the stated result is 25% of expenditures per variety), lambda*L = 1 (one expected new idea per variety per year). These yield: B_w = 1.14 (infinite patent length), v(B_w) = 0.22 (22% of labor in R&amp;amp;D), and an asymptotic maximum real wage growth rate of 4.8%. The optimal breadth B_w is most sensitive to theta: lowering theta (fatter tail, larger average improvements) raises B_w substantially. Under a finite patent length of Omega = 20, the results change minimally: B_w = 1.13, v(B_w) = 0.21.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-model-handle-the-possibility-that-economies-with-low-innovative-capacity-might-not-innovate-at-all-without-policy"&gt;Q9. How does the model handle the possibility that economies with low innovative capacity might not innovate at all without policy?&lt;/h3&gt;
&lt;p&gt;When lambda&lt;em&gt;L &amp;lt; rho&lt;/em&gt;theta, the economy has no R&amp;amp;D in the decentralized equilibrium at B = 1 (v(1) &amp;lt; 0 per equation 22). However, Proposition 3 shows that if lambda&lt;em&gt;L falls in the intermediate range (rho&lt;/em&gt;(theta-1)&lt;em&gt;(theta^2/(theta^2-1))^theta &amp;lt; lambda&lt;/em&gt;L &amp;lt; rho*theta), there exists a range of binding NIS values B &amp;gt; 1 that can shift the economy from zero to positive growth. Setting B = B_v achieves this transition. This is because the profit effect of introducing a binding NIS can more than offset the hurdle effect in this regime, making it profitable for some workers to engage in R&amp;amp;D.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-key-welfare-improving-scope-conditions-for-the-nis-policy"&gt;Q10. What are the key welfare-improving scope conditions for the NIS policy?&lt;/h3&gt;
&lt;p&gt;The welfare gain from a binding NIS requires Assumption 1: lambda&lt;em&gt;L &amp;gt; rho&lt;/em&gt;theta. This ensures the economy already features positive R&amp;amp;D at B = 1, and that the innovative capacity is large enough so the dynamic gains from raising B above 1 exceed the static consumer surplus losses. Without this condition, the NIS may either fail to generate R&amp;amp;D (if lambda*L is very low) or may tip the economy into R&amp;amp;D via Proposition 3&amp;rsquo;s mechanism, but welfare-optimality of the NIS still requires the economy be in a regime where the profit effect dominates for small B. Additionally, the NIS must remain below B_v to generate any dynamic gain.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-model-relate-to-japans-narrow-patent-breadth-policy-from-1960-1993"&gt;Q11. How does the model relate to Japan&amp;rsquo;s narrow patent breadth policy from 1960-1993?&lt;/h3&gt;
&lt;p&gt;The paper cites Ordover (1991) and Maskus-McDaniel (1999) to note that Japan deliberately adopted narrow patent breadth to encourage more incremental innovation and technology catch-up. In the model&amp;rsquo;s terms, Japan was setting B close to 1 (or even at 1) to lower the hurdle for new patents, maximizing the number of patentable ideas. This is consistent with a strategy of maximizing the innovation rate (operating near B_v or even below it), potentially at the cost of some dynamic welfare optimization. The Apple v. Samsung example illustrates that the US tends toward broader patent breadth (higher B) than Japan, consistent with the model&amp;rsquo;s international variation in NIS standards.&lt;/p&gt;
&lt;h3 id="q12-how-does-the-paper-handle-the-price-markup-and-profit-structure-under-the-nis"&gt;Q12. How does the paper handle the price markup and profit structure under the NIS?&lt;/h3&gt;
&lt;p&gt;Under Bertrand competition with limit pricing, the incumbent with the best patentable technology sets price equal to the marginal cost of the second-best technology (the previous patent-holder). The price markup m = Z_k/Z_{k-1} is drawn from a Pareto distribution with shape theta and lower bound 1 (no NIS) or B (with NIS). Flow profits are therefore: Pi = B(1+theta)^{-theta} / [B(1+theta) - theta] &amp;hellip; more precisely from equation (19): Pi = [B(1+theta) - theta] * (B(1+theta))^{-1}. As B rises, Pi increases (higher average markups from higher minimum improvement), which is the profit effect. The expected log productivity of the k-th patentable idea is E[ln Z~_k] = k/theta + k*ln(B), confirming that higher B raises not just the probability threshold but also the expected productivity of successful innovations.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-limitations-and-potential-extensions-noted-by-the-authors"&gt;Q13. What are the limitations and potential extensions noted by the authors?&lt;/h3&gt;
&lt;p&gt;The authors acknowledge several limitations and propose extensions: (1) The model assumes fully cumulative innovation — each idea strictly builds on the frontier. Generalizing to partial cumulativeness (where some ideas are non-cumulative or only partially built on existing knowledge) is flagged as a natural extension. (2) The analysis is confined to a single-country setting. A multi-country extension would allow study of cross-border patent policy spillovers and optimal international IPR harmonization (e.g., under TRIPS). (3) The model does not allow directed research — firms cannot target specific varieties. Relaxing this could introduce additional policy margins. (4) The model abstracts from imitation threats, which Gallini (1992) shows can make broader patent protection optimal.&lt;/p&gt;
&lt;h3 id="q14-how-does-the-paper-compare-to-odonoghue-1998-and-odonoghue-zweimüller-2004"&gt;Q14. How does the paper compare to O&amp;rsquo;Donoghue (1998) and O&amp;rsquo;Donoghue-Zweimüller (2004)?&lt;/h3&gt;
&lt;p&gt;O&amp;rsquo;Donoghue (1998) shows a patentability requirement can raise social welfare in a partial equilibrium setting, and Hunt (2004) finds an inverted-U relationship between innovation rate and requirement strength — both echo Chor-Lai&amp;rsquo;s findings. O&amp;rsquo;Donoghue-Zweimüller (2004) embed patentability in a quality-ladder endogenous growth model but focus more on innovation effects than welfare. The contribution of Chor-Lai relative to these papers is: (i) a fully general equilibrium treatment with explicit welfare analysis; (ii) derivation of both the welfare-maximizing breadth and the innovation-maximizing breadth and proof that Bw &amp;lt; Bv; (iii) extension to jointly optimal patent breadth and length, showing infinite patent length is optimal; and (iv) the Pareto-Gamma tractability that yields closed-form expressions and enables clean comparative statics on three deep parameters.&lt;/p&gt;
&lt;h3 id="q15-what-robustness-checks-does-the-paper-provide"&gt;Q15. What robustness checks does the paper provide?&lt;/h3&gt;
&lt;p&gt;The paper notes in the main text that results are robust to removing the scale effect (the feature that the innovation rate increases in L). An online appendix (referenced but not included in this draft) proves that the main qualitative results — inverted-U in innovation vs. B, unique welfare-maximizing B_w &amp;lt; B_v, and infinite optimal patent length — survive in a model variant without the scale effect. The numerical sensitivity analysis in Section 3.4 also demonstrates robustness of the qualitative findings across wide ranges of rho (0.02 to 0.12) and theta (2 to 6) and lambda*L.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Non-Infringing Inventive Step (NIS) requirement&lt;/strong&gt;: A patent policy parameter B &amp;gt;= 1 stipulating that a new idea must achieve a productivity at least B times that of the current best patent to qualify for a patent and be deemed non-infringing. In the paper&amp;rsquo;s usage, this simultaneously captures both the patentability requirement and the leading breadth (protection of incumbents against near-imitation), and is used interchangeably with &amp;lsquo;patent breadth.&amp;rsquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cumulative innovation&lt;/strong&gt;: An innovation process in which each new idea strictly improves upon the existing technological frontier. Formally, the productivity improvement Z_{k+1}/Z_k is drawn i.i.d. from a Pareto distribution with support [1, infinity), so each arriving idea always delivers a strictly positive productivity gain over the current best technology. This contrasts with non-cumulative models (e.g., Kortum 1997) where draws are from a stationary distribution and may fall below the frontier.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Profit effect (of patent breadth)&lt;/strong&gt;: The mechanism by which a higher NIS requirement B reduces the hazard rate that an incumbent patent-holder is superseded (from lambda&lt;em&gt;v&lt;/em&gt;L to lambda&lt;em&gt;v&lt;/em&gt;L*B^{-theta}), thereby extending the expected duration of monopoly power and raising the value of each patent. This increases R&amp;amp;D incentives by raising expected profits from successful innovation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hurdle effect (of patent breadth)&lt;/strong&gt;: The mechanism by which a higher NIS requirement B reduces the probability that any given arriving idea is patentable (probability B^{-theta}), thereby lowering the expected return to engaging in R&amp;amp;D. This discourages research effort and is the force that eventually dominates when B becomes sufficiently large, causing the innovation rate to fall.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Innovative capacity&lt;/strong&gt;: The product lambda&lt;em&gt;L, where lambda is the per-worker Poisson arrival rate of ideas and L is the total labor endowment. All steady-state outcomes in the model depend on lambda and L only through this product, not their individual values. It is the key parameter determining whether positive R&amp;amp;D equilibrium exists (requires lambda&lt;/em&gt;L &amp;gt; rho*theta) and the magnitude of welfare gains from patent policy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intertemporal spillover externality&lt;/strong&gt;: The sole market failure driving under-investment in R&amp;amp;D in this model&amp;rsquo;s Pareto specification. Because the knowledge embodied in each marketed innovation diffuses freely and becomes the base for subsequent cumulative improvements, private innovators do not internalize the benefit their R&amp;amp;D confers on future innovators. This causes firms to use an effective discount rate of rho + lambda&lt;em&gt;v&lt;/em&gt;L (including the creative destruction hazard) rather than rho alone, leading to strictly less R&amp;amp;D than the social optimum. Appropriability and business-stealing externalities exactly cancel in this model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy complementarity (breadth and length)&lt;/strong&gt;: The property that the welfare-maximizing patent breadth B_w is increasing in patent length Omega: dB_w/d Omega &amp;gt; 0. When the patent authority is constrained to set a shorter patent length, the optimal breadth should also be narrower, and vice versa. This arises because a longer patent length raises the marginal dynamic benefit of providing stronger breadth protection.&lt;/p&gt;</description></item><item><title>Payment data, information disclosure, and privacy</title><link>https://macropaperwarehouse.com/papers/payment-data-information-disclosure-and-privacy/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/payment-data-information-disclosure-and-privacy/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Digital payments generate vast, high-frequency, transaction-level data that several central banks (Bank of Canada, Swiss National Bank, Eurosystem members) already use for nowcasting, and regulatory initiatives (the EU&amp;rsquo;s PSD2, the UK&amp;rsquo;s Open Banking Standard, prospective CBDCs) are broadening system-wide data access. The paper asks how improved aggregate-demand forecasts enabled by payment data affect economic activity and through which channels; what the optimal communication policy for disseminating such forecasts is and how it depends on the monetary-policy stance; whether a competitive market in which private banks produce and sell forecasts is socially optimal; and how privacy concerns over individual transaction data affect optimal policy.&lt;/p&gt;
&lt;p&gt;Model setup: The authors build a Lagos-Wright / Rocheteau-Wright general-equilibrium monetary model with infinitely-lived buyers and sellers (unit measure each) and periods split into a centralized market (CM) and decentralized market (DM). Each period a stochastic fraction theta_t of buyers becomes &amp;lsquo;active&amp;rsquo; and wants the DM good; theta_t takes two values, theta_B &amp;lt; theta_G (bad/good aggregate state) with unconditional mean E[theta_t] = theta-bar. Sellers can pay an effort cost kappa to raise productivity from theta_L to theta_H. Payments use bank deposits fully backed by one-period government bonds costing g &amp;gt; beta (g is the policy variable; r = 1/g - 1). DM terms of trade follow the Kalai (1977) bargaining solution with buyer bargaining power sigma. No agent observes theta_t directly, but aggregating payment data across all banks yields a noisy binary signal s in {o,p} (optimistic/pessimistic), producing an unbiased forecast theta-tilde_t in {theta-tilde_G, theta-tilde_B} with E(theta-tilde_t) = theta-bar.&lt;/p&gt;
&lt;p&gt;Main findings (qualitative, as the paper is theoretical with an illustrative calibration): Disclosing forecasts affects welfare through two channels. (1) Demand channel: buyers hold more deposits when expecting high demand, so disclosure raises deposit-holding volatility; even though buyer utility is strictly concave, aggregate welfare w(theta) can be convex or concave. The sign hinges on the statistic T(x) = [u&amp;rsquo;&amp;rsquo;(x)]^2 / [u&amp;rsquo;&amp;rsquo;&amp;rsquo;(x)(u&amp;rsquo;(x)-1/theta)]: w(theta) is convex if T(x) &amp;lt; 1/3 and concave if T(x) &amp;gt; 1 over the relevant range (Lemma 4). (2) Investment channel: sellers underinvest because they capture only fraction (1-sigma) of DM surplus, so disclosure that encourages (discourages) investment raises (lowers) welfare (Lemma 3, thresholds kappa_1 &amp;lt; kappa_2 &amp;lt; kappa_3). Crucially the welfare effect is state-dependent in the monetary stance: with a low bond price/high deposit rate (low g) disclosure tends to reduce welfare (it mainly adds downside volatility and can weaken investment), while with high g (low deposit rate) disclosure tends to raise welfare (Proposition 1; thresholds g, g-bar). Calibrating to the U.S. economy 2016-19, disclosure improves welfare when the utility curvature parameter gamma is small and g is large; the discrete investment channel is inactive over most of the parameter space (Figure 2).&lt;/p&gt;
&lt;p&gt;Policy/theoretical implications: A central bank that controls disclosure can do better than binary reveal/withhold by sending noisy messages, a form of Bayesian persuasion (Kamenica-Gentzkow 2011): by committing to send the pessimistic message mb even when the forecast is optimistic (P^b &amp;lt; 1), it raises the posterior theta-tilde_b and induces investment, improving welfare when the investment channel is strong (Figures 4-5; numerical cases g = 1.05 and g = 1.00). A competitive market where private banks pay fixed cost C to produce and sell the forecast yields zero profits and always reveals undistorted information; provided C is below a threshold C-bar the forecast is always produced and sold, possibly causing excessive information production relative to the social optimum. Privacy: a fraction eta of buyers with high privacy costs use cash, shrinking recorded transactions and lowering forecast precision, but this need not reduce welfare; concave privacy costs can make deposit buyers&amp;rsquo; preferences less concave, turning welfare convex so disclosure helps via the demand channel, partially but not fully offsetting the privacy cost.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-modelingidentification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the modeling/identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;This is a theoretical general-equilibrium paper, not an empirical identification exercise. The strategy is to embed payment-data-derived forecasting and central-bank communication into a Lagos-Wright/Rocheteau-Wright monetary search model. Aggregate demand theta_t is a two-state random variable realized at the start of the DM; agents make CM decisions (deposit holdings, investment) under a common prior theta-bar unless a forecast is disclosed. The &amp;rsquo;threat&amp;rsquo; analog is robustness of the comparative statics to functional-form and parameter assumptions; the authors discipline curvature via the statistic T(x) and use a CRRA-type utility u(x)=(x+gamma)^{1-sigma_u}&amp;hellip; so that conditions map cleanly into the parameter gamma. They acknowledge agents in reality observe many macro indicators, but assume the only payment-data-based information is the unbiased binary signal, to isolate the informational value of payment data.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-main-channels-and-how-are-they-distinguished"&gt;Q2. What are the two main channels, and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;The demand channel works through buyers&amp;rsquo; deposit holdings: optimistic forecasts raise deposits and DM consumption x, pessimistic forecasts lower them; its welfare sign depends on the convexity/concavity of w(theta), governed by T(x) (convex if T&amp;lt;1/3, concave if T&amp;gt;1). The investment channel works through sellers&amp;rsquo; discrete investment decision: because sellers capture only (1-sigma) of surplus they underinvest, so disclosure that pushes investment up raises welfare and disclosure that pushes it down lowers welfare. They are distinguished analytically by shutting one off: Lemma 4 and Proposition 1 set theta_L = theta_H to isolate the demand channel; Lemma 3 isolates the investment channel via the cost thresholds kappa_1 &amp;lt; kappa_2 &amp;lt; kappa_3.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-welfare-effect-depend-on-monetary-policy-stance"&gt;Q3. How does the welfare effect depend on monetary policy stance?&lt;/h3&gt;
&lt;p&gt;The bond price g (inverse of the deposit rate, r = 1/g - 1) is the key policy variable. When g is small (high deposit rate, cheap to hold deposits), consumption x is near its upper bound x*(theta) already under theta-bar, so an optimistic forecast barely raises x while a pessimistic one sharply lowers it, making welfare locally concave and disclosure welfare-reducing; low g also makes DM surplus large so sellers already invest, and a low theta-tilde_B can discourage investment, hurting welfare. When g is large (low deposit rate, costly deposits), x is low under theta-bar so an optimistic forecast substantially raises trade volume, making welfare convex and disclosure welfare-improving (Proposition 1, thresholds g and g-bar). Hence optimal forecast communication should be designed jointly with conventional monetary policy.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-bayesian-persuasion--noisy-message-result-work"&gt;Q4. How does the Bayesian persuasion / noisy-message result work?&lt;/h3&gt;
&lt;p&gt;Instead of fully revealing theta-tilde_t, the central bank sends messages m in {mg,mb} under a committed, publicly known policy phi, choosing posteriors P^g = P(theta-tilde_G|mg) and P^b = P(theta-tilde_B|mb). Lemma 6 gives the policy implementing constant posteriors (requires P^b + P^g != 1). By lowering P^b below 1, the bank sometimes sends mb even when the forecast is optimistic, raising the posterior theta-tilde_b conditional on mb and encouraging sellers to invest; this can outweigh the demand-channel loss when the investment channel is strong. Lowering P^g below 1 adds beneficial noise via the demand channel when w is concave (low g). Numerical exercises with g = 1.05 (welfare locally convex, full transparency P^g=P^b=1 optimal when only demand channel active) and g = 1.00 (welfare locally concave, noisy messages welfare-improving) illustrate this (Figures 4-5).&lt;/p&gt;
&lt;h3 id="q5-why-do-buyers-and-sellers-always-want-to-buy-the-forecast-even-when-disclosure-can-lower-welfare-and-what-is-the-market-failure"&gt;Q5. Why do buyers and sellers always want to buy the forecast even when disclosure can lower welfare, and what is the market failure?&lt;/h3&gt;
&lt;p&gt;Lemma 5 shows buyers&amp;rsquo; willingness to pay rho^b_t &amp;gt; 0 always and sellers&amp;rsquo; rho^s_t &amp;gt;= 0. Knowing theta-tilde_t lets buyers tailor deposit holdings (avoiding the cost of carrying a fixed level since g &amp;gt; beta) and lets sellers tailor investment, yielding strictly higher private surplus. But neither internalizes the social benefit (the increase in total DM surplus), so private willingness to pay can exceed the social value. Proposition 3 shows that for C &amp;lt;= C-bar the forecast is always produced and sold in the competitive equilibrium (banks earn zero profit), which can lead to excessive information production relative to the social optimum. The market always fully reveals; it cannot replicate the central bank&amp;rsquo;s optimal noisy (persuasion) policy.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-selective-disclosure-result"&gt;Q6. What is the selective-disclosure result?&lt;/h3&gt;
&lt;p&gt;When the production cost C is neither large nor small, the break-even price may exceed only one side&amp;rsquo;s willingness to pay, so the forecast is sold only to buyers or only to sellers (Proposition 3). A buyer-only outcome can improve welfare if the forecast helps via the demand channel but hurts via the investment channel; a seller-only outcome helps if the reverse holds. Online Appendix C.3 shows both are possible, but these market outcomes generally do not coincide with the social optimum, so implementing welfare-improving selective disclosure may require the central bank to control the payment data.&lt;/p&gt;
&lt;h3 id="q7-how-does-forecast-precision-affect-outcomes"&gt;Q7. How does forecast precision affect outcomes?&lt;/h3&gt;
&lt;p&gt;Raising phi_o (precision of the optimistic signal) requires lowering phi_p, sharpening the forecast under both realizations. Through the demand channel, dE[w]/dphi_o = phi-tilde(theta_G-theta_B)[w&amp;rsquo;(theta-tilde_G)-w&amp;rsquo;(theta-tilde_B)], which is positive when w is convex and negative when concave. Through the investment channel, more precision raises theta-tilde_G but lowers theta-tilde_B, which can raise or lower investment depending on kappa. With private banks, Proposition 4 shows buyers&amp;rsquo; and sellers&amp;rsquo; willingness to pay rises with precision, making production (and possible over-production) more likely and selective disclosure less likely. Under Bayesian persuasion, higher precision weakly raises welfare (it expands the feasible policy set); but if private banks also disseminate, the central bank&amp;rsquo;s persuasion is constrained because agents&amp;rsquo; posteriors cannot contain less information than the private forecast.&lt;/p&gt;
&lt;h3 id="q8-how-are-privacy-and-cash-modeled-and-what-is-the-effect-on-welfare"&gt;Q8. How are privacy and cash modeled, and what is the effect on welfare?&lt;/h3&gt;
&lt;p&gt;A fraction eta in (0,1) of buyers (&amp;lsquo;cash buyers&amp;rsquo;) face sufficiently large privacy costs from deposit-based payments and use lower-return cash; the rest (&amp;lsquo;deposit buyers&amp;rsquo;) prefer deposits. Cash use shrinks the share of recorded DM transactions, lowering forecast precision (unless cash and deposit buyers&amp;rsquo; demand is perfectly correlated). By the precision results this can raise or lower welfare; with private production it makes excessive information less likely, while under central-bank noisy-message disclosure lower precision shrinks the feasible policy set and can reduce welfare. If the privacy cost is increasing and concave in DM consumption x, deposit buyers&amp;rsquo; net DM utility becomes less concave, making w more likely convex, so disclosure can improve welfare via the demand channel and the optimal policy may switch from non-disclosure to disclosure. This partially but not fully offsets the negative welfare impact of the privacy cost.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-equilibrium-multiplicity-and-underinvestment-results-in-the-benchmark"&gt;Q9. What are the equilibrium-multiplicity and underinvestment results in the benchmark?&lt;/h3&gt;
&lt;p&gt;With no data sharing, all decisions are state-independent under theta-bar. Strategic complementarity (more sellers investing raises buyers&amp;rsquo; deposits, which raises investment payoff) can generate multiple stationary equilibria (lambda=0, lambda=1, and a mixed lambda in (0,1)) when kappa and theta-bar are intermediate (Figure 1). The lambda=1 equilibrium is highest-welfare and Pareto optimal, and the authors impose a refinement selecting it. Sellers can underinvest: there exists kappa for which lambda=0 is the unique equilibrium even though lambda=1 would be socially better, because sellers receive only (1-sigma) of DM surplus. This underinvestment drives the investment-channel welfare results.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-paper-relate-to-and-differ-from-closely-related-work"&gt;Q10. How does the paper relate to and differ from closely related work?&lt;/h3&gt;
&lt;p&gt;Versus Andolfatto-Berentsen-Waller (2014) and Andolfatto-Martin (2013), where assets pay stochastic dividends and information is disclosed at the start of the DM so nondisclosure is always optimal (consumption smoothing), here the forecast is revealed at the start of the CM and affects deposit and investment decisions, so disclosure can be welfare-positive or -negative. Versus Choi-Liang (2023), whose non-monotonic disclosure effects arise from a money-adoption coordination margin, here non-monotonicity arises from how disclosure shapes marginal deposit holdings and investment. It extends the payment-data literature (Garratt-van Oordt 2021; Garratt-Lee 2020; Kang 2024; Amendola-Araujo-Ferraris 2025; Wang 2020, 2023; Cheng-Izumi 2025; Ahnert-Hoffmann-Monnet 2024) by focusing on the macroeconomic forecasting value of payment data and optimal disclosure, and connects to central-bank communication work (Morris-Shin; Jarocinski-Karadi 2020 information channel; Aruoba-Drechsel forthcoming).&lt;/p&gt;
&lt;h3 id="q11-what-are-the-cbdc-and-privacy-protection-implications-and-their-scope-conditions"&gt;Q11. What are the CBDC and privacy-protection implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;CBDC can serve as an institutional alternative source of payment data: transactions are recorded on a digital ledger, potentially letting the central bank observe flows directly, and can reduce coverage gaps from financial exclusion (the paper cites the 2021 FDIC survey: 4.5 percent of U.S. households, about 5.9 million, were unbanked). CBDC data could improve welfare via the demand and investment channels. Because privacy is a primary public concern, the authors recommend privacy-preserving architectures: adding statistical noise (differential privacy), randomizing data on the buyer&amp;rsquo;s device before transmission, keeping data decentralized with only model updates shared (federated learning), and clear governance/consent. Scope condition: incentivizing a cash-to-deposit/CBDC shift is welfare-improving only under sufficient privacy protection and only under the conditions (e.g., concave privacy cost, high g) that make disclosure beneficial; legal hurdles to central-bank access of payment data remain, which CBDC issuance could circumvent.&lt;/p&gt;
&lt;h3 id="q12-what-extensions-and-robustness-checks-are-reported"&gt;Q12. What extensions and robustness checks are reported?&lt;/h3&gt;
&lt;p&gt;Correlated signals: the central bank and private banks may receive correlated but non-identical signals (e.g., the bank has confidential surveys); Online Appendix B.4 shows this does not change the main results because information affects allocations only through agents&amp;rsquo; beliefs about theta_t at decision time. The model is calibrated to the U.S. 2016-19 (Online Appendix B.2) for the quantitative figures. Online Appendix C.2 provides a continuous-investment version (under which the investment channel is always active and welfare responses are smoother); the paper deliberately presents the discrete-investment case to highlight the channels. Online Appendix C.1 gives additional noisy-message numerical exercises, and C.3 shows selective-disclosure cases. An alternative to the lambda=1 refinement is a government &amp;lsquo;revenue backstop&amp;rsquo; subsidy (Online Appendix B.2).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Demand channel&lt;/strong&gt;: The mechanism by which disclosing the aggregate-demand forecast changes buyers&amp;rsquo; deposit holdings and hence DM consumption volatility; its welfare sign depends on whether aggregate welfare w(theta) is convex or concave, governed by the curvature statistic T(x), not merely by the concavity of buyer utility.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Investment channel&lt;/strong&gt;: The mechanism by which disclosure changes sellers&amp;rsquo; discrete decision to invest in higher productivity; because sellers capture only fraction (1-sigma) of DM surplus they underinvest, so disclosure that encourages investment raises welfare and disclosure that discourages it lowers welfare.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;T(x) statistic&lt;/strong&gt;: A normalized log-curvature measure, T(x) = [u&amp;rsquo;&amp;rsquo;(x)]^2 / [u&amp;rsquo;&amp;rsquo;&amp;rsquo;(x)(u&amp;rsquo;(x)-1/theta)], that disciplines the curvature of w(theta): w is convex when T(x) &amp;lt; 1/3 and concave when T(x) &amp;gt; 1 over the relevant consumption range, capturing how quickly the marginal DM surplus falls as consumption rises.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bayesian persuasion via noisy messages&lt;/strong&gt;: In the paper&amp;rsquo;s sense, the central bank commits to a publicly known communication policy (choosing posteriors P^g and P^b) that deliberately garbles the forecast - e.g., sending the pessimistic message even when the forecast is optimistic - to shift agents&amp;rsquo; expectations (especially to induce socially efficient seller investment), exploiting that Bayes&amp;rsquo; rule constrains only the average posterior.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Excessive information production&lt;/strong&gt;: The outcome under a competitive market for forecasts where, because banks earn zero profit and both buyers and sellers are willing to pay for the forecast even though it may lower aggregate welfare, the forecast is always produced and sold whenever the cost C is below a threshold, over-supplying information relative to the social optimum.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cash buyers / privacy cost&lt;/strong&gt;: Buyers facing sufficiently large privacy costs from deposit-based (recorded) payments who choose lower-return cash; their use reduces recorded transactions and forecast precision, but a privacy cost that is concave in consumption can make deposit buyers&amp;rsquo; preferences less concave, turning welfare convex so that disclosure becomes optimal and partially offsets the privacy cost.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Aggregate state theta_t&lt;/strong&gt;: The two-valued (theta_B bad, theta_G good) random fraction of buyers who become active and demand the DM good, equal to the level of aggregate demand; realized at the start of the DM with unbiased forecast theta-tilde_t derived from aggregated payment data.&lt;/p&gt;</description></item><item><title>Population and Welfare: Measuring Growth when Life is Worth Living</title><link>https://macropaperwarehouse.com/papers/population-and-welfare-measuring-growth-when-life-is-worth-living/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/population-and-welfare-measuring-growth-when-life-is-worth-living/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;The paper asks how much economic progress looks different when one applies a total utilitarian welfare criterion — counting every person&amp;rsquo;s flow utility — rather than the standard per-capita consumption measure. The motivation is both philosophical and practical: philosophers have long debated whether the number of people matters for social welfare alongside average living standards, yet growth economists have almost exclusively used the per-capita approach. The authors do not adjudicate the debate; they quantify its stakes across a broad cross-country sample.&lt;/p&gt;
&lt;p&gt;The framework is parsimonious. Under total utilitarianism, flow social welfare is W = N·u(c). Consumption-equivalent (CE) welfare growth is gλ = v(c)·gN + gc, where gN is population growth, gc is per-capita consumption growth, and v(c) = u(c)/[u&amp;rsquo;(c)·c] is the value of a year of life in units of per-capita consumption. Diminishing marginal utility guarantees v(c) &amp;gt; 1: each percentage point of population growth is worth more than a percentage point of per-capita consumption growth. The baseline utility function is u(c) = ū + log(c). The key parameter ū is calibrated to the U.S. Environmental Protection Agency&amp;rsquo;s Value of Statistical Life (VSL) of $7.4 million (2006 prices): dividing by remaining life expectancy (~40 years) and U.S. per-capita consumption of $38,000 gives v(c_US,2006) ≈ 4.87. Normalizing U.S. 2006 consumption to 1 sets ū = 4.87. Under log utility, v(c) = ū + log(c) rises with living standards: it averaged roughly 2 in 1820 for the U.S. and nearly 5 by 2019, and ranges from about 2 (Ethiopia) to 5 (U.S.) across countries in 2019. The world-sample average of v(c) over 1960–2019 is 2.7.&lt;/p&gt;
&lt;p&gt;Applying the formula to Penn World Table 10.0 data for 101 countries over 1960–2019 yields the following main findings. CE welfare growth averages 6.2% per year versus 2.1% per year for per-capita consumption growth; at 2.1% growth per-capita consumption doubles every 33 years, but under the CE measure social welfare doubles every 12 years. Population growth (averaging 1.8% per year) accounts for 66% of CE welfare growth unweighted across countries, and 51% weighting by country population (which gives China a large weight). For the United States specifically, CE welfare growth averages 6.5% per year versus 2.2% for per-capita consumption growth. Country rankings shift dramatically. Mexico rises from the 35th to the 88th percentile (CE welfare growth: 8.6% per year; population contribution: 79%). South Africa and Kenya similarly move up sharply. Germany falls to the 11th percentile, Japan to the 32nd, and China to the 44th — all below the United States. The cross-country correlation between CE welfare growth and per-capita consumption growth is 0.51; with population growth, 0.29. Over the very long run (1500–2018, Maddison data), per-capita consumption rose 20-fold (0.6% per year) while CE welfare rose 3,700-fold (1.6% per year) due to population growing at 0.5% per year scaled by v(c).&lt;/p&gt;
&lt;p&gt;Robustness checks confirm the core result. Halving the baseline VSL (setting ū = 2.4) still leaves population contributing 38% of CE welfare growth on average. Incorporating within-country consumption inequality under a log-normal distribution lowers CE welfare growth by an average of just 10 basis points (from 6.1% to 6.0% for 1980–2007). Attributing migrants to source rather than destination countries produces a correlation of 0.92 between adjusted and baseline CE welfare growth rates. Decomposing population growth, roughly three-quarters of actual population growth in a 24-country subsample reflected increases in the number of lives lived (i.e., births), not longevity extension — so the welfare contribution of births exceeds that of rising longevity.&lt;/p&gt;
&lt;p&gt;An extended model adds leisure, parental altruism toward children&amp;rsquo;s consumption and human capital, and endogenous fertility. Using time-use data from six countries (U.S. 2003–2019; Netherlands 1975–2006; Japan 1991–2016; South Korea 1999–2019; Mexico 2006–2019; South Africa 2000–2010), the extension modestly reduces the population share of CE welfare growth in most countries. The main reason is that parental altruism &amp;ldquo;double-counts&amp;rdquo; children&amp;rsquo;s consumption in the social welfare function, making consumption growth relatively more valuable and thus scaling down the weight on population growth. Rising quality of children (human capital) roughly offsets falling fertility in most countries, leaving net CE welfare growth little changed. Mexico is the sharpest exception: under extended preferences, CE welfare growth falls from 6.5% to 3.3% because of sharply declining leisure and little offsetting gain in children&amp;rsquo;s quality. Japan and South Korea also see smaller population shares under the extended model. The qualitative conclusion — that population growth is a major contributor to CE welfare growth — survives across all specifications.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-identification-strategy-and-what-does-it-rely-on"&gt;Q1. What is the paper&amp;rsquo;s identification strategy and what does it rely on?&lt;/h3&gt;
&lt;p&gt;This is a welfare accounting exercise rather than a causal identification exercise. There is no identification problem in the traditional econometric sense: the authors are computing a welfare index given a social welfare function and observed data on population and consumption. The two key inputs are (1) data on population and consumption per capita from the Penn World Table 10.0 for 101 countries over 1960–2019, and (2) a calibrated value of the parameter ū, which is the value of a year of life measured in units of per-capita consumption. The calibration of ū is anchored to external VSL estimates (EPA&amp;rsquo;s $7.4 million in 2006 prices), divided by life expectancy and per-capita consumption. The paper is explicit that it cannot make causal policy recommendations because it says nothing about the production side of the economy or externalities (pollution, ideas, human capital spillovers).&lt;/p&gt;
&lt;h3 id="q2-what-is-vc-and-why-does-it-matter-so-much-for-the-results"&gt;Q2. What is v(c) and why does it matter so much for the results?&lt;/h3&gt;
&lt;p&gt;v(c) = u(c)/[u&amp;rsquo;(c)·c] is the value of a year of life measured in consumption-equivalent units — specifically, how many years&amp;rsquo; worth of per-capita consumption an individual would require as compensation for losing one year of life. Under log utility u(c) = ū + log(c), v(c) = ū + log(c), so it rises with the log of consumption. The key implication is that each percentage point of population growth is worth v(c) percentage points of per-capita consumption growth. Since v(c) empirically ranges from about 2 (Ethiopia) to 5 (rich countries), and averages 2.7 over 1960–2019 across 101 countries, population growth receives a substantial weight in the CE welfare measure. Without this amplification (i.e., if v = 1 so CE welfare equals aggregate consumption growth), population would still account for 36% of all growth.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-distinguish-its-approach-from-simply-using-aggregate-total-consumption-growth"&gt;Q3. How does the paper distinguish its approach from simply using aggregate (total) consumption growth?&lt;/h3&gt;
&lt;p&gt;Using aggregate consumption growth is equivalent to setting v(c) = 1 in the CE welfare formula — that is, weighting population growth and consumption growth equally. The paper shows that, under a total utilitarian welfare function with diminishing marginal utility, the correct weight on population growth is v(c) &amp;gt; 1, not 1. So aggregate consumption growth systematically understates the contribution of population growth to welfare: in a country with average v(c) = 2.7, a percentage point of population growth should receive 2.7 times the weight of a percentage point of consumption growth, not equal weight.&lt;/p&gt;
&lt;h3 id="q4-what-threats-to-the-baseline-calibration-of-vc-does-the-paper-address"&gt;Q4. What threats to the baseline calibration of v(c) does the paper address?&lt;/h3&gt;
&lt;p&gt;The paper addresses four main threats. First, VSL uncertainty: it considers halving and raising the baseline VSL by 50%, yielding ū = 2.4 and ū = 7.3 respectively. Population&amp;rsquo;s share of CE welfare growth remains 38% even under the low VSL. Second, functional form: it considers CRRA utility with risk-aversion γ = 2 rather than log (γ = 1), which lowers the population share to 40% (from 53% baseline, population-weighted). Third, whether v(c) should be constant rather than income-varying: rows 7–9 of Table 3 test constant v = 4.87, v = 2.7, and v = 1. Even v = 1 (aggregate consumption growth) gives population a 36% share. Fourth, whether the marginal VSL used to calibrate the model overstates the average value of a birth (since a birth produces a new life from the start, not an added year for a middle-aged person). The paper acknowledges this concern but treats the calibration as a natural baseline and explores lower ū as a robustness check.&lt;/p&gt;
&lt;h3 id="q5-how-does-within-country-inequality-affect-the-results"&gt;Q5. How does within-country inequality affect the results?&lt;/h3&gt;
&lt;p&gt;Under log utility and a log-normal distribution of individual consumption, CE welfare growth becomes gλ = [ū + log(c_t) - (1/2)σ²_t]·gN + gc - σ²_t·gσ, where σ² is the cross-sectional variance of log consumption. Inequality enters in two ways: it reduces the weight on population growth (because average utility is lower than utility of average consumption under concavity), and increases in inequality directly reduce CE welfare growth. Implementing this for 90 countries over 1980–2007, the mean adjustment is -10 basis points (6.1% to 6.0%), with a mean absolute deviation of 18 basis points. The adjustment is sizable for South Africa (−0.83 pp, due to very high inequality relative to U.S. 2006 baseline) and small or positive for Brazil (falling inequality over the period) and Ethiopia.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-treat-migration-and-does-it-matter"&gt;Q6. How does the paper treat migration, and does it matter?&lt;/h3&gt;
&lt;p&gt;The baseline credits population growth to the country of residence. The migration-adjusted measure reassigns migrants to their country of birth: it adds the flow utility of out-migrants (at destination-country consumption levels) and subtracts the flow utility of in-migrants (at destination-country consumption levels) from each country&amp;rsquo;s welfare. Using the World Bank Global Bilateral Migration Database for 81 countries over 1960–2000, migration-adjusted and baseline CE welfare growth rates have a correlation of 0.92. The adjustment matters most for specific countries — it raises welfare growth for net out-migrant countries like Mexico and the Philippines (since their emigrants consume more abroad) and lowers it for net in-migrant countries — but it does not alter the broad conclusion that population growth matters greatly.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-decomposition-of-population-growth-into-fertility-and-longevity-effects"&gt;Q7. What is the decomposition of population growth into fertility and longevity effects?&lt;/h3&gt;
&lt;p&gt;For a 24-country subsample (from the Human Mortality Database combined with World Bank migration data), the authors compute counterfactual population growth holding age-specific death rates constant at their initial-period values. Population-weighted, actual annual population growth is 0.72% versus a counterfactual of 0.53% with fixed longevity. So roughly three-quarters of population growth (and therefore three-quarters of the CE welfare contribution of population growth) reflected an increase in the number of lives lived (births minus deaths under fixed mortality), not gains in longevity. Italy and Japan are outliers: falling death rates (i.e., longevity gains) account for about three-quarters of their population growth. For context, Jones and Klenow (2016) attribute ~1% per year of CE welfare growth to rising longevity for 128 countries over 1980–2007; the total population growth contribution here (~3% per year, population-weighted from Table 1) substantially exceeds the longevity-only benchmark.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-extended-model-with-parental-preferences-add-and-what-are-its-main-results"&gt;Q8. What does the extended model with parental preferences add, and what are its main results?&lt;/h3&gt;
&lt;p&gt;The extended model incorporates adult leisure, parental altruism toward children&amp;rsquo;s consumption and human capital, endogenous fertility, and children&amp;rsquo;s utility as separate welfare contributors. Social welfare is W = N_p·u(c_p, l, c_k, h_k, b) + N_k·ũ(c_k), where b is fertility per adult, l is adult leisure, and h_k is children&amp;rsquo;s human capital. CE welfare growth is computed using first-order conditions from parents&amp;rsquo; utility maximization — specifically, the MRS between leisure/fertility/human capital and consumption can be measured from time-use data, which provides the welfare weights on each term. Key parameters: parental altruism weight α = 2/3 (calibrated to USDA household spending data), diminishing-returns-to-fertility parameter θ = 0.8, and children&amp;rsquo;s human capital elasticity η = 0.21 (from Mincer estimates in Lee, Roys, and Seshadri 2024). Main results: (1) Population growth remains an important contributor to CE welfare growth in most countries. (2) The population share falls somewhat because parental altruism double-counts children&amp;rsquo;s consumption, raising the relative weight on consumption growth. (3) Rising children&amp;rsquo;s quality (human capital, measured via real wage growth) roughly offsets falling fertility in most countries. (4) Mexico is the main exception: CE welfare growth drops from 6.5% to 3.3% due to falling leisure and little offset from rising children&amp;rsquo;s quality. (5) Japan&amp;rsquo;s population share falls further, turning slightly negative in some specifications.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-philosophical-foundation-and-what-is-the-repugnant-conclusion-objection"&gt;Q9. What is the philosophical foundation and what is the &amp;lsquo;repugnant conclusion&amp;rsquo; objection?&lt;/h3&gt;
&lt;p&gt;The total utilitarian social welfare function W = N·u(c) follows from three axioms: same-number Pareto (welfare ordering respects Pareto improvements for fixed populations), non-anti-egalitarianism (society does not prefer inequality), and mere addition (adding a person who values living, holding others&amp;rsquo; utilities constant, does not reduce welfare). These axioms, as surveyed by Kuruc, Budolfson, and Spears (2022), together imply total utilitarianism and rule out diminishing-returns-to-population approaches (e.g., W = N^α·u(c) for α &amp;lt; 1). The repugnant conclusion (Parfit 1984) holds that total utilitarianism could justify very large populations of people whose lives are barely worth living. The authors respond that their calculations are local — reflecting only actual births and deaths over 1960–2019 — not arbitrary expansions. They also note that 29 philosophers and economists (Zuber et al. 2021) have argued the repugnant conclusion is not a reason to reject totalism. The per-capita approach has its own problems: it implies one should remove people whose utility is valuable but below average (and implies the &amp;lsquo;sadistic conclusion&amp;rsquo; under certain conditions).&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-jones-and-klenow-2016"&gt;Q10. How does this paper relate to Jones and Klenow (2016)?&lt;/h3&gt;
&lt;p&gt;Jones and Klenow (2016) is the closest predecessor. That paper computes CE welfare measures incorporating consumption, leisure, life expectancy, and inequality, but in a per-capita framework — it measures individual living standards, not aggregate social welfare. The key difference here is moving from per-capita utility to total utilitarian welfare by multiplying individual utility by population, which introduces the v(c)·gN term. The current paper&amp;rsquo;s baseline is also simpler (consumption only) with an extended version that adds leisure and parental preferences. Jones and Klenow attribute ~1% per year of CE welfare growth to rising longevity for 128 countries over 1980–2007; the present paper shows total population growth (birth + longevity channels combined) contributes ~3% per year (population-weighted), substantially more.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper explicitly states it cannot make policy recommendations because it says nothing about the production side of the economy or about externalities (pollution, ideas externalities, human capital spillovers). Whether fertility rates are &amp;rsquo;too low&amp;rsquo; or the demographic transition raised or reduced social welfare requires estimating these externalities, which is beyond the paper&amp;rsquo;s scope. The paper is a measurement exercise, not an optimal policy analysis. Nonetheless, the results have implications for policy questions that depend on which welfare criterion is adopted: optimal fertility policy, the welfare cost of HIV/AIDS or other mortality shocks, the assessment of China&amp;rsquo;s One Child Policy, the welfare calculus of climate change mitigation, and the social returns to nonrival knowledge (which benefit a larger future population under totalism). The scope condition throughout is that the paper evaluates actual births and deaths over a historical period; the results do not directly speak to the desirability of population expansion beyond what occurred.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-main-robustness-checks-run-and-what-do-they-show"&gt;Q12. What are the main robustness checks run and what do they show?&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;VSL calibration: halving (ū = 2.4) or raising by 50% (ū = 7.3) the baseline VSL — population share falls to 38% or rises to higher levels, but population remains important in all cases. 2. CRRA utility with γ = 2 (more concave): population share falls to 40% population-weighted (from 53%). 3. Constant v(c): results with v = 4.87 (U.S. 2006 level), v = 2.7 (world average), and v = 1 (aggregate consumption growth) all confirm that population growth matters, with v = 1 still giving a 36% population share. 4. Inequality: mean absolute adjustment of 18 basis points; largest adjustment for South Africa (−0.83 pp). 5. Migration: correlation 0.92 between adjusted and baseline. 6. Birth vs. longevity decomposition: ~75% of population growth (population-weighted) is from net new lives, not longevity. 7. Extended preferences (time-use data): qualitative results survive; population share falls modestly except for Japan, South Korea, and Mexico.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="q13-what-heterogeneity-across-countries-and-time-periods-is-documented"&gt;Q13. What heterogeneity across countries and time periods is documented?&lt;/h3&gt;
&lt;p&gt;Cross-country: CE welfare growth ranges from just above 2% per year for the slowest-growing countries to more than 10% per year for the fastest. The correlation between CE welfare growth and per-capita consumption growth is 0.51; with population growth, 0.29. The value of v(c) ranges from about 2 (Ethiopia) to 5 (U.S.) in 2019, tracking consumption levels. Countries with high population growth (Mexico, Brazil, South Africa, Kenya, Sub-Saharan Africa more broadly) move up sharply in the growth rankings; countries with slow population growth (Germany, Japan, China) fall sharply. Within time: Japan shows CE welfare growth falling from 9.7% per year in the 1960s to −0.3% in the 2010s as both consumption growth and population growth slowed and then turned negative. China&amp;rsquo;s CE welfare growth fell more modestly from a 7.0% peak in the 1990s to 5.7% in the 2010s because rising v(c) partly offset slower population growth. Sub-Saharan Africa maintained stable population growth (~2.5% per year across all decades) and saw rising consumption in the 2000s and 2010s, producing CE welfare growth above 8% in the 2010s. Extended-model results (six-country sample): Mexico is a major outlier with falling leisure driving CE welfare growth down to 3.3% (from 6.5% baseline); Japan and South Korea have very small or near-zero population shares under extended preferences.&lt;/p&gt;
&lt;h3 id="q14-what-does-the-very-long-run-historical-exercise-show"&gt;Q14. What does the very long-run historical exercise show?&lt;/h3&gt;
&lt;p&gt;Using Maddison Project data (de Pleijt and van Zanden 2020) from 1500 to 2018 for the world as a whole, per-capita consumption rose by a factor of 20 (0.6% per year) and aggregate consumption rose by a factor of 163 (1.1% per year). CE welfare rose by a factor of 3,700 — a 1.6% per year average annual growth rate. The power of compounding over 500 years causes a difference of only 1 percentage point per year between CE welfare growth (1.6%) and per-capita consumption growth (0.6%) to produce a ratio of 185:1 in cumulative outcomes (3,700-fold versus 20-fold). Population growth accounts for 61% of CE welfare growth over this very long run.&lt;/p&gt;
&lt;h3 id="q15-what-data-sources-are-used-and-what-are-the-sample-restrictions"&gt;Q15. What data sources are used and what are the sample restrictions?&lt;/h3&gt;
&lt;p&gt;Baseline: Penn World Table 10.0 (Feenstra, Inklaar, and Timmer 2015) for 101 countries over 1960–2019 (starting from 111 countries and dropping those flagged as outliers in any year). Consumption is private plus government consumption. Inequality data: Jones and Klenow (2016), 90 countries, 1980–2007. Migration data: World Bank Global Bilateral Migration Database (1960, 1970, 1980, 1990, 2000), 81 countries. Birth/death decomposition: Human Mortality Database combined with World Bank migration data, 24 countries. Long run: Maddison Project (de Pleijt and van Zanden 2020), 1500–2018, with consumption proxied as 0.8 times per-capita GDP. Extended model: time-use surveys for U.S. (2003–2019), Netherlands (1975–2006), Japan (1991–2016), South Korea (1999–2019), Mexico (2006–2019), South Africa (2000–2010); Penn World Table for population, consumption, and hours worked; World Bank for number of children (0–19 years); USDA spending-on-children data (Lino 2011) for parental altruism calibration; Lee, Roys, and Seshadri (2024) Mincer estimates for human capital elasticity.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Consumption-Equivalent (CE) Welfare Growth&lt;/strong&gt;: The rate at which per-capita consumption would need to grow, holding population constant, to produce the same increase in total utilitarian social welfare as the observed combination of population growth and per-capita consumption growth. Formally gλ = v(c)·gN + gc. It is analogous to GDP in being a flow measure (not a present-discounted sum across generations).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Value of a Year of Life, v(c)&lt;/strong&gt;: The ratio u(c)/[u&amp;rsquo;(c)·c], equal to individual utility divided by the marginal utility of consumption times consumption. It converts the value of being alive for one year into consumption-equivalent units. Under log utility u(c) = ū + log(c), v(c) = ū + log(c), so it rises with living standards. It is calibrated from Value of Statistical Life estimates: v(c_US,2006) ≈ 4.87, meaning a year of American life in 2006 was worth approximately 4.87 years&amp;rsquo; worth of per-capita consumption.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Total Utilitarian Social Welfare Function&lt;/strong&gt;: W = N·u(c): social welfare is the sum of all individuals&amp;rsquo; flow utilities. It treats every person&amp;rsquo;s utility symmetrically and linearly in population, so adding a person who values living always increases welfare. This contrasts with the per-capita (average utilitarian) approach (which implicitly sets the welfare weight on population to zero) and intermediate approaches that weight population with diminishing returns (W = N^α·u(c), α &amp;lt; 1).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mere Addition (Axiom)&lt;/strong&gt;: One of three axioms (with same-number Pareto and non-anti-egalitarianism) whose conjunction implies the total utilitarian social welfare function for variable populations. It states that, holding the utilities of existing persons constant, adding a new person who values living does not reduce social welfare. The axiom directly rules out the average (per-capita) utilitarian criterion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Repugnant Conclusion&lt;/strong&gt;: Parfit&amp;rsquo;s (1984) critique of total utilitarianism: maximizing the sum of utilities could in principle justify an extremely large population of people whose lives are just barely worth living (positive but tiny utility), since total utility could exceed that of a smaller population with high per-capita welfare. The paper responds that its welfare calculations are local (reflecting actual historical births and deaths), not global maximization exercises, and cites the Zuber et al. (2021) consensus that this is not a decisive objection.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parental Altruism Weight (α, θ)&lt;/strong&gt;: Parameters governing how parents value children&amp;rsquo;s consumption relative to their own in the extended model. Under Assumption 1, u includes the term αb^θ·log(c_k): α governs the overall weight on children&amp;rsquo;s consumption and θ governs diminishing returns as the number of children rises. Calibrated to α = 2/3 (from USDA household spending ratios with two children) and θ = 0.8 (from cross-family variation in per-child spending). Parental altruism causes double-counting of children&amp;rsquo;s consumption in the social welfare function, upweighting consumption growth and reducing the relative contribution of population growth to CE welfare.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Double-Counting of Children&amp;rsquo;s Consumption&lt;/strong&gt;: When parents are altruistic, their utility depends on children&amp;rsquo;s consumption as well as their own; and children receive direct utility from their consumption too. In the CE welfare growth formula, this means a rise in children&amp;rsquo;s consumption raises welfare through two channels simultaneously (parental and child utility), so each unit of consumption growth is &amp;lsquo;worth more&amp;rsquo; relative to population growth. This is why the extended model&amp;rsquo;s population term is smaller than the baseline&amp;rsquo;s: consumption growth is valued more heavily under parental altruism, scaling down the consumption-equivalent weight on population growth.&lt;/p&gt;</description></item><item><title>Price Setting and Volatility: Evidence from Oil Price Volatility Shocks</title><link>https://macropaperwarehouse.com/papers/price-setting-and-volatility-evidence-from-oil-price-volatility-shocks/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/price-setting-and-volatility-evidence-from-oil-price-volatility-shocks/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether increases in aggregate volatility reduce the effectiveness of monetary policy by making aggregate prices more flexible. The motivation is concrete: policymakers worry that during episodes of high volatility, prices may become more synchronized in their adjustment, reducing monetary non-neutrality and limiting the ability of nominal stimulus to raise real output.&lt;/p&gt;
&lt;p&gt;The empirical strategy exploits variation in oil price volatility as a plausibly exogenous source of aggregate cost volatility. Oil price volatility is measured using a stochastic volatility model estimated on monthly WTI spot prices from 1986 to 2014 (Bayesian MCMC with particle filter). The key identification device is a Bartik-style interaction: an industry&amp;rsquo;s pre-determined oil input share (from the 1997 Input-Output Use Table, expressed as oil spending relative to value added) is interacted with the time-varying aggregate oil price volatility. Industries more dependent on oil should respond more strongly to oil price volatility shocks, while the time fixed effects absorb any aggregate confounders. The micro-price data are confidential item-level Producer Price Index records from the BLS covering 81 four-digit NAICS manufacturing industries from January 1998 to December 2014, with roughly 100,000 prices collected monthly from about 25,000 reporters.&lt;/p&gt;
&lt;p&gt;Two price-setting moments are the main outcomes: price change frequency (fraction of items with non-zero price change within an industry-month) and price change dispersion (standard deviation of non-zero price changes within an industry-month).&lt;/p&gt;
&lt;p&gt;The main empirical findings, from Table 6 (industry-specific oil demand variable regressions with both industry and time fixed effects):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;A one standard deviation increase in oil price volatility raises price change dispersion by approximately 2 percent relative to the mean for a 90th-percentile oil-share industry relative to a 10th-percentile oil-share industry (coefficient of 4.511, significant at 1 percent). This finding is robust to alternative oil price series (WTI, Brent, RAC), alternative volatility measures (stochastic volatility, GARCH, realized volatility), exclusion of the 2008 crisis period, and alternative dispersion measures (interquartile range).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The same cross-industry comparison shows that a one standard deviation increase in oil price volatility reduces price change frequency by approximately 1 percent relative to the mean for high-oil versus low-oil industries (coefficient of -2.486, significant at 5 percent in Table 6 column 1). This negative frequency result holds inside and outside the financial crisis period.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The time-series correlation between price change dispersion and oil price volatility for the top-10-percent oil-share industries is 0.45, versus only 0.08 for the bottom-10-percent industries, previewing the cross-sectional identification.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These findings contrast sharply with what the literature documents for idiosyncratic volatility (Vavra 2014), where both frequency and dispersion rise together. For aggregate (oil) volatility, dispersion rises but frequency does not, implying a different mechanism.&lt;/p&gt;
&lt;p&gt;To interpret these facts, the paper constructs and calibrates a general equilibrium state-dependent pricing model. Firms produce using labor and oil (Cobb-Douglas), face menu costs, and receive idiosyncratic productivity shocks with leptokurtic draws. The key modeling choice is random menu costs (drawn each period from a non-degenerate distribution, following Dotsey, King, and Wolman 1999 and Luo and Villar 2020) rather than fixed menu costs. With random menu costs, the selection of which prices adjust is attenuated relative to the common shock: many firms will not adjust because they drew a high menu cost regardless of the oil shock, keeping the mix of adjusting prices more disperse. A fixed-menu-cost model (Appendix A.3) produces a counterfactual negative relationship between oil price volatility and price change dispersion, because the strong selection effect causes prices to bunch in the direction of the cost shock.&lt;/p&gt;
&lt;p&gt;The calibrated one-sector random menu cost model matches the positive empirical link between oil price volatility and dispersion, with a muted frequency response. The multisector model (eight sectors calibrated to oil-share octiles of PPI industries) is fed the actual observed oil price and volatility series from 1998 to 2014, and the regression run on model-generated data matches the empirical coefficient on dispersion within one standard error of the data estimate (model: 3.876 versus data: 4.511). The model cannot replicate the empirical negative frequency response.&lt;/p&gt;
&lt;p&gt;The key quantitative implication for monetary policy: in the multisector model, a permanent increase in log nominal output of 0.002 (doubling one month&amp;rsquo;s growth rate) translates 59.1 percent into real output at baseline oil price volatility, and 58.8 percent after a one standard deviation increase in oil price volatility. The ability of nominal stimulus to raise consumption on impact falls by only 0.5 percent. The average decline across the full historical distribution of oil price volatility (1998-2014) is 1 percent lower at peak volatility (e.g. 2009) than at trough volatility (e.g. 2013). Supporting aggregate evidence using state-dependent local projections with Romer-Romer monetary shocks (1974-2007) confirms that the price level response to identified monetary shocks is not significantly different across high and low oil price volatility states.&lt;/p&gt;
&lt;p&gt;The policy implication is direct: the output-inflation tradeoff is nearly time-invariant with respect to aggregate volatility. Policymakers who respond to periods of high aggregate volatility by increasing nominal stimulus under the belief that policy effectiveness has declined would be overreacting and would generate unnecessary inflation. The source of volatility — aggregate versus idiosyncratic — matters critically for the price-setting implications and thus for the correct policy response.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The strategy is a Bartik-style interaction: each industry&amp;rsquo;s pre-determined oil input share (oil spending as a share of value added, from the 1997 Input-Output tables, before the sample period) is interacted with aggregate time-varying oil price volatility. Industry fixed effects absorb time-invariant heterogeneity; time fixed effects absorb all aggregate shocks common to all industries in a given month. Identification of the oil price volatility effect is thus from within-industry variation over time, scaling by the pre-existing oil dependence. The main threats are: (1) the interaction term could be correlated with unobserved shocks that are industry-specific and vary with oil price volatility; (2) oil prices could respond to aggregate U.S. economic conditions, threatening exogeneity. The paper defends against (2) by arguing that large oil price movements over the sample can be traced to external events (Middle East conflicts, Venezuelan oil strike, Asian demand expansion, Libyan uprising) rather than U.S. conditions, and that individual industries are price takers in the global oil market. For (1), the paper adds controls for industrial production growth, industry inflation, excess bond premium, and realized stock volatility within industries, and shows results are unchanged.&lt;/p&gt;
&lt;h3 id="q2-what-two-mechanisms-operate-in-a-menu-cost-model-when-common-volatility-increases-and-how-do-they-differ-from-idiosyncratic-volatility"&gt;Q2. What two mechanisms operate in a menu cost model when common volatility increases, and how do they differ from idiosyncratic volatility?&lt;/h3&gt;
&lt;p&gt;Two effects operate. The real options effect: higher volatility increases the option value of waiting, so firms expand the inaction band, decreasing frequency. The volatility effect: larger common shocks push more firms outside the band, but because it is a common shock, the resulting price changes are synchronized in the direction of the cost shock, which compresses dispersion. For idiosyncratic volatility, the volatility effect pushes price changes in both directions symmetrically, so both frequency and dispersion rise. This asymmetry is why aggregate and idiosyncratic volatility have different implications for monetary non-neutrality.&lt;/p&gt;
&lt;h3 id="q3-why-is-a-random-menu-cost-model-necessary-and-what-does-a-fixed-menu-cost-model-predict-instead"&gt;Q3. Why is a random menu cost model necessary, and what does a fixed menu cost model predict instead?&lt;/h3&gt;
&lt;p&gt;A fixed menu cost model (as in Golosov and Lucas 2007) features too strong a selection effect. When oil price volatility rises, more firms are pushed outside the action bands and they all move in the direction of the common cost shock, compressing price change dispersion (model predicts a 2.7 percent decline in dispersion per one standard deviation volatility increase) while frequency rises by 8.1 percent. This is the opposite of the empirical finding. Random menu costs break the tight link between the common shock and which firms adjust, because each firm draws a random menu cost each period. A substantial fraction of firms draw very large menu costs and never adjust regardless of the oil shock, while firms that do adjust include those reacting to idiosyncratic shocks (low menu cost draws), keeping the mix of price changes disperse even when aggregate volatility is high.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented-across-industries"&gt;Q4. What heterogeneity is documented across industries?&lt;/h3&gt;
&lt;p&gt;The main documented heterogeneity is in oil input intensity. The 10th percentile oil share is approximately 0.001 (oil spending equals 0.1 percent of value added) and the 90th percentile is 0.022 (2.2 percent of value added), with the average at 0.8 percent. The top-10-percent oil-share industries (e.g. Basic Chemical Manufacturing at 16.1 percent, Railroad Rolling Stock Manufacturing at 5.1 percent) show substantially stronger responses to oil price volatility shocks than low-oil industries. In terms of price setting statistics, across the eight octile sectors used in the multisector calibration, price change frequency ranges from 0.10 to 0.27, average size from 0.17 to 0.28, and standard deviation from 0.10 to 0.15 — heterogeneity that the multisector model replicates closely. There is no documented differential effect of oil price volatility between durable and non-durable goods industries.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-pass-through-estimates-from-oil-prices-to-producer-prices-and-why-do-they-matter-for-the-main-analysis"&gt;Q5. What are the pass-through estimates from oil prices to producer prices, and why do they matter for the main analysis?&lt;/h3&gt;
&lt;p&gt;The paper first establishes that oil prices actually pass through to producer prices, validating the cost-channel story. The short-run pass-through (impact month) is 1.0 percent (significant at 1 percent), meaning a 1 percent change in real oil prices raises producer price inflation by 1 percent in the same month. The 12-month cumulative pass-through is 8.6 percent (significant at 1 percent). These estimates are obtained from an industry-level panel regression with industry fixed effects and 12 lags of real oil price changes. The large pass-through relative to the average oil share of 0.8 percent is attributed to indirect transmission through input-output linkages. Pass-through establishes that oil is a relevant cost shifter for manufacturing producers, supporting the premise that oil price volatility would affect price-setting decisions.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run-on-the-main-empirical-findings"&gt;Q6. What robustness checks are run on the main empirical findings?&lt;/h3&gt;
&lt;p&gt;The paper conducts extensive robustness checks: (1) Alternative oil price series: WTI, Brent Crude, and Composite Refined Acquisition Cost — all give qualitatively and often quantitatively similar results. (2) Alternative volatility measures: stochastic volatility, GARCH(1,1), and realized volatility (within-month standard deviation of daily log price changes) — all produce consistent findings. (3) Crisis period: splitting the sample into 2008 crisis and non-crisis periods shows the dispersion result holds equally inside and outside the crisis. (4) Alternative dispersion measure: interquartile range of price changes in place of standard deviation — results unchanged. (5) Long-run oil usage: averaging the oil share across 1997, 2002, and 2007 IO tables rather than using only 1997 — dispersion results remain significant. (6) Trimming sensitivity: including all observations regardless of few price changes per industry-month does not change results. (7) Industry-level idiosyncratic volatility control: adding median realized stock volatility within the industry does not alter coefficients on oil price volatility.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-vavra-2014"&gt;Q7. How does this paper relate to and differ from Vavra (2014)?&lt;/h3&gt;
&lt;p&gt;Vavra (2014) studies idiosyncratic volatility and finds that both price change frequency and dispersion are countercyclical using CPI data. He matches these facts with a standard menu cost model with second-moment idiosyncratic productivity shocks. Klepacz differs by studying aggregate (oil price) volatility rather than idiosyncratic volatility, using PPI data, and finding that dispersion rises but frequency does not. These are the opposite implications from the mechanism standpoint: Vavra&amp;rsquo;s model would predict decreased dispersion when common volatility rises (because more prices synchronize), which is why Klepacz needs to modify the model with random menu costs. Klepacz then confirms that his random menu cost model can also reproduce Vavra&amp;rsquo;s idiosyncratic volatility facts when augmented with time-varying idiosyncratic volatility, with price change dispersion rising 1.2 percent and frequency rising 0.5 percent per one standard deviation idiosyncratic volatility shock. This shows the models are complementary, not contradictory.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-model-imply-for-the-magnitude-of-the-change-in-monetary-policy-effectiveness-across-the-full-empirical-distribution-of-oil-price-volatility"&gt;Q8. What does the model imply for the magnitude of the change in monetary policy effectiveness across the full empirical distribution of oil price volatility?&lt;/h3&gt;
&lt;p&gt;Beyond the 0.5 percent decline per one standard deviation oil price volatility increase, the paper simulates the full 1998-2014 oil price and volatility series through the model. At each point, it computes the on-impact output response to a 0.002 permanent log nominal output shock. The average monetary policy efficacy is 1 percent lower on impact during periods of the highest observed oil price volatility (such as 2009) relative to periods of the lowest oil price volatility (such as 2013). The cumulative consumption response is reduced by less than 1 percent throughout the first year following the monetary shock. These magnitudes are small enough that the paper concludes changes in aggregate volatility do not substantially alter the output-inflation tradeoff.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-aggregate-time-series-evidence-on-monetary-policy-effectiveness-across-oil-price-volatility-states"&gt;Q9. What is the aggregate time-series evidence on monetary policy effectiveness across oil price volatility states?&lt;/h3&gt;
&lt;p&gt;Section VI uses state-dependent local projections (Auerbach and Gorodnichenko 2013) with Romer-Romer (2004) monetary policy shocks over 1974-2007. The transition function equals one when the three-month moving average of oil price volatility exceeds the sample median. Controls include two lags of the monetary shock, current and two lags of the federal funds rate, log industrial production index, unemployment rate, log PPI, and log real oil price. Results show that the impulse response of the PPI price level to an expansionary monetary shock is not significantly different between high and low oil price volatility states. The high-volatility state estimates are less precise but are consistent with the linear model response, supporting the model&amp;rsquo;s implication that monetary policy effectiveness is not a function of oil price volatility.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper implies that policymakers should not systematically increase nominal stimulus in response to high aggregate volatility on the grounds that policy is less effective. The output-inflation tradeoff is nearly time-invariant. If policymakers over-stimulate believing effectiveness has declined, the result is unnecessary inflation. However, this conclusion is specific to aggregate (common) volatility shocks, not idiosyncratic volatility — the source of volatility matters for the direction of price-setting response and hence for the policy implications. The paper explicitly states that the analysis applies to oil price volatility but extends conceptually to policy uncertainty, exchange rate volatility, and global demand volatility. One scope condition: the model abstracts from a monetary policy reaction function that responds directly to oil prices (as in Kilian and Lewis 2011 or Bodenstein et al. 2012), so the quantitative results apply to the partial equilibrium price-setting channel rather than to the full general equilibrium policy transmission.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-business-cycle-properties-of-price-change-moments-in-the-ppi-and-how-do-they-compare-to-cpi-findings"&gt;Q11. What are the business cycle properties of price change moments in the PPI, and how do they compare to CPI findings?&lt;/h3&gt;
&lt;p&gt;Table 1 shows that the standard deviation of PPI price changes is countercyclical: the recession dummy adds 0.008 to the mean dispersion of 0.127 (significant at 5 percent). Price change frequency rises during recessions by 0.017 but the coefficient is not statistically significant. These patterns are qualitatively consistent with Vavra (2014) and Bachmann et al. (2019). Comparing PPI and CPI (Table 2): both have frequency around 15 percent and average absolute size around 7-8 percent. The main difference is that the PPI has a higher fraction of small price changes (22 percent vs. 12 percent in the CPI), reflecting a higher frequency of very small adjustments. Price change dispersion is higher in the PPI (standard deviation 0.13) than the CPI (0.08). Monthly inflation correlation between the two series is 0.80 over 1998-2014.&lt;/p&gt;
&lt;h3 id="q12-what-caveats-or-limitations-does-the-paper-acknowledge"&gt;Q12. What caveats or limitations does the paper acknowledge?&lt;/h3&gt;
&lt;p&gt;The main caveats are: (1) The model does not feature a monetary policy reaction function for oil prices, abstracting from the general equilibrium feedback between oil shocks and interest rate policy. (2) The multisector model replicates the positive relationship between oil price volatility and price change dispersion but cannot match the empirically negative frequency response — the model predicts higher relative frequency for high-oil sectors during volatility episodes, while the data show lower relative frequency. (3) The time-varying idiosyncratic volatility extension uses a simplifying assumption that idiosyncratic volatility is perfectly negatively correlated with oil prices, primarily for computational tractability. (4) The model focuses on manufacturer producer prices (the PPI) and on oil as a non-produced input, abstracting from oil in the household consumption function.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Price change dispersion&lt;/strong&gt;: The within-industry standard deviation of non-zero price changes in a given month, measuring how spread out the price changes are in the cross-section of items. A more disperse distribution means price changes are scattered across a wide range of sizes and directions, so a monetary shock shifts fewer prices past the adjustment threshold and has larger real effects. The paper measures it as the square root of the mean squared deviation of item-level price changes from the industry mean, computed only over non-zero changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Real options effect&lt;/strong&gt;: One of two mechanisms through which higher volatility affects price-setting in a menu cost model. Higher volatility increases the value of waiting before paying the menu cost to adjust, because the expected loss from being at a suboptimal price for one more period is smaller relative to the cost of adjusting when future shocks are large and uncertain. This pushes the action and inaction bands outward, reducing the frequency of price adjustment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Volatility effect&lt;/strong&gt;: The second mechanism through which higher volatility affects price-setting. For idiosyncratic volatility, larger idiosyncratic shocks push prices outside the inaction bands in both directions, increasing both frequency and dispersion. For common (aggregate) volatility, larger common shocks push prices outside the bands mostly in one direction, increasing frequency but decreasing dispersion (in a fixed-menu-cost model). In a random menu cost model, this synchronization is attenuated, allowing dispersion to rise.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Random menu costs&lt;/strong&gt;: A modeling device where each firm draws an i.i.d. menu cost each period from a non-degenerate distribution (specifically, a transformation of an exponential distribution as in Luo and Villar 2020) rather than paying a single fixed cost. The distribution has fat tails, giving substantial probability of very low or very high cost draws. This randomness breaks the tight selection effect of fixed-menu-cost models: which firms adjust depends not only on how far their price is from optimal but also on their menu cost draw, so many firms do not adjust even when their price gap is large. This attenuates the synchronization of price changes in response to a common shock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Industry-specific oil demand variable&lt;/strong&gt;: A Bartik-style instrument constructed by multiplying an industry&amp;rsquo;s pre-determined oil input share (oil spending as a fraction of value added from the 1997 IO tables) by aggregate oil price volatility or oil price inflation. The pre-determined share measures the industry&amp;rsquo;s structural sensitivity to oil, while the aggregate oil volatility provides exogenous time variation. The interaction captures the differential exposure of high-oil industries to aggregate oil price volatility shocks, enabling identification via cross-industry variation after controlling for time and industry fixed effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stochastic volatility of oil prices&lt;/strong&gt;: A latent volatility process estimated from real WTI oil prices using an AR(1) model for the log oil price level and a mean-reverting AR(1) process for the log standard deviation of oil price innovations. Estimated via Bayesian MCMC with a particle filter (Sequential Importance Resampling) to handle the nonlinearity, using data from 1986-2014. Produces a smoothed series of time-varying oil price uncertainty. Key estimated parameters: oil price persistence ρ_o = 0.999, volatility persistence ρ_σ = 0.887, unconditional mean log-volatility σ = -2.607 (implying standard deviation of oil price shock ≈ 7.4 percent), and volatility shock size φ = 0.127.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Selection effect&lt;/strong&gt;: In state-dependent pricing models, the mechanism by which the prices that actually change are not a random subset but are selected based on how far they are from their optimal level. A strong selection effect (as in Golosov and Lucas 2007) means that only prices far from optimal change, so average price change size is large and price change frequency is low. Under a common volatility shock with a strong selection effect, more prices are pushed far from optimal in the same direction, causing them all to adjust together — compressing dispersion and increasing frequency. Random menu costs weaken the selection effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Monetary non-neutrality&lt;/strong&gt;: The degree to which a change in the money supply (or nominal spending) affects real output rather than just the price level. In menu cost models, non-neutrality arises because not all prices can adjust instantaneously: a monetary shock shifts the desired price change distribution, but only firms near the adjustment threshold respond, leaving real prices for the others unchanged. After conditioning on price change frequency, higher price change dispersion implies fewer prices are near the threshold, so a given monetary shock affects fewer prices in one direction and has larger real effects (greater non-neutrality). This is the key channel linking the paper&amp;rsquo;s empirical findings to monetary policy effectiveness.&lt;/p&gt;</description></item><item><title>Property rights, fiscal capacity, and social capacity: The lasting impact of the Taiping Rebellion</title><link>https://macropaperwarehouse.com/papers/property-rights-fiscal-capacity-and-social-capacity-the-lasting-impact-of-the-taiping-rebellion/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/property-rights-fiscal-capacity-and-social-capacity-the-lasting-impact-of-the-taiping-rebellion/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: How do civil wars affect long-term development, and through which institutional mechanisms? The paper studies the Taiping Rebellion (1850-1864) in Qing China, one of history&amp;rsquo;s deadliest civil wars (at least ~20 million deaths, with some estimates of 70-100 million), as a critical juncture in China&amp;rsquo;s path to modernity. It matters because the rebellion generated large, persistent regional institutional variation that can help explain what the authors call the &amp;ldquo;Intra-China Divergence&amp;rdquo; — regional GDP-per-capita gaps as large as 27-to-1 (Dongguan vs. Tianshui, 2010) that rival the world&amp;rsquo;s largest inter-regional gaps.&lt;/p&gt;
&lt;p&gt;Data and design: A prefecture-level (occasionally county-level) panel covering 266 prefectures in China proper (1820 delineation). 55 prefectures fell under Taiping control (treatment) — split into 37 &amp;ldquo;Early Taiping&amp;rdquo; prefectures (occupied up to 1859, in Anhui/Jiangxi/Hubei, ambiguous land rights) and 18 &amp;ldquo;Late Taiping&amp;rdquo; prefectures (occupied from 1860, in Jiangsu/Zhejiang, stronger land rights) — and 211 control prefectures. Population is observed at seven points (1820, 1851, 1880, 1910, 1953, 1982, 2000). The core strategy is difference-in-differences (1820 reference year, prefecture and year fixed effects), supplemented by propensity-score matching (135-prefecture matched sample), a spatial autoregressive (SAR) model, and an instrumental-variable strategy using the longitude of the prefectural seat (motivated by the Taiping Navy&amp;rsquo;s eastward-along-the-Yangtze military strategy; first-stage F-statistics above 20).&lt;/p&gt;
&lt;p&gt;Main quantitative findings (with scope conditions): (1) Population: The rebellion caused large, permanent population losses. The Taiping DID coefficient is -0.45 in 1880 (a 36% lower population growth rate vs. control) and -0.51 in 1953 (40% lower) — no convergence. Crucially, in the matched sample Late Taiping areas recovered (no significant long-run population gap vs. control) while Early Taiping areas did not (an immediate ~30% drop in 1880 plus further decline). (2) Property rights: In 1915 county data, the idle-land share is 3.6 percentage points higher in Early Taiping than control counties, while Late Taiping is not significantly different from control — supporting the property-rights hypothesis. (3) Fiscal capacity (likin): Taiping areas collected ~12 times (e^2.5) as much likin per 1,000 sq km as control areas in 1869-1879, still 3.7 times as much in 1922-1925. Late Taiping areas had even higher intensity (22.2x in 1869-1879; 6.1x in 1922-1925) than Early Taiping (9.0x; 2.7x). (4) Social capacity (charities): On average the rebellion had no significant effect, but Late Taiping areas saw charity growth ~56 percentage points (44 log points) above control by 1880, rising to ~78 percentage points (58 log points) by mid-20th century. (5) Long-term development: Driven entirely by Late Taiping areas — 1982 agricultural+industrial output per capita 90% higher (64 log points), 2010 GDP per capita 87% higher (63 log points), and 2010 fiscal revenue per capita 203% higher (111 log points) than control; Early Taiping is statistically indistinguishable from control. Late Taiping counties also show higher post-1895 industrial firm entry. (6) Civic outcomes and resilience: Using CGSS 2010, Late Taiping residents show higher trust in personal networks and greater civic engagement (political attention, local participation). During the Great Famine (1959-1961), Taiping areas had 6.9% larger survivor cohorts; the effect is 28% stronger in Late Taiping (8.4%) than Early Taiping (6.5%).&lt;/p&gt;
&lt;p&gt;Implications: Violent conflict can leave lasting positive institutional imprints — through property rights, decentralized local fiscal capacity (&amp;ldquo;war made the state&amp;rdquo; at the local level), and elite-led social capacity — conditional on favorable initial conditions (strong gentry, wealthier commercial regions). The authors argue cultivating civil society and social capacity could yield large payoffs given China&amp;rsquo;s strong-state/weak-society configuration.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the core identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The baseline is a difference-in-differences comparing Taiping vs. control prefectures over 1820-2000, with prefecture and year fixed effects and 1820 as the reference year. Identification rests on parallel pre-trends: the Taiping coefficient in 1851 (pre-rebellion) is small and insignificant, indicating no differential selection conditional on controls. The main threats are: (i) the binary Taiping measure aligning with provincial boundaries and picking up broad regional dynamics; (ii) control-group contamination because some control prefectures were temporarily conquered (but not governed) by the Taiping Army; (iii) spatial spillovers between neighbors (Tobler&amp;rsquo;s law / Kelly 2019 critique); (iv) omitted subsequent historical events; and (v) omitted variables differing systematically between treated and control areas. The authors address these with dosage measures (battles, occupation months), matching, a SAR model, an IV (longitude), explicit controls for the Taiping conquest, an adjacent-treatment indicator, leave-one-province-out checks, and controls for many other historical events.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-instrumental-variable-strategy-work-and-why-might-longitude-be-valid"&gt;Q2. How does the instrumental-variable strategy work and why might longitude be valid?&lt;/h3&gt;
&lt;p&gt;Longitude of the prefectural seat instruments for the Taiping dummy. Relevance: the Taiping leaders&amp;rsquo; July 1852 military plan was to march eastward along the Yangtze, capture Jiangning (Nanjing), and expand from there using their dominant navy — so eastern (higher-longitude) prefectures were far more likely to fall under Taiping rule (Table 1 confirms Taiping prefectures have significantly larger longitudes; first-stage F-statistics above 20, Shea&amp;rsquo;s partial R-squared above 0.1). Exclusion: prefecture fixed effects absorb time-invariant geographic advantages, and year-dummy interactions with key geography (distances to coastline, Grand Canal, Yangtze) allow flexible time-varying geographic effects; conditional on these, longitude is argued to be excludable. IV estimates are larger in magnitude than OLS but qualitatively confirm a persistent negative population effect (robust to Anderson-Rubin weak-IV inference). The authors caution that omitted determinants correlated with longitude cannot be fully ruled out.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-four-hypotheses-and-how-are-they-distinguished-empirically"&gt;Q3. What are the four hypotheses and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;(1) Property-rights hypothesis: Late Taiping areas (post-1860 &amp;lsquo;direct tenant payment&amp;rsquo; system creating de facto/de jure tenant ownership) had better-defined land rights than Early Taiping areas (collapsed landlord system, lost deeds, anti-rent movements), so should have less idle land and faster population recovery — tested via the 1915 idle-land cross-section and the Early-vs-Late population DID. (2) Likin-as-fiscal-capacity hypothesis: Qing fiscal decentralization and the likin tax (introduced 1853) strengthened local fiscal capacity, persistently higher in Taiping (especially Late Taiping) areas — tested via the likin-intensity DID. (3) Social-change hypothesis: elite-led militias and reconstruction spurred charities (&amp;lsquo;benevolent halls&amp;rsquo;/shantang) as bridging social capital, especially in Late Taiping areas — tested via charity-stock DID and by adding charities as a mediator in long-term regressions. (4) Social-cohesion-and-civic-engagement hypothesis: forged social capital persists, raising modern trust/civic engagement and reducing Great Famine deaths — tested via CGSS 2010 and famine-survivor cohort ratios.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented"&gt;Q4. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;The central heterogeneity is Early vs. Late Taiping. Early Taiping areas (Anhui/Jiangxi/Hubei) suffered permanent population loss, higher idle land (+3.6pp), only modest likin gains, no charity growth, no long-term development advantage, and weaker famine resilience. Late Taiping areas (Jiangsu/Zhejiang) recovered population, had no excess idle land, far higher likin intensity (22x early period), large charity growth (+56 to +78pp), strong long-term development gains (90%/87%/203% in output/GDP/fiscal revenue), higher modern trust and civic engagement, and the strongest famine resilience (8.4% vs 6.5%). Industrialization heterogeneity is also temporal: no Early/Late firm-entry difference before 1895, but after the 1895 Treaty of Shimonoseki liberalized private industry, Late Taiping counties had more entry and Early Taiping fewer.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;For the population results: dosage interactions (log battles, log occupation months); excluding six most-intense-fighting prefectures (Wuchang, Songjiang, Anqing, Jiangning, Suzhou, Hangzhou); controlling for newly selected jinshi (civil-service quota channel); a SAR spatial model (after Pesaran cross-sectional-dependence tests); PSM matched sample; longitude IV with Anderson-Rubin inference; controls for seven other historical events (Guangxu Drought, Hui Revolt, Nian Rebellion, early-Republic conflicts, Sino-Japanese War, Chinese Civil War, missionary activity); explicit controls for Taiping conquest vs. regime; an adjacent-treatment indicator (Butts 2021) for spillovers; and leave-one-province-out exclusion. Long-term development results add SAR, matching, historical-event controls including the Cultural Revolution, and an &amp;lsquo;intermediate-term&amp;rsquo; 1930s industrialization check. Famine results are robust to alternative famine-severity measures, SAR, matching, and historical-event controls.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-mediation-analysis-handled-and-what-does-it-show"&gt;Q6. How is the mediation analysis handled and what does it show?&lt;/h3&gt;
&lt;p&gt;The authors add likin intensity (1880) and average charities (1880-1941) to cross-sectional long-term regressions, explicitly flagging these as endogenous &amp;lsquo;bad controls&amp;rsquo; (Angrist-Pischke 2009; Imai et al. 2011) to be interpreted cautiously as descriptive mediation. Findings: a one-SD increase in likin intensity is associated with +1.7pp middle-school completion, +4.8pp literacy, +5.3% schooling, and +12.2% (11.5 log points) GDP per capita in 2010. A one-SD increase in charities is associated with +15% 1982 output, +20% 2010 GDP, and +55% 2010 fiscal revenue per capita. Once charities are netted out, Late Taiping advantages in output, GDP, and fiscal revenue are attenuated by about 17%, 14%, and 22% respectively — highlighting the social-capacity channel.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-great-famine-resilience-result-connect-to-the-rebellion"&gt;Q7. How does the Great Famine resilience result connect to the rebellion?&lt;/h3&gt;
&lt;p&gt;Famine severity is measured by &amp;lsquo;Famine Control&amp;rsquo; = ratio of cohort size born during the famine (1959-1961) to cohort size born pre-famine (1954-1957) from the 1990 census 1% sample (higher = less severe). Taiping areas had a 6.9% larger survivor cohort than non-Taiping; the effect is 8.4% in Late Taiping vs. 6.5% in Early Taiping. Back-of-envelope, the Late Taiping experience would have &amp;lsquo;saved&amp;rsquo; ~31,374 people in an average prefecture (17% of the 1959-1961 cohort) vs. ~24,145 (13%) for Early Taiping. Controlling for political radicalism (reverse party-member density, -1*PMD, after Yang 1996) does not change the result. The mechanism: higher social capital made local officials more sympathetic/less radical in grain procurement and citizens better able to act collectively (paralleling Cao-Xu-Zhang 2022 on clan density and Hu-Yao-You 2023 on home-county officials).&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q8. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Prior Taiping studies examined narrower consequences: civil-service exam quotas (Li 2014), demographic and industrialization effects (Li and Ma 2016), migration and public goods (Hao and Xue 2017), and late-Qing power distribution (Bai, Jia, and Yang 2023). None addressed the rebellion&amp;rsquo;s enduring impacts on modern development, social trust, and Great Famine responses, nor the property-rights/fiscal-capacity/social-capacity mechanism triad. It complements Xue (2021) on Qing charities, generalized trust, and political participation, but extends to development outcomes. Against the European state-building literature (war strengthens central state capacity via centralization), this paper&amp;rsquo;s distinctive claim is that the Taiping Rebellion strengthened LOCAL fiscal capacity through DECENTRALIZATION, and expanded local social capacity that constrained the central state.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The benefits of war-induced institutions are conditional, not universal: they appeared chiefly in Late Taiping areas with a strong gentry class and favorable initial conditions for modern sectors (the wealthier, more commercial Lower Yangtze). The likin/fiscal-capacity benefits are explicitly stated to be conditional on strong gentry and good modern-sector initial conditions. The broad implication is that, given China&amp;rsquo;s very strong state but still weak society today, cultivating civil society and strengthening social capacity could yield particularly large long-term payoffs. The authors also caution (Appendix F.1) that likin could be distortionary taxation rather than fiscal capacity, arguing the fiscal-capacity interpretation is more relevant for long-term development.&lt;/p&gt;
&lt;h3 id="q10-what-significant-caveats-does-the-paper-acknowledge"&gt;Q10. What significant caveats does the paper acknowledge?&lt;/h3&gt;
&lt;p&gt;Long-term mechanisms cannot be exhaustively identified — likin and charities are endogenous outcomes, so mediation magnitudes are descriptive, not causal. History contains near-infinite interrelated events, so confounding cannot be fully eliminated (a fundamental limitation of all history-based work). The IV may have omitted correlates of longitude. Some 2SLS estimates for development outcomes were largely insignificant. The charity-stock measure assumes charities persisted once founded (no closure dates in the data). On property-rights persistence: using 2005 World Bank Enterprise Survey data they find no association between modern firms&amp;rsquo; perceived property-rights protection and Taiping regimes, suggesting the channel works through income effects rather than persistence of property rights per se.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Early vs. Late Taiping areas&lt;/strong&gt;: Early Taiping = prefectures occupied by the rebels up to 1859 (Anhui, Jiangxi, Hubei), where the old landlord system collapsed and land rights stayed ambiguous; Late Taiping = prefectures occupied from 1860 (Jiangsu, Zhejiang), where the Taiping introduced a &amp;lsquo;direct tenant payment&amp;rsquo; (作佃交粮) system and issued new deeds, granting tenants de facto/de jure ownership. This distinction is the paper&amp;rsquo;s central source of institutional variation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Likin (lijin)&lt;/strong&gt;: A local tax on trade and commerce introduced in 1853 (a transit tax on travelling merchants&amp;rsquo; goods plus a business tax on resident merchants), collected in a decentralized, province-specific way. In the paper it is the operational measure of local fiscal capacity (likin revenue per 1,000 sq km), not central state capacity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social capacity&lt;/strong&gt;: In the paper&amp;rsquo;s sense, the ability of society to act collectively, constrain the state, and empower its members — operationalized empirically by the stock of local charity organizations (&amp;lsquo;benevolent halls&amp;rsquo;/shantang) that functioned as bridging social capital across classes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Likin-as-fiscal-capacity hypothesis&lt;/strong&gt;: The claim that the rebellion-induced likin system durably raised LOCAL fiscal capacity (an instance of Tilly&amp;rsquo;s &amp;lsquo;war made the state&amp;rsquo; operating locally rather than centrally), which improved public-goods provision and long-run development — conditional on strong gentry and favorable modern-sector initial conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stationary bandit (applied to Late Taiping rulers)&lt;/strong&gt;: Borrowing Olson (1993): in Late Taiping areas the consolidated, longer-horizon Taiping regime behaved like a stationary bandit, lowering effective tax rates, encouraging land registration, and securing tenant property rights to expand the tax base and promote production, unlike the looting/confiscation of the early stage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Famine Control&lt;/strong&gt;: The paper&amp;rsquo;s local famine-severity measure: the ratio of the cohort born during the Great Famine (1959-1961) to the cohort born pre-famine (1954-1957) in the 1990 census; a higher value means less severe famine and more survivors, and it is less vulnerable to government understatement of famine deaths.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intra-China Divergence&lt;/strong&gt;: The authors&amp;rsquo; term for China&amp;rsquo;s persistent, very large regional disparities in economic performance (up to 27-to-1 in GDP per capita) despite all regions historically sharing similar Malthusian income levels — the macro puzzle the rebellion&amp;rsquo;s institutional legacy helps explain.&lt;/p&gt;</description></item><item><title>Remote Work and City Structure</title><link>https://macropaperwarehouse.com/papers/remote-work-and-city-structure/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/remote-work-and-city-structure/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Monte, Porcher, and Rossi-Hansberg ask why remote work surged abruptly and permanently after COVID-19 despite information-technology advances raising it only marginally between 1980 and 2019, why the change was so heterogeneous across cities, and what the welfare consequences are. Their answer is a coordination mechanism: working downtown (the CBD) yields productive interactions with other in-office workers but entails commuting/congestion costs, while remote work avoids those costs but forgoes agglomeration benefits. Because workers do not internalize the spillovers they confer, a worker prefers the office only if others commute too — generating, in a dynamic discrete-choice model with idiosyncratic preferences and fixed switching costs, the possibility of MULTIPLE stationary equilibria with different permanent commuter shares. A temporary shock (the pandemic) that drives commuters near zero can then select the low-commuting equilibrium permanently.&lt;/p&gt;
&lt;p&gt;The model is a dynamic monocentric city (disk-shaped, radially symmetric CBD, absentee landlords, Cobb-Douglas utility, Gumbel idiosyncratic shocks). Multiplicity arises (Proposition 4.3) when agglomeration forces are strong enough — the net strength delta + xi exceeds a threshold above theta + gamma/(2mu) — AND remote-work productivity relative to office productivity z/A lies in an intermediate &amp;ldquo;cone of multiplicity&amp;rdquo; (neither too low nor too high). The authors quantify city-specific parameters for U.S. CBSAs using pre-2019 data (Census/ACS 1980-2023, NLSY79 panel of 4,147 individuals 1998-2022, SafeGraph cell-phone mobility, Zillow ZHVI zip-code house prices). Estimation: transition elasticity s = 0.30 (elasticity of transitions into remote work = 3.09), fixed switching cost F = 1.78 (equivalent to giving up 83% of a year&amp;rsquo;s earnings); agglomeration externality delta with mean 0.067 (SD 0.022, 619 CBSAs); the amenity-vs-congestion difference xi - theta is statistically insignificant and set to zero.&lt;/p&gt;
&lt;p&gt;Stylized facts. Predicted remote-work share (controlling for composition) rose in the ACS from under 1% (1980) to 2.6% (2019), jumped to 12% (2020), peaked at 15% (2021), and fell to 11% (2023); NLSY shows a parallel path (1.4% in 1998 to 3.7% in 2018, 9.2% in 2020, 7.8% in 2022). The remote-work wage premium rose steadily but did NOT jump post-2018: ACS discount of 44.5% in 1980 became a 6.5% premium by 2022; NLSY discount fell from 18.5% (2000) to 3.1% (2022). A stable premium alongside a sudden quantity jump argues against pure productivity/preference shocks.&lt;/p&gt;
&lt;p&gt;Mobility/housing facts. All cities dropped to ~20% of pre-pandemic CBD trips in spring 2020 (about a 75% drop, unrelated to city size). Recoveries diverged: the 25 largest CBSAs (employment &amp;gt; 1.5M) stabilized at ~60% of January-2020 trips, while the 663 smallest (&amp;lt; 150K) returned fully to pre-pandemic levels by early 2021. New York and San Francisco stabilized near 40%; Madison, WI recovered fully. House-price distance gradients flattened ~0.01 everywhere by January 2021; the flattening persisted and stabilized around 0.095 by end-2024 in large cities but reversed in small ones.&lt;/p&gt;
&lt;p&gt;Results and welfare. Of 278 estimated CBSAs, 208 were inside their cone of multiplicity pre-pandemic; larger cities are systematically more likely to be inside (probit on log employment significant). The cone indicator predicts trip shortfalls (R-squared 0.144 alone, retaining significance with controls) and gradient flattening. Welfare: comparing high- vs low-commuting stationary equilibria for the 208 cone cities, the loss from switching is positive but modest — mean 2.3%, median 2.2%, range 1.2% to 4.0% (Table 3). Average wages fall sharply (15-35%) but option-value and commuting-cost savings offset most of it; net strength delta - gamma/(2mu) predicts the loss with R-squared 0.85. Cities with trips at 60% or less of pre-pandemic levels have an average welfare loss of 2.7%.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-economic-mechanism-and-how-does-it-generate-multiple-equilibria"&gt;Q1. What is the core economic mechanism, and how does it generate multiple equilibria?&lt;/h3&gt;
&lt;p&gt;Office work confers productivity spillovers and CBD amenity value that rise with the mass of in-office workers (L-tilde-c), but workers do not internalize these external benefits. So each worker prefers the office only if enough others commute. In a dynamic setting with idiosyncratic Gumbel preference shocks and fixed switching costs F, this coordination can produce multiple stationary equilibria: a high-commuting and a low-commuting one (with an unstable equilibrium E2 between them). Multiplicity requires (Prop 4.3) static agglomeration forces (delta + xi) above a threshold eta_min &amp;gt; theta + gamma/(2mu), AND relative remote productivity z/A in an intermediate interval Z — the &amp;lsquo;cone of multiplicity.&amp;rsquo; If z/A is too low, the high-commuting equilibrium is unique; if too high, only the remote equilibrium survives.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identificationquantification-strategy-and-its-main-threats"&gt;Q2. What is the identification/quantification strategy and its main threats?&lt;/h3&gt;
&lt;p&gt;To avoid taking a stand on which equilibrium generated the data, the authors rely ENTIRELY on pre-2019 data (when every city was plausibly in the high-commuting equilibrium) and on model relationships that hold in any equilibrium. Four steps: (1) transition elasticity s and cost F from NLSY79 transition probabilities via a CCP/log-linear regression (eq. 21), using past wage ratios as an instrument for future ratios to address measurement error / forward-looking expectations (IV eta0 = -0.47, eta1 = 3.09); (2) agglomeration externality delta_j from commuter-wage changes instrumented by 1980 occupational composition interacted with economy-wide occupation-specific commuter-share changes (shift-share IV, eq. 26-28), with five industry groups; (3) remote/office productivity z_j, A_j from occupation-level remote-work premia (NLSY, 22 occupation groups) reweighted by city occupation shares; (4) transport-cost elasticity gamma_j from CBSA-specific housing rent-distance gradients (ACS block-group rents 2015-2019). Main threats: selection of workers into remote work on unobservables (addressed by NLSY individual fixed effects), endogeneity of commuter shares to local productivity shocks (addressed by the shift-share IV), and the assumption that all cities were in the high-commuting equilibrium in 2019; tau_j is calibrated to match each city&amp;rsquo;s 2019 Lc/L.&lt;/p&gt;
&lt;h3 id="q3-how-do-the-authors-rule-out-competing-explanations-pure-productivitypreference-shocks-congestion-establishment-size-occupational-shift"&gt;Q3. How do the authors rule out competing explanations (pure productivity/preference shocks, congestion, establishment size, occupational shift)?&lt;/h3&gt;
&lt;p&gt;National productivity/preference shocks: would be expected to leave some lasting imprint even in small cities, but small CBSAs reverted fully, and at least 34% of jobs remain teleworkable even in fully-reverting cities (Dingel-Neiman teleworkable share ranges 25-55% across CBSAs), so low telework capacity cannot explain reversion; cities with permanent 40%+ trip declines have only a modestly higher 43% teleworkable share. The wage premium shows no differential evolution across high- vs low-teleworkable occupations over the pandemic. Congestion: if congestion drove the shift, large cities should show lower CBD propensity pre-pandemic, but the opposite holds (30.6% of trips to CBD in large vs 15.6% in small CBSAs in late 2019). Establishment concentration: employment is LESS concentrated in smaller cities, so big-employer return-to-office decisions cannot explain reversion. Occupational shift: teleworkable employment share rose only ~5% post-pandemic, and rose MORE in smaller CBSAs (7.9%) than larger (5.8%) by end-2023, the wrong direction to explain the heterogeneity.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-across-cities-is-documented-and-how-does-it-map-to-the-theory"&gt;Q4. What heterogeneity across cities is documented and how does it map to the theory?&lt;/h3&gt;
&lt;p&gt;Large cities (high agglomeration, high net strength delta - gamma/(2mu), which rises with size: doubling size raises net strength ~0.004 off a mean 0.049) are disproportionately inside the cone of multiplicity (208 of 278 estimated cities in-cone; probit on log employment positive and significant). These cities show permanent CBD-trip declines (stabilizing ~60% for the 25 largest) and persistent gradient flattening (~0.095 by 2024). Small cities are mostly outside the cone, with unique equilibria, and revert fully. The cone indicator is also positively associated with delta_j and z_j/A_j and negatively with gamma_j, as the theory predicts.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Estimates of s and F are similar using restricted-use county-geocoded NLSY and under an alternative city-partition definition (two days/week remote). Main results are robust to lower delta_j and higher gamma_j calibrations (Appendix A.17). A CES production function in remote/in-person labor yields very large substitution elasticities, motivating the linear specification. An endogenous-housing-supply model yields a nearly identical rent gradient (because commuters were a high share of employment pre-2020). Office-trip-only versions of the mobility figures (workplace visits) show similar patterns. The cone indicator retains significance in Table 2 after adding teleworkable share, pre-pandemic CBD-trip share, industry value-added shares, and total employment; results hold for an alternative binary &amp;lsquo;returned to office&amp;rsquo; indicator 1back(5,20). Multiple DYNAMIC equilibria were not found in numerical exercises (Appendix B.6).&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-differ-from-closely-related-prior-work"&gt;Q6. How does this paper differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Unlike Davis, Ghent &amp;amp; Gregory (2024) (remote productivity via adoption externalities), Parkhomenko &amp;amp; Delventhal (2024) (amenity value of remote work), and Duranton &amp;amp; Handbury (2023) (exogenous changes in who may work remotely), this paper does NOT rely on exogenous productivity or amenity/preference shocks to explain the large persistent jump. Instead a temporary commuter shock SELECTS among pre-existing multiple equilibria. Liu &amp;amp; Su (2023) document a falling urban wage premium for remote-amenable occupations (consistent with weaker agglomeration). The paper&amp;rsquo;s documented divergence of residential rent-distance gradients between large and small cities is, to the authors&amp;rsquo; knowledge, a new fact, interpreted structurally. Owens, Rossi-Hansberg &amp;amp; Sarte (2020) similarly use coordination/residential externalities (Detroit neighborhoods).&lt;/p&gt;
&lt;h3 id="q7-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q7. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Because the coordination failure operates partly OUTSIDE firm boundaries, individual firms&amp;rsquo; return-to-office mandates may be insufficient to restore the high-commuting equilibrium. City-level interventions — taxing remote work or subsidizing commuting — could in principle move a city back, since the only active externality in the quantification is a positive agglomeration externality (implying too little commuting relative to the efficient benchmark in all equilibria). However, the authors stress these welfare effects and the effectiveness of policy remain open questions; their welfare numbers depend on estimation details and the abstraction from a system-of-cities with migration.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-main-caveats-and-abstractions"&gt;Q8. What are the main caveats and abstractions?&lt;/h3&gt;
&lt;p&gt;The model treats each city as a CLOSED economy: no inter-city migration, trade, or investment links, though the authors note large cities show a small differential population drop (Appendix A.9), attributed to low migration elasticities. Remote work is &amp;lsquo;partial&amp;rsquo; with a FIXED fraction mu = 3/5 of days at home, not chosen. Occupational heterogeneity is abstracted from (justified by rare occupation transitions). The amenity (xi) vs congestion (theta) externalities are not separately identified and set to zero (difference insignificant). Spillovers are not internalized by firms in the model. The welfare ranking (high-commuting preferred) is intuited from the single positive externality rather than formally proven.&lt;/p&gt;
&lt;h3 id="q9-why-is-there-a-discrepancy-between-the-abstracts-welfare-figures-and-per-city-numbers"&gt;Q9. Why is there a discrepancy between the abstract&amp;rsquo;s welfare figures and per-city numbers?&lt;/h3&gt;
&lt;p&gt;The abstract and revised Table 3 report a mean welfare loss of 2.3% (median 2.2%, range 1.2%-4.0%) across the 208 cone cities, and state cities with permanently low commuting (60% or less of pre-pandemic trips) experience average losses of 2.3% (2.7% in the text). The introduction additionally quotes specific city losses (about 3.7% for Los Angeles and San Jose, 3.2% for New York, 2.8% for San Francisco, 2% for Phoenix); these are the largest cities and lie within or near the upper part of the distribution, consistent with welfare loss rising in net agglomeration strength (R-squared 0.85 of loss on delta - gamma/(2mu)).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;!-- flags: Welfare magnitudes: the final/revised headline figures are mean 2.3%, median 2.2%, range 1.2-4.0% (Table 3, 208 cities). The Introduction also cites larger per-city losses (3.7% LA/San Jose, 3.2% NYC, 2.8% SF, 2% Phoenix); these are consistent with the distribution (loss rises with net agglomeration strength) but appear to be from a specific large-city calibration table, not the summary distribution. Reported both, flagged for reviewer., Paper is a Nov 2025 revision of NBER WP 31494 (orig. July 2023); some figures span data through end-2024/Nov-2024, later than the original draft. --&gt;</description></item><item><title>Returns to experience and the elasticity of labor supply</title><link>https://macropaperwarehouse.com/papers/returns-to-experience-and-the-elasticity-of-labor-supply/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/returns-to-experience-and-the-elasticity-of-labor-supply/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: A large empirical literature uses micro data to estimate the intertemporal elasticity of substitution (IES) of labor supply, a parameter crucial for understanding business-cycle fluctuations in hours and labor-supply responses to tax policy. Standard micro studies, which regress log hours on log wages, typically obtain small estimates (in the range of 0-0.4), leading much of the profession to conclude labor-supply elasticities are small. These studies assume wages evolve exogenously. The authors argue that when wages rise with work experience (learning-by-doing, LBD), the marginal return to an hour of work exceeds the wage because it also includes the discounted increase in all future earnings from added experience. Because the wage is only one component of total remuneration, a given percentage wage increase raises the total marginal return by a smaller percentage, so regressing hours on wages produces a downward-biased estimate of the IES. Critically, the omitted variable (the ratio of total remuneration to the wage) is mechanically related to the wage, so the bias cannot be corrected by instrumental variables or natural experiments.&lt;/p&gt;
&lt;p&gt;Model and strategy: The authors extend a MaCurdy (1981) life-cycle model of consumption and labor supply to include LBD, where the wage equals marginal return to human capital times a human-capital stock that grows with experience. They derive a log-linear labor-supply equation with an extra term capturing future returns to work, which is negatively correlated with the wage. Their key insight: for individuals whose future returns to experience are negligible (the term F approaches zero, e.g., at end of working life or at very high human-capital stocks), the standard regression yields an unbiased IES estimate, allowing them to remain agnostic about the human-capital accumulation process.&lt;/p&gt;
&lt;p&gt;Data: They use daily labor-supply records of Florida spiny lobster trap fishermen from the Florida Fish and Wildlife Conservation Commission, covering the 1986 through 2007 seasons (a 22-year panel), restricted to the first 70 days of each season. Analysis samples are drawn from fishermen active 2001-2005. Wage variation is exogenous and partly predictable because lobster catch rates rise around the new moon (and with rough weather). The moon phase is the key instrument. The preferred sample of &amp;ldquo;retiring fishermen&amp;rdquo; (at least 60 years old, at least 15 years of experience, exiting at season&amp;rsquo;s end) has 50 individuals. A &amp;ldquo;naive&amp;rdquo; full sample has 639 fishermen; an &amp;ldquo;entering fishermen&amp;rdquo; sample (new entrants remaining at least two more seasons) has 29 individuals.&lt;/p&gt;
&lt;p&gt;Main findings: Estimating intensive (hours) and extensive (daily participation) margins via a type-2 Tobit and summing them, the preferred total IES for retiring fishermen is 2.65 (hours elasticity 0.249, participation elasticity 2.401). Across retiring-fishermen specifications, the total IES ranges roughly 2.3 to 3.1, and the headline estimate stated in the abstract and discussion is 2.7. The naive full-sample estimate is 1.27 (about 1.3), implying that accounting for LBD bias more than doubles the IES (relative bias factor about 2.1). For entering fishermen, the IES is approximately zero (-0.068). Earnings per hour are about 40% higher during a new moon than a full moon. Returns to experience are positive, significant, and plateau around 15 years.&lt;/p&gt;
&lt;p&gt;Implications: Results support using relatively large labor-supply elasticities in representative-agent macro models and provide model-free evidence that LBD matters. Because LBD breaks the equivalence of IES, Frisch, Hicks, and Marshall elasticities, a Frisch estimate no longer bounds welfare effects of tax changes, and permanent tax changes can have larger short-run labor-supply effects than transitory ones, undermining transitory tax cuts as stimulus.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-theoretical-mechanism-generating-the-bias"&gt;Q1. What is the core theoretical mechanism generating the bias?&lt;/h3&gt;
&lt;p&gt;In a life-cycle model with learning-by-doing, the wage equals the marginal return to human capital times the human-capital stock (w = w-tilde times k), and human capital grows with hours worked. The intra-temporal first-order condition shows total remuneration for an hour of work is w + F, where F is the discounted marginal increase in all future earnings from one additional hour of experience. The log-linear labor-supply equation thus contains an extra term, omega times ln(1 + F/w). Since F is non-negative and negatively correlated with the wage, omitting it (the standard model, where gh=0 so F=0) produces omitted-variable bias that pushes the estimated IES downward. The Frisch elasticity equals omega times w/(w+F), which is weakly less than omega.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q2. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Identification rests on (1) selecting fishermen for whom future returns to experience are negligible (F approximately 0), so the standard regression is unbiased, and (2) using the lunar cycle as an instrument for the wage, since catch rates and hence hourly earnings vary predictably with the moon phase but the moon plausibly does not affect tastes for or opportunity costs of work (fishermen fish in daylight, are not affected by tides, and other relevant fisheries are closed during the studied window). A type-2 Tobit (Amemiya 1984) corrects for selection because earnings and hours are observed only when fishermen participate; exclusion restrictions for the selection equation include weekend indicators, their interactions with age and age-squared, and a hurricane-preparation indicator. The main threat: that something other than returns to experience makes the samples respond differently to wage variation. Because the omitted variable is mechanical, IV cannot fix the bias in the biased samples, but it is not needed in the retiring sample where F is approximately 0.&lt;/p&gt;
&lt;h3 id="q3-how-do-they-validate-the-key-exclusion-restrictions"&gt;Q3. How do they validate the key exclusion restrictions?&lt;/h3&gt;
&lt;p&gt;For weekend indicators, prices and landings must not vary with the day of week; they regress daily lobster prices on Saturday/Sunday indicators with season and dealer fixed effects and find the coefficients extremely small and insignificant. Landings are argued independent of day-of-week because trap catch does not depend on aggregate participation. For the hurricane-preparation indicator, they regress daily prices on hurricane indicators with season and dealer fixed effects and find the hurricane-preparation coefficient very small and insignificant. Lobsters being storable/transportable and Florida supplying only 4-7% of the global annual spiny lobster catch supports price exogeneity.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-evidence-that-returns-to-experience-matter-in-this-industry"&gt;Q4. What is the evidence that returns to experience matter in this industry?&lt;/h3&gt;
&lt;p&gt;They estimate two restrictive wage specifications: one with years of experience, its square, and an indicator for having one or more years of experience; another with eighteen indicators for each experience level. Both (Figure 1) show returns to experience are positive and statistically significant, with cumulative returns plateauing around 15 years (consistent with the model&amp;rsquo;s assumption that gh approaches 0 at high human capital and with the 15-year experience criterion for retiring fishermen) and a sizable drop in marginal returns between zero and some experience.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-headline-elasticity-magnitudes"&gt;Q5. What are the headline elasticity magnitudes?&lt;/h3&gt;
&lt;p&gt;Preferred retiring sample (15+ seasons): hours elasticity 0.249 (SE 0.062), participation elasticity 2.401 (SE 0.548), total IES 2.650. The 10+ seasons retiring sample gives total IES 2.309 (smaller because returns to experience may not yet be negligible below 15 years). Across specifications retiring estimates span about 2.3 to 3.1, with 2.7 as the headline. Full (naive) sample: hours 0.046, participation 1.226, total 1.272 (about 1.3). Entering fishermen (preferred): total -0.068, i.e., approximately zero; expanded entering sample also small and insignificant. New moon earnings about 40% above full moon.&lt;/p&gt;
&lt;h3 id="q6-how-do-they-rule-out-that-sample-differences-other-than-experience-drive-the-results"&gt;Q6. How do they rule out that sample differences other than experience drive the results?&lt;/h3&gt;
&lt;p&gt;They re-estimate using a placebo sample of fishermen who meet the retiring-sample criteria (at least 60 years old, at least 15 years experience) but are at least two years from retirement, so they share age and career history but still have non-negligible returns to experience. Estimates for these older, experienced, non-retiring fishermen (Table 3) are very similar to the full sample and notably smaller than for retiring fishermen, indicating the elasticity difference is driven by returns to experience, not age or career history. They also note (footnote 27) that a flat cumulative return after 15 years is consistent with significant human-capital depreciation, so marginal returns can remain non-negligible until the final pre-retirement season.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-address-the-wage-prediction-instrument-being-estimated-separately-per-sample"&gt;Q7. What robustness checks address the wage-prediction (instrument) being estimated separately per sample?&lt;/h3&gt;
&lt;p&gt;Because estimating equation (11) separately per sample lets the moon-phase coefficient vary across samples, they run two pooled alternatives. Alternative #1 predicts earnings from the full sample of fishermen; the preferred retiring IES falls slightly (to about 2.06) because the moon coefficient is larger in absolute value, but entering-fishermen estimates stay small and insignificant. Alternative #2 pools entering and retiring fishermen in estimating (11), interacting all variables with an entering-fisherman indicator to limit selection-bias contamination; this raises retiring IES somewhat. Both confirm the cross-sample differences come from different responses to wage variation, not from different wage predictions.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-relate-to-and-differ-from-prior-structural-and-reduced-form-work"&gt;Q8. How does the paper relate to and differ from prior structural and reduced-form work?&lt;/h3&gt;
&lt;p&gt;Beginning with Imai and Keane (2004), a literature jointly estimates labor supply and human-capital accumulation in fully structural models (Imai and Keane 2004 IES 3.8; Wallenius 2011 IES 1.1; Keane and Wasi 2016 IES 2). Structural models control for wage endogeneity and allow counterfactuals but require fully specifying the wage and choice environment, are complex, and it can be unclear which moments identify the IES. This paper&amp;rsquo;s complementary, largely model-free approach exploits negligible end-of-career returns to experience, remaining agnostic about human-capital accumulation. Their estimates lie within (at the high end of) the structural range. Their relative bias (2.1) nearly matches Wallenius (2011) and is below Imai and Keane&amp;rsquo;s 8-12 (whose sample of 20-36 year-old males has high returns to experience; bias falls to 3.2 for a 20-64 simulated sample with outliers removed). The closest prior approach is Rogerson and Wallenius (2013), who infer an IES lower bound from rationalizing retirement; both approaches are robust to LBD but use very different identification.&lt;/p&gt;
&lt;h3 id="q9-what-alternative-explanations-do-they-consider-and-reject"&gt;Q9. What alternative explanations do they consider and reject?&lt;/h3&gt;
&lt;p&gt;Two. (1) Borrowing/credit constraints (Domeij and Floden 2006) also bias the IES downward and could differ across samples if retiring fishermen are less constrained; but the authors study daily decisions, and fishermen own a collateralizable vessel and almost certainly have credit or liquid assets for day-to-day purchases, so daily credit constraints are implausible. (2) Reference dependence with daily income targets and loss aversion (Camerer et al. 1997; tested by Farber 2015 on NYC taxi drivers, who also finds elasticities rising with experience): reference-dependent behavior should appear only when realized wages deviate from expected wages, but here identification comes from the perfectly predictable lunar cycle, so it cannot drive the results. The much larger participation elasticity for retiring fishermen (a decision based on anticipated wages) further argues against it; moreover Farber (2015) and Haggag, McManus and Paci (2017) find LBD in NYC taxis, so the experience-elasticity correlation there may itself reflect LBD.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Results support relatively large labor-supply elasticities in calibrated representative-agent macro models (their IES falls within aggregate hours elasticities of 1.9 to 4 reported by Chetty et al. 2011). But extrapolation to macro requires care: the IES-to-labor-supply-elasticity link is broken under LBD, and aggregate elasticities depend on long-run labor-force participation and aggregation across life-cycle stages, not the daily participation margin estimated here; a fully structural model is still needed for life-cycle and aggregate predictions. On taxes, because LBD breaks the standard ordering (IES = Frisch, Frisch &amp;gt; Hicks &amp;gt; Marshall), a Frisch estimate no longer bounds welfare effects of tax changes. Permanent tax changes can have larger short-run labor-supply effects than transitory ones (which only affect the current wage), undermining transitory tax cuts as ideal short-term stimulus; permanent changes also have amplified long-run effects because reduced current labor lowers future wages.&lt;/p&gt;
&lt;h3 id="q11-what-modeling-choices-and-caveats-accompany-the-estimates"&gt;Q11. What modeling choices and caveats accompany the estimates?&lt;/h3&gt;
&lt;p&gt;They model a daily period, so omega is the IES over hours within a working day; the total elasticity comparable to annual data is the sum of the hours elasticity (delta from the intensive-margin equation) and the daily participation elasticity (from the probit). For retiring fishermen, individual fixed effects equal individual-by-season fixed effects (each appears one season), flexibly controlling for the human-capital stock. They do not correct standard errors for the generated regressor (predicted log wage) but, citing Miles (1997) and Benito (2006), judge it unlikely to render estimates insignificant; standard errors are clustered by calendar date. A potential dynamic concern (lobsters accumulating in traps) is dismissed because catch per trap stops rising after a few days of soak time (and average soak times of 7-15 days exceed that), so daily catch depends on environmental conditions, not past fishing. The exit-date inference rule drops less than 3% of observations with virtually identical results.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Sovereign Debt Restructuring and Reduction in Debt-to-GDP Ratio</title><link>https://macropaperwarehouse.com/papers/sovereign-debt-restructuring-and-reduction-in-debt-to-gdp-ratio/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/sovereign-debt-restructuring-and-reduction-in-debt-to-gdp-ratio/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Sovereign debt restructuring is a central tool for countries in debt distress, yet surprisingly little evidence exists on whether it actually reduces the debt-to-GDP ratio — the metric used in virtually every debt sustainability analysis. This paper fills that gap. The debt-to-GDP ratio is not a simple pass-through from restructuring: the numerator (debt stock) only falls at the completion of a restructuring episode, while the denominator (GDP) can be depressed from the start of the crisis. Cash flow relief and face value reductions affect the numerator along different timelines, and fiscal consolidation — or its absence — can erode or reinforce whatever gains restructuring provides. These complexities make the net effect on the ratio genuinely non-obvious.&lt;/p&gt;
&lt;p&gt;The authors compile a novel, highly comprehensive dataset covering 709 restructuring events across 115 emerging market and developing economies from 1950 to 2021, encompassing private external creditors, Paris Club bilateral creditors, China, and domestic creditors — broader coverage than any prior study. Country-level macroeconomic data (GDP, general government debt, primary balances, inflation, exchange rates) come from the IMF World Economic Outlook October 2022 vintage. The sample excludes advanced economies, which almost never restructure (the three AE episodes — Slovenia 1992–96, Greece 2011–12, Cyprus 2013 — are dropped because the structural features of AE debt differ markedly from EMEs and LICs).&lt;/p&gt;
&lt;p&gt;Identification addresses the core problem that restructuring is endogenous to macroeconomic conditions: countries restructure precisely when growth is weak and fiscal positions are deteriorating. Following Jorda and Taylor (2016), the authors employ an Augmented Inverse Probability Weighted (AIPW) estimator. A first-stage saturated probit model estimates each country-year&amp;rsquo;s propensity score using lagged GDP growth, debt-to-GDP levels (interacted with country dummies to allow heterogeneous thresholds), primary and current account balances, US short and long interest rates, effective interest rates, and prior restructuring history. The predicted propensity scores feed a second-stage local projection of debt-to-GDP changes on the restructuring dummy and covariates across horizons 0–5 years. The AIPW is doubly robust: consistency requires only that the first stage or the second stage (not necessarily both) be correctly specified. The propensity model achieves an AUROC above 0.85.&lt;/p&gt;
&lt;p&gt;The main finding is that a typical sovereign debt restructuring event reduces the debt-to-GDP ratio by 3.8 percentage points in the first year (statistically significant), rising to a cumulative 7.2 percentage points after five years. The effect is negative and significant at every horizon from year 0 through year 5, and extends beyond five years (robustness checks to 10-year horizon show consistently negative effects, though standard errors widen with smaller samples). An important robustness check using debt level (percent change in debt stock) as the outcome shows the restructuring reduces debt by about 7 percent on impact and over 35 percent after five years — establishing that the ratio result is not mechanically driven by GDP movements alone.&lt;/p&gt;
&lt;p&gt;Heterogeneity across restructuring types and accompanying policies is substantial. When restructuring coincides with fiscal consolidation (positive average cyclically adjusted primary balance during the episode), the debt-to-GDP decline ranges from 4.7 percentage points in year 1 to 11.9 percentage points in year 5 — roughly double the average effect in the long run. Restructurings that include a face value reduction show an immediate impact of 8.9 percentage points in year 1 (versus 3.8 for the average), but the long-run effect after five years converges toward 5.0 percentage points — smaller than the fiscal consolidation pathway. Large-scale creditor coordination under the HIPC/MDRI initiatives produces ATEs of 5.4 percentage points in year 1 and 6.4 percentage points in year 5. These results collectively indicate that the long-run depth of the debt reduction is most reliably achieved when restructuring is paired with sustained fiscal effort, whereas face value reduction and creditor coordination are particularly potent in the short run.&lt;/p&gt;
&lt;p&gt;A novel finding concerns cash flow relief only (maturity extension and/or coupon rate reduction, without face value reduction): normalizing by the size of treatment (the average present-value reduction in the debt ratio, estimated at 2.8 percentage points of GDP for private external restructurings, compared to 6.0 percentage points for face value reduction events), the ATE per unit of treatment for cash flow relief converges to roughly the same magnitude as for face value reduction after four to five years. This suggests that, conditional on treatment depth, the form of restructuring does not determine long-run effectiveness — what matters is that the intervention provides sufficient fiscal space for subsequent adjustment.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses an Augmented Inverse Probability Weighted (AIPW) estimator following Jorda and Taylor (2016). The first stage is a saturated probit model predicting the propensity score for restructuring entry using: two lags of the treatment dummy, GDP growth, and change in debt-to-GDP; one lag of exchange rate change, inflation, global output gap, US short and long rates, effective interest rate, primary balance, and current account balance; and the level of debt-to-GDP interacted with country dummies (to allow heterogeneous restructuring thresholds). The second stage is a local projection of the change in debt-to-GDP regressed on the treatment dummy, its interaction with covariates, and country plus year fixed effects, across horizons 0–5. The AIPW ATE formula re-weights observed outcomes by propensity scores and adds augmentation terms from the outcome model, yielding double robustness. The main identification threat is selection-on-unobservables: countries that restructure may have systematically different unobserved growth prospects that simultaneously affect the debt ratio. The authors address one specific form of this concern — that countries and creditors time resolution to coincide with favorable growth — by including 1- and 2-year ahead IMF GDP forecasts as controls in a robustness check, finding similar results. Observations with propensity scores outside [10^-4, 1−10^-4] are excluded to avoid extreme weight instability. Significant overlap between treatment and control propensity score distributions (both approaching full support in [0,1]) is verified.&lt;/p&gt;
&lt;h3 id="q2-why-is-the-timing-of-restructuring-start-vs-end-relevant-for-the-debt-ratio"&gt;Q2. Why is the timing of restructuring start (vs. end) relevant for the debt ratio?&lt;/h3&gt;
&lt;p&gt;Prior papers (Reinhart and Trebesch 2016; Cheng et al. 2019) measure the impact from the end of the restructuring episode or the resolution of the debt crisis. This paper instead measures from the start of the restructuring event (the onset of debt crisis). The distinction matters because: (i) the debt stock is only formally reduced at the completion of restructuring (once a deal is struck and recorded), so the numerator of the debt ratio moves discontinuously at the end of the episode; (ii) GDP, however, can be negatively affected from the outset of the crisis, compressing the denominator before any debt relief is delivered. About one-third of restructuring episodes last two or more years, so the distinction is empirically non-trivial. Measuring from the start captures the full dynamic path — including the initial GDP drag and the later debt relief — without conditioning on crisis resolution, which could itself be endogenous.&lt;/p&gt;
&lt;h3 id="q3-what-does-the-dataset-cover-and-how-does-it-differ-from-prior-work"&gt;Q3. What does the dataset cover and how does it differ from prior work?&lt;/h3&gt;
&lt;p&gt;The dataset covers 709 restructuring events in 115 emerging market and developing countries from 1950 to 2021. It includes four creditor classes: private external creditors (sourced from Asonuma and Trebesch 2016), official bilateral external creditors under the Paris Club (from Paris Club database and Horn et al. 2022), official bilateral creditors outside the Paris Club including China (from Horn et al. 2022), and domestic creditors (from IMF 2021). The paper also covers restructurings that occur outside sovereign defaults, including preemptive restructurings where payments are not missed. Prior literature focused primarily on post-default restructurings with external private or Paris Club creditors. The 310 EM restructuring events break down as 85.8% cash flow relief only and 14.2% face value reduction; 58.4% are preemptive, 21.6% post-default, and 20% both or unidentified. For LICs, 396 events are recorded, with 73.5% cash flow relief only and 26.5% face value reduction. Macroeconomic controls come from the IMF WEO October 2022 vintage.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-propensity-models-predictive-performance-and-what-does-it-reveal-about-the-determinants-of-restructuring"&gt;Q4. What is the propensity model&amp;rsquo;s predictive performance, and what does it reveal about the determinants of restructuring?&lt;/h3&gt;
&lt;p&gt;The first-stage probit achieves an AUROC above 0.85 and a pseudo R-squared of 0.295 on 1,233 observations. Key findings: the lagged treatment dummy is negative and significant (countries that recently restructured are less likely to do so again soon, possibly because creditors resist multiple sequential restructurings); lagged changes in debt-to-GDP are negative in the two years preceding restructuring (reflecting that countries often pursue fiscal consolidation before resorting to restructuring as a last resort); global output gap and GDP growth have the expected signs (restructurings more likely when global conditions are favorable and domestic growth is low), though p-values are near 0.10; US interest rate coefficients have opposite signs for short vs. long rates and are statistically insignificant. The propensity score distributions show significant overlap between treatment and control groups, supporting the common support assumption.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-ate-per-unit-of-treatment-analysis-reveal-about-cash-flow-relief-vs-face-value-reduction"&gt;Q5. What does the ATE per unit of treatment analysis reveal about cash flow relief vs. face value reduction?&lt;/h3&gt;
&lt;p&gt;The ATE per unit of treatment is constructed by dividing the estimated ATE by the average size of treatment. For face value reduction events, the size is the average annual face-value-reduction-to-GDP ratio, approximately 6.0 percentage points. For cash flow relief only events (restricted to private external restructurings where present-value data are available from Asonuma et al. 2023), the size is estimated using a back-of-envelope calculation scaling the FVR size by the ratio of present-value debt reduction for cash flow relief (5 percent) to that for FVR (10.6 percent), yielding 2.8 percentage points. Table 4 shows: for FVR, the ATE in year 0 is -10.6 pp (per unit: -1.77), falling to -5.0 pp in year 5 (per unit: -0.83) — a frontloaded and then diminishing profile. For cash flow relief, the ATE is +3.6 pp in year 0 (per unit: +1.29), moving to -5.7 pp in year 5 (per unit: -2.04) — a monotonically increasing profile. The per-unit effects converge by around year 4, supporting the conclusion that treatment depth rather than treatment type is what determines long-run effectiveness.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-interaction-between-restructuring-and-fiscal-consolidation-defined-and-what-does-the-heterogeneity-analysis-show"&gt;Q6. How is the interaction between restructuring and fiscal consolidation defined and what does the heterogeneity analysis show?&lt;/h3&gt;
&lt;p&gt;Fiscal consolidation is defined as a positive average cyclically adjusted primary balance during the duration of the restructuring episode. The AIPW model is re-estimated using only the subset of restructuring events meeting this criterion as the treatment group, while keeping all non-restructuring observations as the control group. The estimated ATE ranges from 4.7 percentage points in year 1 to 11.9 percentage points in year 5 — substantially exceeding the 3.8 and 7.2 pp average effects. The long-run amplification relative to the average is larger than the short-run amplification, underscoring that sustained fiscal effort is the dominant factor in durable debt ratio reduction. A robustness check using a weaker definition of fiscal consolidation (positive year-on-year change in the cyclically adjusted primary balance, which can still leave the primary balance negative) shows a larger initial impact but a declining cumulative effect after a few years, consistent with the interpretation that only episodes maintaining a positive (not just improving) fiscal stance sustain the gain.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-heterogeneity-analysis-show-for-creditor-coordination-hipcmdri-versus-the-average"&gt;Q7. What does the heterogeneity analysis show for creditor coordination (HIPC/MDRI) versus the average?&lt;/h3&gt;
&lt;p&gt;Restricting the treatment group to restructuring events under the Heavily Indebted Poor Country Initiative and the Multilateral Debt Relief Initiative, the paper finds ATEs of 5.4 percentage points in year 1 and 6.4 percentage points in year 5. Both exceed the average effects (3.8 and 7.2 pp, respectively) in year 1, though the five-year effect is slightly smaller than the average (6.4 vs. 7.2 pp). The authors contrast this with Easterly (2002), who argued that HIPC countries remained heavily indebted even after two decades of debt relief and concessional financing (1980–1997). The paper&amp;rsquo;s result suggests that more comprehensive HIPC/MDRI programs produce meaningful and durable reductions in the debt ratio, at least within the five-year window studied.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-analysis-imply-about-gdp-dynamics-during-restructuring"&gt;Q8. What does the analysis imply about GDP dynamics during restructuring?&lt;/h3&gt;
&lt;p&gt;The paper establishes that debt levels fall more in percentage terms than the debt ratio does. In the baseline, the average debt-to-GDP ratio falls 3.8 pp in year 1 while the debt level falls about 7 percent in year 1. A back-of-the-envelope calculation (holding the average debt ratio at roughly 1, so the ratio change approximately equals the percent change in debt minus the percent change in GDP) implies that GDP falls by roughly 3.8 percent after one year of restructuring relative to the year prior, after controlling for selection. Over five years, the debt level falls over 35 percent while the debt ratio falls 7.2 pp, implying cumulative GDP losses that moderate the ratio improvement. The authors confirm this via a robustness check using GDP forecasts as additional controls, finding similar results to the baseline.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-performed-and-what-do-they-show"&gt;Q9. What robustness checks are performed and what do they show?&lt;/h3&gt;
&lt;p&gt;Six main robustness checks are reported: (1) Extending the horizon from 5 to 10 years — effects remain negative throughout, though standard errors widen due to smaller samples. (2) Using the change in debt level (percent) as the outcome instead of the change in the debt ratio — the restructuring reduces debt by about 7 percent on impact and over 35 percent after 5 years, confirming the ratio result is not purely a GDP-denominator artifact. (3) Including 1- and 2-year ahead IMF GDP forecasts as additional controls — results are similar to baseline. (4) Removing interaction terms between the treatment dummy and covariates from equation (1) — results are similar to baseline. (5) Comparing AIPW ATE to a plain OLS local projection (setting the ATE equal to the coefficient on the treatment dummy, without AIPW weighting) — the AIPW attenuates the estimated impact compared to OLS, as expected given upward selection bias: countries in worse shape are more likely to restructure, so naive estimates understate the baseline counterfactual. (6) Alternative probit subsetting for FVR events: removing top/bottom 10% of FVR-to-GDP from the treatment group (to address outliers) produces robust results; alternatively, using the predicted probability of FVR occurrence (based on pre-restructuring information only) to define treatment group membership yields similar findings.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-and-differ-from-prior-work-on-debt-restructuring-and-debt-ratios"&gt;Q10. How does this paper relate to and differ from prior work on debt restructuring and debt ratios?&lt;/h3&gt;
&lt;p&gt;The closest prior papers are Reinhart and Trebesch (2016) and Cheng et al. (2019). Reinhart and Trebesch compare simple pre/post means across 18 AEs (1920–1939) and 35 EMs (1978–2010) — limited by small samples, no causal identification, focus on private external creditors, and measurement from the end of the restructuring episode. Cheng et al. study 93 EMs and LICs (1956–2015) using local projections but cover only Paris Club official creditors and focus on the end of the crisis. The present paper adds: coverage of 115 countries over 1950–2021; a broader set of creditors (private, Paris Club, China, domestic); timing from the start rather than the end of the episode; causal identification via AIPW; and heterogeneity analysis across fiscal consolidation, face value reduction, creditor coordination, and treatment size. The finding that cash flow relief per unit of treatment converges to face value reduction in the long run is novel; prior literature mostly emphasized nominal haircuts. The positive result for HIPC/MDRI also directly contradicts Easterly (2002).&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The key policy implication is that debt restructuring is an effective tool for reducing debt ratios in EMEs and LICs — this is not automatic or mechanical, as GDP effects partially offset the debt stock relief, yet the net effect on the ratio is statistically significant and long-lasting. Scope conditions: (i) The results apply to emerging market economies and low-income countries; advanced economies rarely restructure and the three AE episodes in the sample are excluded as structurally different. (ii) The effectiveness is substantially amplified when restructuring is accompanied by sustained fiscal consolidation (positive average cyclically adjusted primary balance), implying that restructuring alone, without accompanying fiscal effort, provides a smaller and less durable reduction. (iii) Face value reduction is more potent in the short run but converges to cash flow relief in the long run (per unit of treatment), suggesting that deep rescheduling without nominal haircuts can be comparably effective as long as it provides sufficient fiscal space. (iv) The HIPC/MDRI creditor coordination framework is associated with larger-than-average impacts. (v) Preemptive restructurings (without outright default) are included and common, suggesting the results are not limited to post-default episodes. The paper informs current IMF and policymaker discussions on how to manage the post-COVID sovereign debt overhang.&lt;/p&gt;
&lt;h3 id="q12-what-stylized-facts-characterize-the-types-of-restructuring-in-the-dataset"&gt;Q12. What stylized facts characterize the types of restructuring in the dataset?&lt;/h3&gt;
&lt;p&gt;Based on Table 2: among EMs, 85.8% of restructurings involve cash flow relief only (no face value reduction) and 14.2% involve face value reduction; 58.4% are preemptive, 21.6% post-default. The most common creditor type in EMs is private external (54.8%), followed by Paris Club (48.1%). Among LICs, 73.5% involve cash flow relief only and 26.5% face value reduction; 54.3% are preemptive and 31.1% post-default; Paris Club is dominant (73.5%). Domestic debt restructurings are rare across both groups; when they occur, they tend to involve smaller face value reductions than external restructurings. The paper also notes that 60% of restructuring events are preceded by an increase in the primary-balance-to-GDP ratio, indicating fiscal effort before crisis resolution is common.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Augmented Inverse Probability Weighted (AIPW) Estimator&lt;/strong&gt;: A two-stage causal estimator that first models the propensity score (probability of treatment) and then uses it to re-weight observed outcomes in a local projection, with an augmentation term from the predicted outcome model. It is doubly robust: the average treatment effect is consistently estimated if either the propensity model or the outcome model is correctly specified, but not necessarily both.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Face Value Reduction (FVR)&lt;/strong&gt;: A cut in the nominal (principal) amount of the outstanding debt instruments, also called a nominal haircut. In the paper, the average FVR-to-GDP ratio during restructuring events with FVR is approximately 6 percent per year. FVR events constitute 14.2% of EM restructurings and 26.5% of LIC restructurings in the dataset.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cash Flow Relief&lt;/strong&gt;: Debt rescheduling without reduction in face value — encompassing maturity extension and/or coupon rate reduction — that alters the stream of future payments without changing the nominal amount owed. This is the predominant form of restructuring (85.8% of EM events). The present-value size of treatment for cash flow relief is estimated at 2.8 pp of GDP for private external restructurings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Average Treatment Effect (ATE) per Unit of Treatment&lt;/strong&gt;: The estimated ATE divided by the average size of the treatment (e.g., face-value-reduction-to-GDP for FVR events, or estimated present-value reduction for cash flow relief events). Used to compare the effectiveness of different restructuring modalities on a common scale, revealing that FVR has a larger per-unit impact in the short run but converges to cash flow relief by year 4–5.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Preemptive Restructuring&lt;/strong&gt;: A restructuring implemented before any missed payments occur (no legal default), or with only briefly missed payments over a short window after negotiations begin, without a unilateral default. Distinguished from post-default restructurings, which involve unilateral cessation of payments prior to any creditor agreement. Preemptive restructurings account for 58.4% of EM events and 54.3% of LIC events in the dataset.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Doubly Robust Estimator&lt;/strong&gt;: In the paper&amp;rsquo;s context, an estimator (the AIPW) whose consistency holds as long as at least one of its two component models — the propensity score model (first stage) or the outcome model (second stage) — is correctly specified. This provides a safeguard against misspecification in one stage, unlike single-model approaches such as simple IPW or plain OLS local projections.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HIPC/MDRI Creditor Coordination&lt;/strong&gt;: The Heavily Indebted Poor Country Initiative and the Multilateral Debt Relief Initiative, which provide structured large-scale debt relief programs with coordinated participation by multiple official creditors. In the paper, restructuring events under HIPC/MDRI constitute a treatment subgroup showing ATEs of 5.4 pp (year 1) and 6.4 pp (year 5), exceeding the average year-1 effect but roughly in line with the average year-5 effect.&lt;/p&gt;</description></item><item><title>Strapped for Cash: The Role of Financial Constraints for Innovating Firms, Misallocation and Aggregate Productivity Growth</title><link>https://macropaperwarehouse.com/papers/strapped-for-cash-the-role-of-financial-constraints-for-innovating-firms-misallocation-and-aggregate-productivity-growth/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/strapped-for-cash-the-role-of-financial-constraints-for-innovating-firms-misallocation-and-aggregate-productivity-growth/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Firms that invest heavily in intangible assets — patents, R&amp;amp;D, software — face a structural financing disadvantage: intangibles offer limited collateral value to banks, so intangible-intensive firms can be cut off from credit even when their marginal revenue product of capital (MRPK) exceeds the going interest rate. The paper asks how binding this collateral constraint is in practice, what relaxing it does to firm behavior, and how large the aggregate productivity and misallocation consequences are.&lt;/p&gt;
&lt;p&gt;The empirical setting is a 2015 Norwegian legal reform that, for the first time, allowed firms to pledge patents as stand-alone collateral. Before the reform, a patent could serve as collateral only in conjunction with a physical asset or if it was actively generating revenue; the reform removed both conditions as of 1 July 2015. The change was introduced specifically to ease financing for innovative firms and was narrow in scope — not part of a broader financial reform.&lt;/p&gt;
&lt;p&gt;The empirical analysis draws on matched administrative panel data covering the universe of Norwegian private non-financial joint-stock companies (about 85 percent of all firms with employees) over 2005–2018. The five linked data sets provide annual firm accounts, loan-level bank lending records (firm-bank-year), shareholder and equity issuance records, and the universe of patent applications to the Norwegian Patent Office. The pre-reform window runs 2010–2015; the post-reform window 2015–2018; the 2005–2010 period is used for placebo tests.&lt;/p&gt;
&lt;p&gt;The identification strategy is difference-in-differences. The treatment group consists of firms with at least one patent application in the five years before the reform (2010–2015); the control group consists of firms without a patent portfolio but with similar observable characteristics (size, tangible assets, intangible intensity, profitability, public-funding status), all within the same 2-digit NACE industry. Firm fixed effects and industry-by-year fixed effects are included throughout; control variables are measured pre-reform and interacted with year dummies.&lt;/p&gt;
&lt;p&gt;Firm-level results confirm that treated firms were collateral constrained: (i) the probability of having a bank loan rose by 5.1 percentage points; (ii) the bank debt-to-sales ratio rose by 1.5 percentage points; (iii) the share of short-term debt fell by 2.7 percentage points, consistent with conversion to longer-term collateralized debt; (iv) the number of bank connections rose by 0.144; and (v) the interest rate was unchanged. Simultaneously, the capital stock (total fixed assets) rose by 0.20 log points, employment rose by 0.051 log points, and MRPK fell significantly (–0.224), satisfying the necessary and sufficient conditions for collateral constraint under the theoretical framework. Sales showed no significant change, which the authors attribute to the short post-reform window (only three years). Pre-trend tests using placebo reform years (2010) and pre-2010 periods yield insignificant estimates, supporting parallel trends.&lt;/p&gt;
&lt;p&gt;For young firms (six years old or younger in 2015), there are additional effects: a larger employment response (+0.181 log points for the interaction term) and positive effects on equity issuance (the equity issue dummy rises by 0.137 for young treated firms) and number of shareholders (+0.225 log points). The improvement in debt access appears to have signaled creditworthiness and improved terms of access to equity for young firms. Innovation also rose: the probability of filing at least one patent in 2016–2018 increased by 21.7 percentage points for treated firms relative to the control group, and the count of patent applications increased by 0.936.&lt;/p&gt;
&lt;p&gt;For aggregate quantification, the authors develop a model of monopolistic competition with heterogeneous firms and credit constraints (following Hsieh and Klenow, 2009 and Melitz, 2003). Each constrained firm faces an implicit capital cost of τ times the market interest rate, where τ ≥ 1. The model is solved in changes using exact hat algebra. Under the small-open-economy assumption (capital supply infinitely elastic), removing the constraint raises labor productivity through two channels: (1) reduced within-industry misallocation as firms equalize MRPKs, and (2) capital deepening as constrained firms invest more. The key advantage of the methodology is that the friction τ is identified directly from the DiD capital stock estimate (0.20 log points) combined with observed capital shares (mean α = 0.30) and an elasticity of substitution σ = 4 (from Broda and Weinstein, 2006), sidestepping the need to estimate revenue TFP.&lt;/p&gt;
&lt;p&gt;The median treated firm faces a credit friction of τ = 1.12, implying an implicit capital cost 12 percent above the market rate. Industry output per worker increases by up to 3 percent, concentrated in sectors where treated (innovative) firms hold a large initial market share. The dominant source of this gain is capital deepening: the ratio of economy-wide labor productivity growth to TFP growth is 39:1, meaning within-industry misallocation reduction accounts for only a small fraction of the productivity gain. The aggregate price index falls by 0.6 percent (P-hat = 1.006 in output-per-worker terms), translating to an increase in total output of 6.4 billion NOK (approximately 0.62 billion USD). A back-of-the-envelope calculation using the implicit cost r(τ-1)K yields 7.5 billion NOK, consistent with the model estimate. For comparison, Norway&amp;rsquo;s main innovation subsidy agency disbursed 5.3 billion NOK in 2021, putting the collateral reform&amp;rsquo;s welfare gain in the same order of magnitude.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses difference-in-differences: the treatment group is firms with at least one patent application in 2010–2015; the control group is all other firms matched on size, tangible assets, intangible intensity, profitability, and public-funding status within the same 2-digit NACE industry. Identification requires parallel trends in the absence of the reform. Three tests are conducted: (1) visual inspection of pre-reform trends in the bank loan dummy after residualizing on controls and fixed effects shows broadly similar trajectories; (2) a placebo regression using 2010 as the fake reform year over 2005–2015 yields insignificant coefficients across most credit access measures; (3) a second placebo uses the same 2010–2015 treatment group but compares the pre-2010 period against 2010–2015, again finding insignificant pre-trends. A residual threat is that treated and control firms may differ in unobservable ways that generate differential post-2015 trends unrelated to the reform. The authors address this by conditioning on a rich set of pre-reform firm characteristics interacted with year dummies, but general equilibrium spillovers (e.g., control firms affected by increased competition from treated firms) mean the DiD cannot cleanly capture the aggregate effect, which is why the structural model is needed.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-authors-establish-that-observed-effects-reflect-collateral-constraints-rather-than-mere-debt-substitution"&gt;Q2. How do the authors establish that observed effects reflect collateral constraints rather than mere debt substitution?&lt;/h3&gt;
&lt;p&gt;The theoretical framework makes a sharp prediction: if a firm is unconstrained, an increase in available funding will leave the capital stock and MRPK unchanged (the firm simply substitutes between funding sources). Only a constrained firm will simultaneously (i) increase borrowing, (ii) increase the capital stock, and (iii) show a decline in MRPK as capital is brought closer to its optimal level. The paper documents all three outcomes for treated firms — 5 pp higher probability of bank debt, 0.20 log-point higher capital, and –0.224 significant decline in MRPK — satisfying the necessary and sufficient conditions for collateral constraint. The unchanged interest rate rules out credit becoming cheaper as a confound.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-two-channels-through-which-removing-collateral-constraints-raises-aggregate-productivity-and-how-large-is-each"&gt;Q3. What are the two channels through which removing collateral constraints raises aggregate productivity, and how large is each?&lt;/h3&gt;
&lt;p&gt;The model decomposes industry labor productivity growth (Ys-hat/Ls-hat) into two multiplicative components: (1) TFP growth (TFPs-hat) reflecting reduced within-industry misallocation as capital is reallocated toward previously constrained firms with high MRPK, and (2) capital deepening (Ks-hat/Ls-hat)^alpha reflecting an increase in the aggregate capital-labor ratio as constrained firms invest more. Quantitatively, capital deepening dominates: economy-wide labor productivity growth is 39 times larger than TFP growth. This is because Norway is treated as a small open economy where capital supply is elastic at a fixed world interest rate, so aggregate capital expands substantially when constraints are removed. Under the alternative closed-economy assumption (capital supply fixed, interest rate endogenous), capital deepening would be muted and misallocation reduction would play a larger relative role.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-by-firm-age-is-documented-and-why-does-it-arise"&gt;Q4. What heterogeneity by firm age is documented, and why does it arise?&lt;/h3&gt;
&lt;p&gt;Young firms (six years old or younger in 2015) show larger employment responses (the triple interaction P_t x P_i x Young_i is 0.181, significant at 5%) and are the primary drivers of the shift from short-term to long-term debt (triple interaction –0.114, significant at 1%). Young treated firms also gain more in equity access: equity issuance probability rises by 0.137 (significant at 1%) and number of shareholders rises by 0.225 log points (significant at 10%) compared to older treated firms. The authors argue that for young firms the collateral constraint is more binding — consistent with the broader literature — and that improved bank access signals creditworthiness to equity investors, alleviating information asymmetries. For innovation outcomes, there is no strong differential effect by age.&lt;/p&gt;
&lt;h3 id="q5-how-is-the-structural-credit-friction-τ-identified-from-the-reduced-form-estimates"&gt;Q5. How is the structural credit friction τ identified from the reduced-form estimates?&lt;/h3&gt;
&lt;p&gt;From the structural model, the capital stock of a treated firm changes relative to a control firm as K-hat_si = τ^[α_s(σ-1)+1] x P-hat_s^(σ-1). Inverting this expression (Proposition 1 in the paper) yields τ as a function of the observed capital growth K-hat (from the DiD estimate of 0.20 log points), the capital share α_s (measured from the data as 1 minus wage costs over total costs, mean 0.30), and the elasticity of substitution σ (set to 4 from Broda and Weinstein, 2006). Because the DiD estimate is well-identified from a quasi-natural experiment, τ is identified directly from causal variation rather than from cross-sectional dispersion in MRPK as in the traditional misallocation literature (Hsieh-Klenow). This avoids the measurement error and production function estimation problems inherent in that approach.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-distribution-of-the-credit-friction-τ-across-treated-firms"&gt;Q6. What is the distribution of the credit friction τ across treated firms?&lt;/h3&gt;
&lt;p&gt;Since τ in Proposition 1 varies only with the industry capital share α_s (the other inputs — the DiD estimate and σ — are uniform), variation in τ across firms is entirely driven by cross-industry variation in α_s. The density of τ is concentrated between roughly 1.06 and 1.14. The median treated firm has τ = 1.12, implying an implicit capital cost 12 percent above the market interest rate.&lt;/p&gt;
&lt;h3 id="q7-how-are-aggregate-gains-computed-and-how-large-are-they"&gt;Q7. How are aggregate gains computed and how large are they?&lt;/h3&gt;
&lt;p&gt;The aggregate output gain is computed as 1 minus the aggregate price index P-hat. Using initial expenditure shares β_s and the industry price indices from equation (5), the authors obtain P-hat = 1.006 — a 0.6 percent fall in the aggregate price level, equivalently a 0.6 percent rise in output per worker and real wages. Multiplied by aggregate value added in the data, this yields 6.4 billion NOK (approximately 0.62 billion USD). A separate back-of-the-envelope calculation using the formula r(τ-1)K — the total implicit cost of the constraint — gives 7.5 billion NOK (approximately 0.73 billion USD), with median r = 0.07 and median τ = 1.12. The proximity of the two estimates is offered as a consistency check. These gains accrue over the three post-reform years (2015–2018) and are described as substantial, comparable in magnitude to Norway&amp;rsquo;s main innovation subsidy program (5.3 billion NOK in 2021).&lt;/p&gt;
&lt;h3 id="q8-what-does-the-paper-find-regarding-the-impact-on-innovation-and-why-is-the-innovation-regression-different-from-the-other-regressions"&gt;Q8. What does the paper find regarding the impact on innovation, and why is the innovation regression different from the other regressions?&lt;/h3&gt;
&lt;p&gt;Post-reform innovation (2016–2018) is measured using a patent dummy (equals 1 if the firm files at least one application) and a patent count. The paper finds a 21.7 percentage point increase in the patent dummy and a 0.936 increase in the patent count for treated firms. These regressions are cross-sectional (estimated on the 2015 cross-section) rather than panel DiD, because using patenting pre-reform to define treatment and then examining patenting post-reform as an outcome would create a mechanical correlation. There is no strong age heterogeneity in the innovation response (the interaction with Young is negative for patent count at –0.469, marginally significant, but the patent dummy interaction is insignificant at 0.054).&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-differ-methodologically-from-the-standard-hsieh-klenow-misallocation-approach"&gt;Q9. How does this paper differ methodologically from the standard Hsieh-Klenow misallocation approach?&lt;/h3&gt;
&lt;p&gt;Hsieh and Klenow (2009) infer capital misallocation from cross-sectional dispersion in MRPK across firms, computed from observed factor shares and revenue. This approach requires estimating production functions and is subject to measurement error in capital stock and revenue TFP. The present paper instead identifies the credit friction τ from a quasi-natural experiment (the DiD capital growth estimate), which directly measures the within-sector relative capital response for constrained firms. This sidesteps production function estimation, avoids TFPR measurement issues, and produces a transparent mapping from reduced-form estimates to model primitives. The trade-off is that results are specific to the type of friction being studied (collateral constraints on intangible-intensive firms) rather than summarizing aggregate misallocation.&lt;/p&gt;
&lt;h3 id="q10-what-capital-market-assumption-is-used-in-the-baseline-and-what-is-the-alternative"&gt;Q10. What capital market assumption is used in the baseline, and what is the alternative?&lt;/h3&gt;
&lt;p&gt;The baseline assumes that Norway is a small open economy with an infinitely elastic capital supply at a fixed world interest rate r (exogenous r). Under this assumption, relaxing constraints allows constrained firms to expand their capital stock without crowding out capital from unconstrained firms, generating large capital-deepening gains. The appendix solves the model under the alternative closed-economy assumption where aggregate capital supply is fixed and the interest rate adjusts endogenously. Under the closed-economy assumption, capital deepening is muted (constrained firms can expand only at the expense of unconstrained ones), and the misallocation reduction channel plays a larger relative role. The authors argue the small open economy assumption is more appropriate for Norway.&lt;/p&gt;
&lt;h3 id="q11-what-complementarities-between-debt-and-equity-funding-are-documented-and-what-mechanism-is-proposed"&gt;Q11. What complementarities between debt and equity funding are documented, and what mechanism is proposed?&lt;/h3&gt;
&lt;p&gt;For young treated firms, improved access to bank debt (pledging patents as collateral) is associated with a higher probability of equity issuance (coefficient 0.137) and more shareholders (0.225 log points). The proposed mechanism has two parts: (1) the investment financed by bank loans improves firm profitability and return on equity, attracting investors; (2) obtaining a bank loan credibly signals firm quality to equity investors who face information asymmetries about intangible-intensive firms, facilitating equity access that would not have occurred without the debt catalyst. This complementarity is concentrated in young firms, consistent with information asymmetries being most severe early in the firm life cycle.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-paper-find-about-the-funding-structure-beyond-total-borrowing"&gt;Q12. What does the paper find about the funding structure beyond total borrowing?&lt;/h3&gt;
&lt;p&gt;Beyond the extensive margin (probability of having bank debt, +5.1 pp) and intensive margin (bank debt-to-sales ratio, +1.5 pp), the paper documents a shift in debt maturity: the share of short-term debt in total debt falls by 2.7 percentage points. This is interpreted as firms converting short-term unsecured debt into long-term debt backed by patent collateral. The number of bank connections also rises by 0.144, indicating that treated firms gained access to additional lenders (credit lines) after the reform. The interest rate on bank debt shows no significant change, ruling out a price effect — the reform operated through quantity of credit rather than its cost.&lt;/p&gt;
&lt;h3 id="q13-how-does-this-paper-relate-to-the-broader-intangible-capital-finance-literature"&gt;Q13. How does this paper relate to the broader intangible-capital finance literature?&lt;/h3&gt;
&lt;p&gt;Mann (2018) studies the US, where patent pledging is already common, and finds that strengthened creditor rights over patents raise debt and innovation. Hochberg et al. (2018) show that thicker secondary markets for patents improve debt access. Farre-Mensa et al. (2020) find that getting a patent granted raises the probability of a patent-backed loan. Falato et al. (2022) show that rising intangible intensity explains the trend decline in US corporate debt capacity. Brown et al. (2009) document the importance of financial constraints for R&amp;amp;D financing among young US firms. The present paper differs by: (a) using a reform-based quasi-experiment rather than exploiting existing cross-sectional variation; (b) covering the universe of firms including startups rather than only listed or patent-filing firms; (c) quantifying the aggregate implications for misallocation and growth, which prior work does not; and (d) documenting complementarities with equity funding and innovation.&lt;/p&gt;
&lt;h3 id="q14-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q14. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The main policy implication is that legal reform to improve the pledgeability of intangible assets — specifically patents — can substantially ease financing constraints for innovative firms, with economy-wide productivity gains of comparable magnitude to direct innovation subsidies. The scope conditions are: (1) gains are concentrated in sectors where innovative, intangible-intensive firms hold large initial market shares; (2) the capital-deepening channel — which dominates — requires an elastic capital supply, making the results most directly applicable to small open economies integrated into global capital markets; (3) the reform&amp;rsquo;s effectiveness depended on the prior absence of patent collateral rights (Norway was late relative to other OECD countries where 38% of patenting US firms had already pledged patents by 2013); (4) the short post-reform observation window (three years) may understate long-run effects on sales and productivity, since capital investment takes time to translate into revenue. The results underscore the importance of financial regulation — beyond direct subsidy programs — as a tool for promoting innovation and growth.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Collateral constraint&lt;/strong&gt;: In this paper&amp;rsquo;s framework, a firm is collateral constrained if it holds less capital than it would choose at the interest rate it currently pays — formally K_si &amp;lt; K*_si — because limited pledgeable collateral restricts its access to bank credit. The constraint is parameterized as an implicit capital cost markup τ ≥ 1 above the market rate r, so the firm equates MRPK to τr rather than r.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stand-alone patent collateral&lt;/strong&gt;: The legal status introduced by Norway&amp;rsquo;s 2015 reform under which a firm can pledge patents as collateral independently of any physical asset and regardless of whether the patent is generating current revenue. Before the reform, Norwegian law required patents to be bundled with physical assets or actively used in production before they could serve as collateral.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implicit capital cost (τ)&lt;/strong&gt;: The paper&amp;rsquo;s measure of the severity of a firm&amp;rsquo;s credit constraint: the ratio of the firm&amp;rsquo;s effective cost of capital (MRPK) to the market interest rate r. A firm with τ = 1 is unconstrained (MRPK = r); τ &amp;gt; 1 implies the firm would invest more if it could obtain capital at the prevailing rate. The median treated firm has τ = 1.12, meaning a 12% implicit cost premium.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital deepening (as a source of productivity growth)&lt;/strong&gt;: In the model, removing credit constraints allows previously constrained firms to expand their capital stock, raising the aggregate capital-to-labor ratio without proportionally reducing unconstrained firms&amp;rsquo; capital (under elastic capital supply). This increase in capital intensity per worker raises labor productivity independently of any improvement in allocative efficiency or TFP.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Within-industry misallocation (TFP_s)&lt;/strong&gt;: Following Hsieh and Klenow (2009), the paper defines industry-level TFP as the efficiency loss from heterogeneous MRPKs across firms within a sector. When firms face different implicit capital costs (τ_si), capital is misallocated: some firms use too little capital relative to their productivity. Removing constraints equalizes MRPKs and raises TFP_s, but in the paper&amp;rsquo;s quantitative results this channel is small relative to capital deepening.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pledgeability of intangible assets&lt;/strong&gt;: The extent to which a firm&amp;rsquo;s intangible assets (patents, R&amp;amp;D, goodwill, licenses) can be legally accepted as collateral for bank loans. The paper treats low pledgeability as a market friction specific to intangible-intensive firms — distinct from general credit risk — that results in those firms being systematically credit rationed even when their MRPK exceeds the interest rate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exact hat algebra&lt;/strong&gt;: A solution method due to Dekle, Eaton, and Kortum (2008) in which the model is solved entirely in terms of relative changes (hat variables, e.g., x-hat = x&amp;rsquo;/x) using observed pre-reform values in place of calibrated level parameters. This approach avoids the need to estimate unobservable structural parameters and is used here to compute counterfactual industry and aggregate outcomes after the credit friction is removed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Debt–equity complementarity&lt;/strong&gt;: The paper&amp;rsquo;s term for the finding that improved access to bank debt (via patent collateral) also raises equity issuance and the number of shareholders, especially for young firms. The proposed mechanism is that new bank loans signal creditworthiness to equity investors who face information asymmetries about intangible-intensive firms, making debt and equity complements rather than substitutes in the financing of innovative young firms.&lt;/p&gt;</description></item><item><title>Technology Sophistication Across Establishments</title><link>https://macropaperwarehouse.com/papers/technology-sophistication-across-establishments/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/technology-sophistication-across-establishments/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: How sophisticated are the technologies establishments actually use, and how close are they to the world frontier? Traditional measures (since Ryan-Gross 1943 and Griliches 1957) characterize technology by the presence of one or a few advanced technologies, which (i) cover too few technologies and unrepresentative tasks, (ii) say nothing about how non-adopters produce or how far they are from the frontier, and (iii) ignore the intensity with which a technology is used. The authors argue intensity of use matters for explaining income divergence (Comin-Mestieri 2018), so they build a direct, comprehensive measure of technology sophistication.&lt;/p&gt;
&lt;p&gt;Data and design: The authors construct &amp;ldquo;the grid,&amp;rdquo; a two-dimensional structure with business functions (BF) on the horizontal axis and technologies ranked by sophistication (simplest to world frontier) on the vertical axis. The grid spans 63 business functions (7 general business functions [GBF] relevant to all sectors plus 56 sector-specific business functions [SSBF] across 12 sectors) and a total of 305 technologies. More than 50 industry experts built and ranked the grid before survey administration. The grid is implemented in the Firm Adoption of Technology (FAT) survey, fielded 2019-2023 to 21,055 randomly selected establishments forming nationally representative samples (for establishments with 5+ workers) in 15 countries spanning all income levels (Korea, Poland, Croatia, Chile, Brazil-Ceara, Georgia, Vietnam, four Indian states, Ghana, Bangladesh, Kenya, Cambodia, Senegal, Ethiopia, Burkina Faso), representing a universe of about 2.1 million establishments. The median establishment has 9 workers (mean 34); 20% of workers hold a college degree, 17% are exporters, 18% are multinational-affiliated. FAT records, per BF, which grid technologies are used and which one is &amp;ldquo;most widely used.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Two measures are built at the BF-establishment level on a [1,5] affine scale: MAX (sophistication of the most advanced technology used, reflecting adoption) and MOST (sophistication of the most widely used technology, reflecting both adoption and intensity/diffusion within the firm). Establishment-level measures are simple averages across in-house BFs. Cardinalization is validated three ways: linearity of the sophistication-productivity relationship; correlation above 0.98 with a z-score cardinalization (Bloom-Van Reenen 2007); and median correlation 0.95 with an independent productivity-based (&amp;ldquo;Q&amp;rdquo;) cardinalization for 18 BFs.&lt;/p&gt;
&lt;p&gt;Main findings with magnitudes: (1) Establishments underutilize their most sophisticated adopted technology. In 63% of BFs where multiple technologies are used, MOST is not the most sophisticated available; the MAX-MOST gap appears in 62% of multi-technology BFs. (2) MAX and MOST are distinct upgrading processes: a one-unit rise in the number of technologies (NUM) raises MAX by 0.84 but MOST by only 0.25; MAX explains just 34% of within-establishment MOST variance. (3) Gaps are persistent, not transitory: only weakly related to age (cross-decile correlation -0.29; individual -0.01) and unrelated to time since adoption. (4) Gap frequency falls with income (country-level 51% in Korea to 83% in Burkina Faso; correlation -0.55 with per-capita income) and rises with input scarcity (low human capital, loan denial) and managerial mistakes (perception bias, family ownership, non-exporting). (5) Within-country dispersion in gaps (0.28) is about three times the between-country dispersion (0.09). (6) Establishment-level MAX and MOST average 2.6 and 2.0; both correlate with income (0.78 for MAX, 0.94 for MOST) and with size, human capital, management, exporter and multinational status. (7) Both productivity and profitability rise with sophistication, more strongly for MOST and for agriculture; the association is not smaller in low-income countries, contradicting the &amp;ldquo;appropriate technology&amp;rdquo; hypothesis.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-max-and-most-and-why-are-they-conceptually-distinct"&gt;Q1. What are MAX and MOST, and why are they conceptually distinct?&lt;/h3&gt;
&lt;p&gt;MAX_{f,j} is the sophistication of the most advanced grid technology establishment j uses in business function f; MOST_{f,j} is the sophistication of the most widely used technology in that function. Both lie in [1,5] with MAX &amp;gt;= MOST by construction, and both measure closeness to the world frontier. They are conceptually different: increases in MAX reflect adoption of a new (to the function) more sophisticated technology, whereas increases in MOST can reflect adoption OR the extension/intensification of an already-adopted technology — closer to Mansfield&amp;rsquo;s (1963) concept of intra-firm technology diffusion. The paper&amp;rsquo;s central empirical claim is that these are driven by distinct upgrading processes.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identification-strategy-and-what-does-the-paper-not-claim"&gt;Q2. What is the identification strategy, and what does the paper NOT claim?&lt;/h3&gt;
&lt;p&gt;This is a descriptive/correlational paper, not a causal one. The authors explicitly state their data do not permit causal inference; the productivity, profitability, and characteristic associations are partial correlations from cross-sectional regressions with country and 2-digit sector fixed effects. The BF-level analyses (MAX-NUM, MOST-NUM, MAX-MOST) use establishment and function fixed effects to absorb establishment- and function-specific levels. The main &amp;lsquo;identification&amp;rsquo; work is measurement validity, not causal identification.&lt;/p&gt;
&lt;h3 id="q3-how-are-max-and-most-shown-to-be-distinct-upgrading-processes-empirically"&gt;Q3. How are MAX and MOST shown to be distinct upgrading processes empirically?&lt;/h3&gt;
&lt;p&gt;Three pieces of evidence. First, regressing MAX on NUM (number of technologies) with establishment and function FE yields a coefficient of 0.84 (s.e. 0.01) — near one-to-one — while regressing MOST on NUM yields only 0.25 (s.e. 0.01). Second, regressing MOST on MAX (with FE) shows MAX explains only 34% of within-establishment MOST variance, so MAX is not a sufficient statistic for MOST. Third, MAX and MOST have different distributions (MOST more skewed), different lifecycle profiles, different correlates, and different associations with productivity.&lt;/p&gt;
&lt;h3 id="q4-is-the-max-most-gap-transitory-or-persistent-and-how-is-this-tested"&gt;Q4. Is the MAX-MOST gap transitory or persistent, and how is this tested?&lt;/h3&gt;
&lt;p&gt;Persistent. Three exercises: (i) across age deciles the gap correlates only -0.29 with age (-0.01 at the individual level), with no clear lifecycle pattern by income or size except a decline only among large establishments aged 16+; (ii) the distribution of years since adopting a top-tier technology is similar for BFs with and without a gap, so time does not close it; (iii) splitting top-tier adopters into early vs. recent adopters yields similar MOST distributions. Together these confirm gaps persist long after adoption.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-two-hypothesized-drivers-of-max-most-gaps-and-what-evidence-supports-each"&gt;Q5. What are the two hypothesized drivers of MAX-MOST gaps, and what evidence supports each?&lt;/h3&gt;
&lt;p&gt;(1) Input constraints — scarcity of skilled labor or finance pushes firms to rely on simpler technologies operable by less-educated workers or needing less capital. Supported by the negative coefficient on human capital (college share) and the positive coefficient on the loan-denied dummy. (2) Managerial mistakes — poor management or biased self-perception of one&amp;rsquo;s own sophistication causes suboptimal underuse. Supported by positive correlations with perception bias and family ownership, and a negative correlation with exporter status (competitive pressure narrows the gap); the management z-score association is weak. Across subsamples, input scarcity is more prominent in low-income countries while managerial-mistake proxies are more salient among large establishments (likely from the complexity of managing scale).&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-in-technology-sophistication-is-documented"&gt;Q6. What heterogeneity in technology sophistication is documented?&lt;/h3&gt;
&lt;p&gt;By income: country averages span 1.53 (MAX) and 1.01 (MOST); within-country dispersion (p80-p20) rises with income, more steeply for MOST (0.95 vs 0.33). By sector: agriculture shows greater cross-establishment dispersion in both MAX and MOST than manufacturing or services. Lifecycle: MAX rises gradually with age in all income/size groups, but MOST flattens beyond ~10 years in low-income countries and among small establishments. Size effects on MOST are stronger in high-income countries; on MAX they are similar across income levels. The performance-sophistication link is strongest in agriculture and weakest in services, and is not weaker in low- than high-income countries.&lt;/p&gt;
&lt;h3 id="q7-how-much-of-the-variation-is-across-vs-within-sectors-and-why-does-that-matter"&gt;Q7. How much of the variation is across vs. within sectors, and why does that matter?&lt;/h3&gt;
&lt;p&gt;Following Syverson (2011), sector dummies explain only 14% (2-digit), 20% (3-digit), and 23% (4-digit ISIC) of cross-establishment variance in sophistication — comparable to their explanatory power for productivity (sales per worker). This implies sophistication variation reflects differences in the technologies used to perform similar tasks, not differences in what tasks/goods establishments produce.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-and-validation-checks-are-run"&gt;Q8. What robustness and validation checks are run?&lt;/h3&gt;
&lt;p&gt;Cardinalization: linear approximation of the sophistication-productivity relation; correlation &amp;gt;0.98 with z-score cardinalization; median 0.95 (p25-p75: 0.90-0.98) with a productivity-based Q-cardinalization across 18 BFs; establishment-level baseline-vs-Q correlations of 0.90 (MAX) and 0.91 (MOST). Ranking validity: three-stage expert validation (functionality/integration/automation; novelty and cost; ChatGPT replication) on 14 BFs plus an independent relative-productivity exercise on 18 BFs. Data quality: response rates 15-86% (high for establishment surveys); no significant non-response differences in employment, sophistication, wages, or skill; a Kenya back-check pilot showing 80.6% consistency for technology-use reports; external validation against Korea (KED) and Brazil (RAIS) with cross-establishment correlations above 0.93 for sales/employment and 0.73 for labor productivity; ERP adoption in Korean manufacturing of 32% vs. 40% in Chung-Kim (2021). Establishment-level results are robust to controlling for the in-house fraction of functions.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q9. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;It generalizes the intra-firm diffusion literature (Mansfield 1963; Battisti-Stoneman 2003), which studied a handful of technologies in a few countries, by showing MAX-MOST gaps are widespread and persistent across 63 functions and 15 countries. It parallels Bloom-Van Reenen (2007) on management practices in method (expert rankings, survey scoring, z-scores) and finds supporting evidence for the Bloom-Sadun-Van Reenen (2012) technology-management complementarity. It differs from the US Advanced Business Survey / Acemoglu et al. (2022), which covered five frontier technologies, by being comprehensive and frontier-relative. It contributes new evidence to the agricultural productivity gap (Caselli 2005; Gollin-Lagakos-Waugh 2014) and to the appropriate-technology debate (Basu-Weil 1998; Acemoglu-Zilibotti 2001).&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Because the sophistication-performance association is not smaller in low-income than high-income countries, advanced technologies appear &amp;lsquo;appropriate&amp;rsquo; across income levels — challenging the appropriate-technology hypothesis that poor countries gain little from sophisticated technology. Policy should target not only adoption (MAX) but also the extension of use/intensity (MOST), since MOST is more strongly tied to productivity and profitability. Scope conditions: associations are correlational, not causal; samples are representative only for establishments with 5+ workers; coverage is the 12 surveyed sectors; and the cross-section cannot trace dynamics (the authors plan a longitudinal extension).&lt;/p&gt;
&lt;h3 id="q11-what-do-the-descriptive-technology-use-patterns-show-about-adoption-behavior"&gt;Q11. What do the descriptive technology-use patterns show about adoption behavior?&lt;/h3&gt;
&lt;p&gt;Establishments use about two technologies per function on average; 62.6% of functions use more than one and 28.3% use at least three. Leapfrogging/skipping is rare: among single-technology functions (37.4% of cases), 52.8% use the least sophisticated grid technology, so only about 18% of functions have fully skipped or abandoned simpler technologies. In 70.4% of multi-technology functions one technology used is the least sophisticated available, and sophistication gaps (non-contiguous use) occur in only 25% of functions (27% GBF, 17% SSBF; most common in payments 48%, business administration 34%, sales 28%). Firms thus typically retain dominated technologies rather than abandon them, which is why MAX proxies the full adoption history well. Only 16% of establishments use an ERP system (the most sophisticated business-administration technology).&lt;/p&gt;
&lt;h3 id="q12-any-notable-caveats-about-the-measures-themselves"&gt;Q12. Any notable caveats about the measures themselves?&lt;/h3&gt;
&lt;p&gt;MAX-MOST gaps are ordinal (cardinalization-free), but establishment-level MAX and MOST are cardinal and could be sensitive to the chosen cardinalization — addressed by the validation exercises. Establishment-level measures use only in-house functions (87% of relevant SSBFs and an overwhelming majority of GBFs are in-house; only 3.9% of GBFs not in-house), and results are robust to controlling for the in-house share. The survey deliberately avoided the words &amp;rsquo;technology&amp;rsquo; and &amp;lsquo;sophistication&amp;rsquo; (using &amp;lsquo;methods&amp;rsquo;/&amp;lsquo;processes&amp;rsquo;) to limit social-desirability bias.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The grid&lt;/strong&gt;: A two-dimensional structure mapping each key business function (horizontal axis, task-based) to the range of technologies that can perform it (vertical axis, ranked by sophistication from simplest to the world frontier). Spans 63 business functions (7 general + 56 sector-specific across 12 sectors) and 305 technologies, built and ranked by 50+ industry experts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MAX&lt;/strong&gt;: The sophistication (on a [1,5] affine scale) of the most advanced technology an establishment uses in a given business function. Increases in MAX reflect adoption of a technology new to that function; near one-to-one with the number of technologies used (coefficient 0.84).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MOST&lt;/strong&gt;: The sophistication (on a [1,5] scale) of the most widely used technology in a business function. Changes in MOST reflect both adoption and the intensification/extension of already-adopted technologies — closer to Mansfield&amp;rsquo;s (1963) intra-firm diffusion than to adoption per se; only weakly tied to the number of technologies (coefficient 0.25).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MAX-MOST gap&lt;/strong&gt;: A binary indicator equal to 1 when MAX &amp;gt; MOST in a function with multiple technologies in use — i.e., the most widely used technology is not the most sophisticated one adopted. Present in 62-63% of multi-technology functions, persistent over time, and associated with input scarcity, managerial mistakes, and lower productivity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;FAT survey&lt;/strong&gt;: The Firm Adoption of Technology survey: a cross-section of 21,055 establishments forming nationally representative samples (5+ workers) in 15 countries (2019-2023), implementing the grid plus modules on financials, employment, management practices, and adoption barriers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Appropriate technology hypothesis&lt;/strong&gt;: In this paper&amp;rsquo;s usage, the claim (Basu-Weil 1998; Acemoglu-Zilibotti 2001) that establishments in poor countries underutilize sophisticated technologies because scarce human and physical capital limits the productivity gains those technologies embody. The paper&amp;rsquo;s finding that the sophistication-performance association is not smaller in low-income countries runs counter to this hypothesis.&lt;/p&gt;</description></item><item><title>TFPR: Dispersion and Cyclicality</title><link>https://macropaperwarehouse.com/papers/tfpr-dispersion-and-cyclicality/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/tfpr-dispersion-and-cyclicality/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates what drives the countercyclical dispersion of TFPR — total factor productivity measured in revenue terms — a pattern that is well documented empirically but poorly understood theoretically. The central motivation is a gap between data measurement and model theory: empirical studies (Kehrig 2011; Bloom, Floetotto, Jaimovich, Eksten, and Terry 2018) document countercyclical dispersion of TFPR, yet the models that seek to explain it routinely conflate TFPR with TFPQ (quantity-based TFP) and treat the two as interchangeable. Cooper and Ozturk argue this conflation is misleading because the distribution of TFPR is endogenous — it depends both on the exogenous distribution of TFPQ and on the endogenous price-setting decisions of firms.&lt;/p&gt;
&lt;p&gt;The paper builds an overlapping generations (OG) model with monopolistic competition and state-dependent pricing (menu costs). Young agents set prices ex ante, observe idiosyncratic productivity shocks, menu cost draws, and aggregate shocks, then decide whether to adjust prices ex post at a fixed cost. Old agents consume a CES bundle of goods produced by the young. The aggregate state includes shocks to the money supply, to the mean (µQ) and dispersion (dispQ) of TFPQ, and to the dispersion of idiosyncratic demand (dispD). The model is solved as a stationary rational expectations equilibrium (SREE) without linear approximations, allowing the nonlinear hazard of price adjustment to propagate to the aggregate.&lt;/p&gt;
&lt;p&gt;The calibration matches three moments: the standard deviation of TFPR (dispR = 0.102 in data, 0.103 in model), the ratio of dispersion in TFPQ to TFPR (1.181 in both), and the monthly frequency of price adjustment (0.110 in data, 0.127 in model), using parameters from Vavra (2014) and Foster, Haltiwanger, and Syverson (2008). The model period is one month. A key structural feature is a U-shaped hazard of price adjustment: firms with very large or very small gaps between actual and desired prices are most and least likely to adjust, respectively.&lt;/p&gt;
&lt;p&gt;The central empirical target is three jointly countercyclical moments: (i) dispersion of TFPR, (ii) dispersion of price changes, (iii) frequency of price adjustment. The paper&amp;rsquo;s first set of findings is negative. Taken individually, no single shock source reproduces all three patterns. Specifically, shocks to dispQ alone produce procyclical TFPR dispersion — output expands when dispersion rises because high-productivity firms can produce more, but TFPR dispersion rises with dispQ (and hence with output), contradicting the data. Money shocks produce procyclical TFPR dispersion and an inverse U-shaped relationship between dispR and the money shock: at extreme shock values, more firms adjust to the common nominal shock, compressing TFPR dispersion; at moderate values, idiosyncratic heterogeneity dominates and dispR is higher. Shocks to µQ alone leave TFPR dispersion nearly flat. Shocks to dispD produce slight countercyclical TFPR dispersion but counterfactually procyclical price adjustment moments.&lt;/p&gt;
&lt;p&gt;Two combinations succeed. First, a joint shock to dispQ and µQ with perfect negative correlation (corr = -1, as in Vavra 2014) generates all three countercyclical moments: as dispQ rises, µQ falls, and output contracts while TFPR dispersion increases; from Table 5, dispR is 0.126 in contraction versus 0.020 in expansion, disp∆p is 0.208 in contraction versus 0.082 in expansion, and freq∆p is 0.328 in contraction versus 0.164 in expansion. Second, a monetary feedback rule where the central bank leans against the wind (ζ = -0.05) — tightening money when dispQ is above average — also replicates all three countercyclical moments (Table 5, leaning-against-the-wind rows).&lt;/p&gt;
&lt;p&gt;Two additional findings emerge. The model generates state-dependent monetary policy effectiveness: the response of output to a monetary shock is larger in expansions (coefficient 0.644) than in contractions (0.578) when business cycle state is measured by output growth, consistent with Tenreyro and Thwaites (2016) only for the growth-based measure. The paper also finds no role for uncertainty distinct from realized dispersion: when Markov-switching uncertainty over TFPQ dispersion is introduced, the ex ante price is essentially unchanged, consistent with Berger, Dew-Becker, and Giglio (2020).&lt;/p&gt;
&lt;p&gt;The theoretical contribution is a TFPR decomposition: Var(tfpr) = Var(tfpq) + Var(ln p) + 2·Cov(ln p, tfpq). In the FHS data, Var(tfpr) = 0.0484, Var(tfpq) = 0.0676, Var(ln p) = 0.0324, Cov(ln p, tfpq) = -0.0258. In recessions, Var(tfpr) rises to 0.0618, driven by an increase in Var(ln p) to 0.0506 while Var(tfpq) stays at 0.0676. This means countercyclical TFPR dispersion can be generated through endogenous price adjustment even holding TFPQ dispersion fixed — a mechanism entirely absent from models that equate TFPR with TFPQ.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-fundamental-measurement-theory-gap-the-paper-identifies"&gt;Q1. What is the fundamental measurement-theory gap the paper identifies?&lt;/h3&gt;
&lt;p&gt;Existing business cycle models (Bloom et al. 2018, Vavra 2014) are calibrated to observed countercyclical dispersion of TFPR but then build theoretical mechanisms around countercyclical dispersion of TFPQ, treating the two as equivalent. Cooper and Ozturk show this is incorrect: TFPR = TFPQ × (p/P), so the TFPR distribution is endogenous, shaped by both the exogenous TFPQ distribution and the endogenous price-setting decisions of firms. Changes in the distribution of prices — through extensive and intensive margins of price adjustment — can move TFPR dispersion independently of TFPQ dispersion.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-og-framework-give-the-model-tractability-advantages"&gt;Q2. Why does the OG framework give the model tractability advantages?&lt;/h3&gt;
&lt;p&gt;In the OG model, young sellers make price decisions within a single period, so the ex post price is independent of the ex ante price. This means the state space is simplified (no lagged own-price), individual choice problems are tractable, the ex post pricing problem is static, and the full SREE can be characterized without log-linear approximations. Crucially, this allows the nonlinear U-shaped price adjustment hazard to propagate to aggregate outcomes exactly, without the approximation errors that would arise in linearized dynamic models.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-main-shocks-in-the-model-and-how-are-they-parameterized"&gt;Q3. What are the main shocks in the model and how are they parameterized?&lt;/h3&gt;
&lt;p&gt;There are four aggregate shocks: (i) money supply shocks x, (ii) shocks to the mean of TFPQ (µQ), (iii) shocks to the dispersion of TFPQ (dispQ, implemented as a mean-preserving spread in z), and (iv) shocks to the dispersion of idiosyncratic demand (dispD). At the individual level, sellers face idiosyncratic productivity shocks z with standard deviation σz = 0.0378 and idiosyncratic demand shocks with σd = 0.0069. Menu costs follow the Dotsey and Wolman (2019) distribution with a fraction ψ = 0.053 of firms having zero adjustment costs.&lt;/p&gt;
&lt;h3 id="q4-why-does-a-dispq-shock-alone-produce-procyclical-not-countercyclical-tfpr-dispersion"&gt;Q4. Why does a dispQ shock alone produce procyclical, not countercyclical, TFPR dispersion?&lt;/h3&gt;
&lt;p&gt;An increase in dispQ expands the tails of the productivity distribution. High-productivity firms can produce more and expand output (reallocating labor to them raises aggregate output), so output rises with dispQ. Simultaneously, higher dispQ directly raises TFPR dispersion because TFPR = (p/P)×TFPQ and the increased heterogeneity in z carries through to TFPR. Since dispR rises when output rises, the cyclicality is procyclical — directly contradicting the empirical pattern. The pricing response (more adjustment for extreme z draws) magnifies rather than offsets this pattern.&lt;/p&gt;
&lt;h3 id="q5-how-do-monetary-shocks-affect-tfpr-dispersion-and-why-is-the-relationship-non-monotone"&gt;Q5. How do monetary shocks affect TFPR dispersion, and why is the relationship non-monotone?&lt;/h3&gt;
&lt;p&gt;Money shocks cause a rightward shift in the price gap distribution rather than a spread. For moderate money shocks (near average), few firms adjust, so non-adjusters retain their ex ante prices and face heterogeneous gaps — TFPR dispersion is high. For extreme money shocks (very high or very low), many firms adjust to align their prices with the common nominal shock, compressing idiosyncratic price dispersion. Combined with U-shaped adjustment frequency, this creates an inverse U-shaped relationship between dispR and the money shock: TFPR dispersion is highest at moderate shocks and lowest at extreme shocks. Consequently, money shocks alone produce procyclical TFPR dispersion on average, but the model can produce countercyclical dispersion for extreme realizations.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-joint-dispq-µq-shock-with-perfect-negative-correlation-work-to-match-the-data"&gt;Q6. How does the joint (dispQ, µQ) shock with perfect negative correlation work to match the data?&lt;/h3&gt;
&lt;p&gt;Following Vavra (2014), the paper assumes corr(dispQ, µQ) = -1: the highest dispQ state is paired with the lowest µQ state and so on. When dispQ rises, µQ falls. The mean productivity drop dominates in determining output (output contracts), while the dispersion increase drives up TFPR dispersion. This creates countercyclical dispR. From Table 5, in contractions: dispR = 0.126, disp∆p = 0.208, freq∆p = 0.328; in expansions: dispR = 0.020, disp∆p = 0.082, freq∆p = 0.164. All three moments are countercyclical, matching the data. The key mechanism is that the two shocks drive a wedge between the movements in mean output (dominated by µQ) and the movements in dispersion (dominated by dispQ).&lt;/p&gt;
&lt;h3 id="q7-how-does-the-monetary-leaning-against-the-wind-feedback-rule-generate-countercyclical-tfpr-dispersion"&gt;Q7. How does the monetary &amp;rsquo;leaning against the wind&amp;rsquo; feedback rule generate countercyclical TFPR dispersion?&lt;/h3&gt;
&lt;p&gt;The central bank sets money growth as Mt+1 = Mt[Φ(st+1) + x̃t+1] where Φ(dispQ) = ζ × (dispQ − µdispQ) with ζ &amp;lt; 0 (specifically ζ = -0.05 in the main experiment). When dispQ is above average, the central bank contracts money supply. Since without this rule increased dispQ raises output (procyclical), the monetary contraction more than offsets this, turning the dispQ shock into a net recessionary force. Meanwhile TFPR dispersion still tracks dispQ and rises. Result: both dispR and recession coincide. Table 5 shows that with leaning against the wind on dispQ shocks, dispR = 0.093 in contraction versus 0.082 in expansion, and all three moments remain countercyclical. A second case (feedback to µQ shocks) also produces countercyclical dispR but fails to match the pricing-frequency moment (which becomes procyclical due to asymmetry in the U-shaped hazard).&lt;/p&gt;
&lt;h3 id="q8-what-are-the-nonlinearities-in-the-model-and-why-does-the-paper-avoid-using-correlations-as-summary-statistics"&gt;Q8. What are the nonlinearities in the model and why does the paper avoid using correlations as summary statistics?&lt;/h3&gt;
&lt;p&gt;The U-shaped price adjustment hazard creates nonlinear aggregate responses: variables can be positively correlated with output in expansions and negatively correlated in contractions, or vice versa. For example, under money shocks the correlation of frequency of price adjustment with output is -0.648 in contractions and +0.977 in expansions (Table 7). The dispersion of TFPR under money shocks also switches sign across states. Standard unconditional correlations average over these sign switches and can give misleading or zero correlations, masking the underlying structure. The SREE is solved exactly without linearization so these nonlinearities are not averaged away in the solution.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-finding-on-the-state-dependence-of-monetary-policy-effectiveness"&gt;Q9. What is the finding on the state-dependence of monetary policy effectiveness?&lt;/h3&gt;
&lt;p&gt;Table 8 reports regressions of log output on the log money shock separately in contractions and expansions. When recessions are defined by output below trend, the coefficient is 0.578 in contractions and 0.644 in expansions — monetary policy is less effective in recessions. When recessions are defined by three consecutive periods of negative output growth (as in Tenreyro and Thwaites 2016), coefficients are 0.589 in contractions and 0.611 in expansions — the same qualitative finding. However, this contrasts with Tenreyro and Thwaites (2016) in that the paper finds the asymmetry holds regardless of whether the cycle state is measured in levels or growth rates, whereas Tenreyro and Thwaites find the effect only for growth-based definitions. The mechanism is that recessions (high dispQ, low µQ) are associated with more frequent price adjustment, which attenuates the real effect of money shocks.&lt;/p&gt;
&lt;h3 id="q10-what-is-found-regarding-the-effects-of-uncertainty-versus-realized-dispersion"&gt;Q10. What is found regarding the effects of uncertainty versus realized dispersion?&lt;/h3&gt;
&lt;p&gt;The paper introduces Markov-switching uncertainty where firms do not know in advance which dispersion regime they are in (high or low dispQ). For the ex ante price setting problem, this amounts to taking an expectation over the future dispersion distribution. The quantitative finding is that the ex ante price is essentially unchanged when uncertainty over the dispersion regime is added versus the baseline without such uncertainty. This confirms that the effects on price adjustment and TFPR dispersion in the model come from the realized dispersion, not from ex ante uncertainty about which regime will prevail — consistent with Berger, Dew-Becker, and Giglio (2020) who find that uncertainty shocks have negligible real effects.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-variance-decomposition-of-tfpr-characterize-the-empirical-patterns"&gt;Q11. How does the variance decomposition of TFPR characterize the empirical patterns?&lt;/h3&gt;
&lt;p&gt;The paper uses the identity Var(tfpr) = Var(tfpq) + Var(ln p) + 2·Cov(ln p, tfpq). In the FHS data: Var(tfpr) = 0.0484, Var(tfpq) = 0.0676, Var(ln p) = 0.0324, Cov(ln p, tfpq) = -0.0258. The covariance is negative (prices are lower for high-productivity firms, consistent with markup compression), which is why Var(tfpr) &amp;lt; Var(tfpq). In recessions: Var(tfpr) rises to 0.0618, Var(tfpq) is held fixed at 0.0676 (by assumption in the thought experiment), Var(ln p) rises to 0.0506 (from Vavra 2014), and Cov(ln p, tfpq) becomes more negative at -0.0282. This decomposition shows that countercyclical TFPR dispersion can be generated by endogenous price changes — through both higher price variance and a larger (absolute) covariance between prices and productivity — even if TFPQ dispersion is fixed.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-role-of-the-u-shaped-adjustment-hazard-in-the-model"&gt;Q12. What is the role of the U-shaped adjustment hazard in the model?&lt;/h3&gt;
&lt;p&gt;The U-shaped hazard (probability of price adjustment as a function of the price gap or idiosyncratic shock z) is a key structural feature inherited from state-dependent pricing. Adjustment probability is near zero for small gaps (moderate z) and rises steeply for large gaps (extreme z). This creates nonlinear responses: a mean-preserving spread in z (dispQ shock) pushes more mass into the tails, sharply increasing adjustment frequency; a mean shift in z (µQ shock) shifts the gap distribution rightward, also raising adjustment but asymmetrically; a money shock shifts all gaps in one direction (rightward for a positive shock). The interaction between the shock type and the hazard shape determines whether the covariance of prices and productivity rises or falls, which in turn determines whether TFPR dispersion moves countercyclically.&lt;/p&gt;
&lt;h3 id="q13-how-does-price-stickiness-create-a-non-degenerate-tfpr-distribution-without-needing-other-frictions"&gt;Q13. How does price stickiness create a non-degenerate TFPR distribution without needing other frictions?&lt;/h3&gt;
&lt;p&gt;In the flexible-price monopolistic competition benchmark (used for comparison), if production is linear in labor (α=1), TFPR = ω/(1-η) and is independent of z — the TFPR distribution is degenerate. In the sticky-price model, non-adjusters set prices ex ante proportional to the money supply, while adjusters set prices that depend on both z and the money shock. The resulting cross-sectional distribution of prices is non-degenerate and generates a non-degenerate TFPR distribution. The coexistence of adjusters and non-adjusters — with prices reflecting both idiosyncratic productivity and aggregate conditions to different degrees — is sufficient to generate TFPR heterogeneity without additional distortions or wedges.&lt;/p&gt;
&lt;h3 id="q14-what-are-the-robustness-checks-and-how-do-they-affect-the-main-findings"&gt;Q14. What are the robustness checks and how do they affect the main findings?&lt;/h3&gt;
&lt;p&gt;Table 6 reports robustness under money shocks alone across three parameter changes: (1) Higher elasticity of substitution ε = 4 (versus baseline 2.37): higher adjustment frequency, lower price change dispersion, but still procyclical TFPR dispersion. (2) Lower labor supply convexity φ = 1.5 (versus baseline 2): moments become nearly acyclical; TFPR dispersion is much higher than baseline. (3) Equal demand and productivity shock dispersion σd = σz: frequency of price adjustment is nearly four times the baseline, but the monetary shock model still fails to generate countercyclical TFPR dispersion. None of these alternatives bring the money-shock-only model into line with the data, confirming that the main positive results (joint dispQ-µQ shock, or monetary feedback) are not artifacts of baseline parameterization. The paper also notes its calibrated ε is lower than Vavra (2014) and Golosov-Lucas (2007), which use higher elasticities and linear labor disutility.&lt;/p&gt;
&lt;h3 id="q15-how-does-this-paper-relate-to-and-differ-from-vavra-2014-and-bloom-et-al-2018"&gt;Q15. How does this paper relate to and differ from Vavra (2014) and Bloom et al. (2018)?&lt;/h3&gt;
&lt;p&gt;Vavra (2014) documents countercyclical dispersion of price changes and frequency, and argues this follows from countercyclical TFPQ dispersion driving volatility of firm-level productivity shocks. He calibrates to TFPR moments but treats TFPQ and TFPR as equivalent. Bloom et al. (2018) combine uncertainty and dispersion shocks to TFPQ to generate aggregate fluctuations, requiring both a rise in dispQ and a fall in mean TFPQ to avoid counterfactual negative correlation between consumption and investment. Cooper and Ozturk differ in three respects: (i) they explicitly model the TFPQ-to-TFPR mapping through state-dependent pricing; (ii) they show that dispQ shocks alone produce procyclical (not countercyclical) TFPR dispersion in their model; (iii) while they confirm that the joint (dispQ, µQ) combination matches data, they attribute the mechanism to the pricing wedge rather than uncertainty — uncertainty per se has no effect in their framework.&lt;/p&gt;
&lt;h3 id="q16-what-are-the-limitations-and-directions-for-future-work-noted-by-the-authors"&gt;Q16. What are the limitations and directions for future work noted by the authors?&lt;/h3&gt;
&lt;p&gt;The OG model&amp;rsquo;s one-period price-setting horizon misses forward-looking dynamics in price adjustment — specifically, the distinction between permanent and temporary adjustment opportunities that matters in infinite-horizon models. However, the authors show the OG model&amp;rsquo;s policy functions and hazard shape closely replicate those from infinite-horizon state-dependent pricing models, so this limitation is argued to be minor. On the data side, the authors note the ideal structural estimation would use high-frequency joint data on prices and quantities at the firm level, which is not yet available. They suggest future work extending the model to incorporate real-options-style wait-and-see behavior (as in Bloom 2009) combined with state-dependent pricing, and point to the value of non-linear empirical methods (analogous to Tenreyro and Thwaites 2016) for studying price adjustment dynamics.&lt;/p&gt;
&lt;h3 id="q17-what-is-the-relationship-between-idiosyncratic-demand-shocks-and-tfpr-dispersion"&gt;Q17. What is the relationship between idiosyncratic demand shocks and TFPR dispersion?&lt;/h3&gt;
&lt;p&gt;Idiosyncratic demand shocks (αi) directly affect a seller&amp;rsquo;s revenue without changing physical productivity z. Under flexible prices they would affect TFPR directly; under sticky prices the adjustment decision interacts with both the demand and productivity shocks. From Table 5, dispD shocks generate slightly countercyclical TFPR dispersion, but the pricing moments (dispersion of price changes and adjustment frequency) are procyclical — inconsistent with the data. Additionally, the dispersion of demand shocks (σd = 0.0069) is calibrated to be about 18% of productivity shock dispersion (σz = 0.0378), so demand shocks play a smaller quantitative role in the baseline. When σd = σz (equal dispersions), adjustment frequency is nearly four times the baseline but the model still fails to match all three target moments.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;TFPR (Revenue Total Factor Productivity)&lt;/strong&gt;: In this paper, TFPR = (p/P) × TFPQ, where p is a firm&amp;rsquo;s price and P is the aggregate price index. It is the revenue-based measure of productivity that is directly observed in plant-level data. Its distribution is endogenous because prices are set by sellers; unlike TFPQ, it is not a primitive of the model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;TFPQ (Quantity Total Factor Productivity)&lt;/strong&gt;: The physical or quantity-based measure of productivity, denoted z in the model. It is exogenous to the individual seller and drawn from a distribution that can shift in mean (µQ) or dispersion (dispQ). TFPQ is the primitive shock; TFPR is derived from TFPQ through the pricing decisions of sellers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;State-Dependent Pricing (SDP)&lt;/strong&gt;: A pricing framework in which firms adjust prices only when the gain from adjustment exceeds a menu cost. In this paper, sellers set prices ex ante and then decide ex post whether to pay a stochastic cost to reset. Price adjustment depends on the realized state (idiosyncratic z, money shock x), creating both extensive margin (who adjusts) and intensive margin (what price to set) decisions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stationary Rational Expectations Equilibrium (SREE)&lt;/strong&gt;: The equilibrium concept used in the paper. It is a set of ex ante prices, ex post prices, critical adjustment costs, and aggregate price levels that are mutually consistent across all aggregate and idiosyncratic states. The SREE is solved exactly without log-linear approximations, allowing the model&amp;rsquo;s nonlinearities to be preserved.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;U-Shaped Adjustment Hazard&lt;/strong&gt;: The probability of price adjustment as a function of the gap (difference between desired and actual log price) is U-shaped: near-zero for small gaps and sharply increasing for large gaps in either direction. This creates nonlinear aggregate responses to shocks — aggregate variables can comove differently in expansions versus contractions — and is a central driver of the model&amp;rsquo;s results on TFPR cyclicality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Leaning Against the Wind (Monetary Feedback Rule)&lt;/strong&gt;: A monetary policy rule in the paper where the central bank contracts the money supply when the dispersion of TFPQ (dispQ) rises above its average (ζ &amp;lt; 0 in the feedback rule). By doing so, the authority converts what would otherwise be a procyclical dispQ shock into a recessionary one, generating countercyclical TFPR dispersion as a byproduct.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;dispQ Shock&lt;/strong&gt;: An aggregate mean-preserving spread in the distribution of idiosyncratic productivity z. It widens the cross-sectional distribution of TFPQ without changing its mean. Taken alone, it produces procyclical TFPR dispersion; combined with a negative shock to µQ (or with monetary tightening), it can produce countercyclical TFPR dispersion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Price Gap&lt;/strong&gt;: The difference between the log of the price a seller would optimally set if adjustment were free and the log of the seller&amp;rsquo;s current ex ante price. The gap is the sufficient statistic for the price adjustment decision: sellers with larger gaps (in absolute value) have larger gains to adjustment and hence higher adjustment probability. The distribution of gaps across sellers responds to aggregate shocks and shapes aggregate price dynamics.&lt;/p&gt;</description></item><item><title>The (In)effectiveness of Targeted Payroll Tax Reductions</title><link>https://macropaperwarehouse.com/papers/the-ineffectiveness-of-targeted-payroll-tax-reductions/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-ineffectiveness-of-targeted-payroll-tax-reductions/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper studies the cost-effectiveness of targeted payroll tax reductions as a tool for stimulating labor demand among marginalized workers, using a natural experiment from Italy. The motivation is policy-relevant: governments routinely deploy targeted payroll tax cuts to combat youth and low-skill unemployment, but such subsidies risk subsidizing inframarginal hiring — employment that would have occurred without the incentive — rather than creating net new jobs. Rigorous evaluation requires two features that are rarely satisfied simultaneously: (1) the subsidy must target genuinely marginalized workers so estimates pertain to the population of interest, and (2) variation in incentives across firms must be quasi-random so firm responses are causally identified. This paper exploits a policy that satisfies both.&lt;/p&gt;
&lt;p&gt;The data are confidential matched employer-employee records from the Italian Social Security Institute (INPS), covering the universe of private non-agricultural firms with at least one employee from January 2003 to December 2009. The main analysis sample comprises 1,015,619 firms with policy-relevant firm size between 3 and 15 employees — the stratum containing the policy threshold. The study period spans 84 months.&lt;/p&gt;
&lt;p&gt;The policy variation is the Italian 2007 Budget Bill (Law 296/2006), which raised employer social security contributions (SSCs) on apprenticeship contracts from a flat rate of 148 euros per year to 10 percent of annual earnings (approximately 1,200 euros per year for an average apprentice earning 12,000 euros). However, firms with at most 9 full-time-equivalent employees (excluding apprentices) received a graduated discount: 1.5 percent of earnings in the first year (180 euros) and 3 percent in the second year (360 euros). This generated a clean discontinuity in incentives at the 9-employee threshold. The discount is equivalent to roughly two months of earnings per apprentice, or about 8 percent of the cost of a typical 19-month apprenticeship.&lt;/p&gt;
&lt;p&gt;The empirical strategy is a difference-in-discontinuities design. For each calendar month, the authors estimate a regression discontinuity specification comparing firms just above and just below the 9-employee threshold, then subtract the estimated baseline discontinuity from January 2006 (before the policy existed). This normalizes away pre-existing size-related differences in outcomes, yielding reduced-form estimates of how the policy-induced difference in SSC costs between small and large firms changed over time. The policy variation is used as an instrument for actual SSC payments to compute IV estimates of jobs supported per euro of foregone revenue.&lt;/p&gt;
&lt;p&gt;The main finding is a precise zero: the SSC discount does not increase the number of apprenticeship contracts. The reduced-form estimates of the policy&amp;rsquo;s effect on apprentice hiring are not statistically different from zero and are tightly estimated. Firms below the threshold pay approximately 25 euros less per month in SSCs than firms above, confirming the policy has fiscal bite (first-stage F-statistic = 230), but this differential generates no detectable behavioral response in employment.&lt;/p&gt;
&lt;p&gt;The policy also does not increase the rate at which apprentices are converted to permanent contracts (&amp;ldquo;transformations&amp;rdquo;). Firms do not adjust apprentice wages, do not substitute toward other contract types, do not churn through more apprentices, do not re-label existing contracts, and do not lower hiring standards for apprentices.&lt;/p&gt;
&lt;p&gt;For cost-effectiveness, the IV estimates imply that each 1 million euros of foregone SSC revenue supports the employment of 29 apprentices for one year — a point estimate not statistically different from zero. The point estimate for supported permanent-contract transformations is negative (point estimate: -2), also indistinguishable from zero. By comparison, directly hiring apprentices at their prevailing wage of 1,050 euros per month would employ 79 apprentices per million euros, making direct hiring 2.7 times more cost-effective than the subsidy. The paper surveys the broader literature and finds that once existing studies&amp;rsquo; employment effects are normalized against fiscal costs, targeted subsidies rarely appear cost-effective; hiring credits that require a new hire may outperform payroll tax cuts because they are harder to claim for inframarginal employment.&lt;/p&gt;
&lt;p&gt;The underlying mechanism is inelastic labor demand for apprentices. Survey evidence from the RIL firm survey confirms that when firms do not hire apprentices, cost is rarely the stated reason — the most common answer is that they do not need more people. When firms do hire apprentices, the most common reason is to provide training before converting them to permanent employees, not to economize on labor costs.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The identification strategy is a difference-in-discontinuities design. In each month, a regression discontinuity (RD) specification compares firms just above and just below the 9-employee SSC eligibility threshold; the authors then subtract the baseline (January 2006, pre-policy) discontinuity estimate to remove pre-existing size-related level differences. The key identifying assumption is a &amp;lsquo;weak parallel trends&amp;rsquo; assumption: the curvature of the conditional expectation function of untreated potential outcomes at the threshold is time-invariant. Threats and the evidence against them: (1) Manipulation of firm size at the threshold — addressed by showing that the CDF of policy-relevant firm size is virtually identical across all 84 months with no bunching at 9 employees before or after the reform; (2) Pre-existing trends — no pre-trends are found in the estimated discontinuity in outcomes for the four years before January 2007; (3) Compositional shifts — covariate balance tests show that firm characteristics (age, type, industry, region) at the threshold do not change over time relative to baseline; the covariate index (predicted apprentice hiring based on time-invariant firm characteristics) fluctuates between -0.0005 and +0.0005 — nearly two orders of magnitude smaller than the employment estimates; (4) Imperfect compliance — handled explicitly: the design estimates an intention-to-treat effect, which is attenuated relative to the treatment on the treated; (5) Measurement error in running variable — addressed by excluding firms within one unit of the threshold in the preferred specification; null results are robust to varying the exclusion window.&lt;/p&gt;
&lt;h3 id="q2-why-is-the-difference-in-discontinuities-design-superior-to-a-standard-difference-in-differences-design-in-this-context"&gt;Q2. Why is the difference-in-discontinuities design superior to a standard difference-in-differences design in this context?&lt;/h3&gt;
&lt;p&gt;The paper provides a formal and empirical case that standard difference-in-differences applied to a continuous firm-size running variable produces spurious results. When the conditional expectation function of outcomes with respect to firm size rotates over time (i.e., the slope changes), a DiD estimator that discretizes firms into treated and control groups will detect this rotation as a treatment effect, even if the true policy effect is zero. This is because the DiD constrains the slopes of the conditional expectation function above and below the threshold to be zero, making them implicit omitted variables. In the Italian data, the conditional expectation function of apprentice hiring with respect to firm size rotates clockwise between 2007 and 2009, coinciding with a general slowdown in hiring during the Great Recession. This rotation would cause a naive DiD analysis to conclude, spuriously, that the subsidy supported hiring. The difference-in-discontinuities design controls flexibly for the running variable in each period and isolates only the variation near the threshold, where firm size cannot proxy for trends unrelated to the policy.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-main-mechanisms-considered-for-why-the-subsidy-has-no-employment-effect-and-how-does-the-paper-distinguish-among-them"&gt;Q3. What are the main mechanisms considered for why the subsidy has no employment effect, and how does the paper distinguish among them?&lt;/h3&gt;
&lt;p&gt;The paper considers and rules out seven alternative explanations before concluding that demand for apprentices is simply inelastic: (1) Measurement error — ruled out because the null holds across specifications with different exclusion windows, and measurement error does not prevent finding significant effects on fiscal outcomes; (2) Subsidy too small — ruled out because the 8% subsidy (960 euros per apprentice per year, up to 1,460 euros at the 95th percentile of earnings) is comparable in magnitude to subsidies that generate large employment effects in Cahuc et al. (2019) and Guo (2024); (3) Low awareness — ruled out because 80% of eligible firms that hire apprentices receive the discount, confirming they must claim it actively; (4) Firms restricting hiring to maintain eligibility — ruled out because apprentices are excluded from policy-relevant firm size, so hiring an apprentice does not risk crossing the threshold; the firm-size distribution also remains stable; (5) Temporary nature of subsidy — ruled out because most apprenticeships last 19 months and the subsidy covers the first two years; moreover, the literature suggests temporary subsidies should be at least as effective as permanent ones; (6) Training requirements — ruled out because training requirements are poorly enforced, and no effects are found even among firms that previously employed apprentices (lower marginal training costs) or firms that rarely cite training costs as a deterrent; (7) Great Recession — ruled out because no effects appear in the year before the recession began, and effects are not larger or smaller for liquidity-constrained firms.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-analyses-are-conducted-and-what-do-they-show"&gt;Q4. What heterogeneity analyses are conducted and what do they show?&lt;/h3&gt;
&lt;p&gt;The authors estimate pooled post-reform difference-in-discontinuities coefficients separately across multiple dimensions and find consistently null effects with no evidence of heterogeneous treatment effects: (1) by industry — estimates across manufacturing, transportation and construction, trading, services, and other sectors are all tightly centered on zero; (2) by region — null across all Italian regions; (3) by baseline apprentice earnings quartile — null across Q1 through Q4 and for firms with no apprentices at baseline; (4) by contemporaneous apprentice earnings quartile — null; (5) by three measures of liquidity constraints (liquid assets to total assets, cash flow to total assets, revenues above/below median) — null in all six groups; and (6) by prior apprenticeship training status — null for both firms that employed at least one apprentice in 2006 and those that did not. The authors note the scope condition: estimates are internally valid for firms in a neighborhood of 9 employees, and effects for substantially larger firms cannot be ruled out to differ.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-conducted-beyond-the-main-heterogeneity-analysis"&gt;Q5. What robustness checks are conducted beyond the main heterogeneity analysis?&lt;/h3&gt;
&lt;p&gt;The main robustness checks are: (1) sensitivity of apprentice hiring effects to the amount of excluded data around the threshold (the &amp;lsquo;donut bandwidth&amp;rsquo;) — the null holds across all exclusion windows (Appendix Figure A.2); (2) placebo tests using the pre-reform periods (January 2003 through December 2006) — no pre-trends in the estimated discontinuity for any outcome; (3) covariate stability tests — the discontinuity in a covariate index predicting apprentice hiring from time-invariant firm characteristics shows no change over time, with point estimates between -0.0005 and +0.0005 versus employment estimates between -0.01 and +0.01; (4) comparison of results to a standard DiD specification — the DiD produces spurious positive effects driven by rotation of the conditional expectation function, while the difference-in-discontinuities estimate remains precisely zero; (5) examination of other outcomes (contract churn, re-labeling, worker quality, contract type substitution, temporary worker stocks) — all null.&lt;/p&gt;
&lt;h3 id="q6-how-is-cost-effectiveness-formally-measured-and-what-does-the-iv-estimate-imply"&gt;Q6. How is cost-effectiveness formally measured and what does the IV estimate imply?&lt;/h3&gt;
&lt;p&gt;Cost-effectiveness is defined as the number of jobs supported per unit of foregone revenue: omega = E[L(1) - L(0)] / E[R(0) - R(1)], where L is employment and R is tax payments. Rather than back-of-the-envelope calculation, the authors estimate this with 2SLS, instrumenting for actual SSC payments with the interaction of being below the eligibility threshold and the post-2007 indicator. This allows them to compute standard errors, which back-of-the-envelope methods do not provide. The first-stage F-statistic is 230, confirming instrument strength. Point estimates from Table 4: 29 apprentice-years supported per 1 million euros of foregone SSC (standard error 58, not significant); 647,237 euros of apprentice compensation supported per 1 million euros (standard error 921,320, not significant); and -2 permanent-contract transformations per 1 million euros (standard error 21, not significant). For context, directly hiring apprentices at 1,050 euros per month would generate 79 apprentice-years per million euros — 2.7 times more than the point estimate from the subsidy.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-paper-benchmark-its-cost-effectiveness-estimates-against-the-broader-literature"&gt;Q7. How does the paper benchmark its cost-effectiveness estimates against the broader literature?&lt;/h3&gt;
&lt;p&gt;The authors normalize employment effects from nine other studies against their fiscal costs to produce a common metric of jobs or job-years per 1 million dollars of foregone revenue. The studies span payroll tax cuts (Egebark and Kaunitz 2013; Saez, Schoefer, and Seim 2021), hiring credits (Cahuc, Carcillo, and Le Barbanchon 2019; Neumark 2013), and fiscal stimulus programs (Bartik 2001; Bartik and Erickcek 2010; Dupor and Mehkari 2016; Dupor and McCrory 2018; Feyrer and Sacerdote 2011; Wilson 2012). The conclusion is that most wage subsidies, including those that generate positive reduced-form employment effects, produce very high costs per job. With two exceptions (Bartik 2001 and Cahuc et al. 2019), cost-effectiveness estimates across the literature are extremely low. The paper argues that hiring credits may be more cost-effective than payroll tax cuts because the requirement to make a new hire makes it harder to subsidize inframarginal employment. Importantly, the Italian study&amp;rsquo;s cost-effectiveness estimates — though imprecisely estimated — are broadly consistent with the cross-study pattern once fiscal costs are accounted for.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-welfare-and-public-finance-implications-of-the-null-employment-effects"&gt;Q8. What are the welfare and public finance implications of the null employment effects?&lt;/h3&gt;
&lt;p&gt;Because the behavioral response is zero and the fiscal cost is non-zero, the policy functions as a pure transfer from the government to firms. The paper invokes the framework of Hendren and Sprung-Keyser (2020) to note that the marginal value of public funds is essentially 1 — there is no distortion introduced but also no welfare gain from resource reallocation. This interpretation cuts in two directions: (1) the pre-reform apprentice SSC subsidies (which were larger than the post-2007 discount) were also essentially transfers with large fiscal costs and no employment-creation value; and (2) the SSC increase imposed on larger firms (those with more than 9 employees) effectively raised revenue without causing meaningful employment losses, since labor demand for apprentices is inelastic. The policy is thus deemed inefficient in the sense that taxpayer revenue is lost without generating the intended social return of increasing employment of marginalized workers.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-scope-conditions-and-limitations-of-the-estimates"&gt;Q9. What are the scope conditions and limitations of the estimates?&lt;/h3&gt;
&lt;p&gt;The difference-in-discontinuities design provides internally valid estimates only for firms in a neighborhood of 9 employees, which in Italy means firms with 3 to 15 employees (90% of Italian firms and 65% of all apprentices). The paper cannot rule out that larger firms respond differently to similar subsidies. The analysis is partial equilibrium: it cannot measure spillovers, general equilibrium effects on wage-setting across the firm-size distribution, or displacement effects between firms. Cost-effectiveness estimates reflect only the direct fiscal cost of foregone SSCs and do not include fiscal externalities (e.g., effects on income tax revenues or social insurance outlays) or administrative and political costs. The exclusion of workers from the public sector means the results pertain solely to private-sector apprenticeships.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-prior-studies-on-payroll-tax-cuts-and-what-distinguishes-it-methodologically"&gt;Q10. How does this paper relate to prior studies on payroll tax cuts, and what distinguishes it methodologically?&lt;/h3&gt;
&lt;p&gt;Prior national studies (e.g., Saez et al. 2019, 2012, 2021; Egebark and Kaunitz 2013; Huttunen et al. 2013; Bozio et al. 2020; Rubolino 2021) estimate labor demand responses by comparing employment of targeted versus untargeted workers, which can overstate policy effectiveness if firms substitute targeted for untargeted workers (a SUTVA violation that would not be detected by parallel pre-trend tests). Cross-regional studies (e.g., Bennmarker et al. 2009; Benzarti and Harju 2021a; Bohm and Lind 1993; Guo 2024) study firms but typically do not target genuinely marginalized workers, so estimates reflect average rather than marginal labor demand. This paper satisfies both requirements simultaneously: the discontinuity in incentives provides quasi-random variation across firms (avoiding SUTVA), and the policy specifically targets apprentices — a non-random, marginalized group — so the estimated elasticities pertain to the actual population of interest. The paper is also the first (to the authors&amp;rsquo; knowledge) to use a formal IV strategy to estimate cost-effectiveness with standard errors, enabling statistical precision comparisons across the distribution of estimates.&lt;/p&gt;
&lt;h3 id="q11-what-does-survey-evidence-from-the-ril-data-contribute-to-the-interpretation"&gt;Q11. What does survey evidence from the RIL data contribute to the interpretation?&lt;/h3&gt;
&lt;p&gt;The RIL (Rilevazione Longitudinale su Imprese e Lavoro), a representative firm survey collected in 2005, provides direct evidence on firms&amp;rsquo; stated reasons for their apprenticeship hiring decisions. Among firms that do not hire apprentices, the most common reason by far is &amp;lsquo;we don&amp;rsquo;t need more people,&amp;rsquo; with cost cited rarely. Among firms that do hire apprentices, the dominant reason is to train workers prior to hiring them as permanent employees; &amp;rsquo;lower labor costs&amp;rsquo; is a secondary consideration. This corroborates the paper&amp;rsquo;s interpretation that demand for apprentices is driven by training-for-retention motives rather than cost arbitrage, which explains why a cost reduction leaves hiring behavior unchanged.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-policy-recommendation-and-its-scope"&gt;Q12. What is the policy recommendation and its scope?&lt;/h3&gt;
&lt;p&gt;The paper urges caution in using payroll tax credits to stimulate employment, particularly for targeted groups with inherently low or inelastic labor demand. The results suggest that, for apprentices, firms hire based on training-and-conversion needs rather than cost considerations, so subsidizing cost does not expand hiring. More broadly, the cross-study cost-effectiveness comparison suggests that hiring credits — which require a new hire as a prerequisite for receiving the subsidy — may be more efficient than payroll tax cuts precisely because they screen out inframarginal firms. The paper does not rule out effectiveness for other worker types or for much larger subsidies, but the documented uniformity of null effects across industries, regions, and firm types suggests the inelasticity finding is robust within the studied population.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Inframarginal hiring&lt;/strong&gt;: Employment that would occur absent the subsidy; when a policy subsidizes inframarginal hiring, it transfers resources to firms without generating net new jobs, making it fiscally costly but behaviorally inert.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Difference-in-discontinuities&lt;/strong&gt;: An empirical design that combines regression discontinuity with difference-in-differences: in each period a discontinuity at the policy threshold is estimated, and the pre-policy baseline discontinuity is subtracted to remove pre-existing size-related level differences and time-invariant non-linearities in the conditional expectation function.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy-relevant firm size&lt;/strong&gt;: As defined by INPS under the 2007 Budget Bill: total full-time equivalent employment minus apprentices, temporary agency workers, workers on leave (unless replaced), and workers on specific on-the-job training contracts; this is the running variable determining SSC eligibility.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cost-effectiveness (jobs per foregone revenue)&lt;/strong&gt;: The number of job-years supported per unit of foregone tax revenue (here, per 1 million euros of lost SSCs), formally estimated via instrumental variables to allow statistical inference — as opposed to back-of-the-envelope calculations that provide no standard errors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inelastic labor demand for apprentices&lt;/strong&gt;: In this paper&amp;rsquo;s sense: firms&amp;rsquo; demand for apprenticeship contracts does not respond to changes in their labor cost, because hiring decisions are driven by training-and-conversion motives (hiring to eventually retain as permanent employees) rather than by cost minimization at the margin.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rotation of the conditional expectation function&lt;/strong&gt;: A change over time in the slope of the relationship between an outcome (e.g., apprentice hiring) and the running variable (firm size); when the slope changes, standard DiD specifications that discretize firms into treated/control groups will spuriously detect a treatment effect even when the true policy effect is zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Transformation (apprentice to permanent contract)&lt;/strong&gt;: The event of a firm converting an existing apprenticeship contract into an open-ended (permanent) employment contract at the end of the apprenticeship; used as an alternative outcome to evaluate whether the subsidy increased the ultimate goal of permanent employment, not just temporary apprenticeships.&lt;/p&gt;</description></item><item><title>The Aggregate Costs of Uninsurable Business Risk</title><link>https://macropaperwarehouse.com/papers/the-aggregate-costs-of-uninsurable-business-risk/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-aggregate-costs-of-uninsurable-business-risk/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; A large literature argues that credit constraints are the dominant financial friction holding private businesses below their optimal scale, so that easing credit access would yield large aggregate efficiency gains. This paper challenges that view. Private businesses are also poorly diversified — their owners bear undiversifiable business-income risk — and the authors argue the macroeconomic costs of this lack of diversification are far larger than those of credit constraints. The crux is that entrepreneurs can limit risk exposure by operating at a smaller scale, so productive-but-poor entrepreneurs choose an inefficiently low scale and are unwilling to borrow to expand. Firm size is thus limited by risk, not by credit availability.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and setup.&lt;/strong&gt; The empirical analysis uses the historical Orbis dataset (Moody&amp;rsquo;s Bureau van Dijk), 1995–2019, focusing on Spain (best coverage; results extend to Italy, France, Norway, Portugal, Slovakia in the appendix). Output is value added; the sample is partnerships and private limited companies, excluding FIRE, public administration, defense, education. The final sample is 622,883 firms (6,298,358 firm-year observations), observed on average 10 years; the mean (median) firm has 12 (5) workers and 486 (151) thousand EUR value added. The Spanish Survey of Household Finances (EFF, 2008–2020) provides entrepreneur wealth/prevalence and consumption data. The model is a small-open-economy model of entrepreneurial dynamics (à la Quadrini 2000; Cagetti–De Nardi 2006) with two frictions: each firm is owned by a single (undiversified) entrepreneur, and a collateral constraint k&amp;rsquo; ≤ a&amp;rsquo;/(1−ξ). Key modeling choices: capital AND labor are chosen before productivity is observed (time-to-build), and productivity has persistent and transitory shocks drawn from fat-tailed mixtures of normals. Parameters are estimated by simulated method of moments (9 parameters, 16 moments; objective 0.013, ~1.3% average deviation).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main quantitative findings.&lt;/strong&gt; Profit shares fluctuate sharply: 5% of firms have losses exceeding 20% of output, against an average profit share of 0.13; the 5th percentile of profit-share deviations is −0.33 and the 95th is +0.47. Output growth is fat-tailed (s.d. 0.48, IQR/s.d. ratio 0.65 vs 1.35 Gaussian; excess kurtosis 10.7). Inputs do not track output: regressing wage-bill growth on output growth gives 0.40 (capital 0.16); restricting to |Δlog y|&amp;lt;0.5 gives 0.58 and 0.31. A change in profit share on output growth has slope 1.56 (0.46 in the restricted sample). The headline result: eliminating both frictions would raise output by 15.8%; eliminating the risk wedge alone raises output by 15.4%, while eliminating the credit wedge alone raises output by only 0.4%. Misallocation losses are 10.8% (11.0% due to risk, 0.2% due to credit). Aggregate wedges are equivalent to a 12.8% tax on labor and 14.9% on capital. Wage losses are 27.8% (26.4% risk, 0.4% credit).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanisms and implications.&lt;/strong&gt; Two wedges distort choices: a risk wedge (from the covariance of consumption and productivity) that distorts both labor and capital, and a credit wedge (from the binding collateral constraint) that distorts only capital. The credit wedge falls quickly with wealth (vanishing once unconstrained), but the risk wedge declines only gradually and persists even for wealthy entrepreneurs. Aggregate losses are governed by the distribution of wedges weighted by efficient firm size (Hopenhayn 2014): risk wedges are large precisely for high-ability entrepreneurs who would be large under efficiency, whereas credit-constrained firms are mostly unproductive with small efficient size. Policy implication: improving credit access has limited impact unless it also improves risk sharing. The findings also imply firm profits largely reflect compensation for risk (75% of the aggregate profit share), and dispersion in returns to business wealth largely reflects risk compensation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-model-and-how-are-parameters-pinned-down"&gt;Q1. What is the identification strategy for the model, and how are parameters pinned down?&lt;/h3&gt;
&lt;p&gt;Parameters ϑ=(β,α,η,ρ,σu,σε,s,p,ϕ) are estimated by simulated method of moments, minimizing a weighted distance between 16 empirical and model moments scaled by 1+empirical moment (objective = 0.013, ~1.3% average deviation). Intuitively: β is pinned by the entrepreneur wealth-to-income ratio (12.5 in data and model); α and η by the capital-output ratio (1.22 vs 1.21), labor share (0.72 vs 0.71) and profit share (0.13 vs 0.14); ρ, σu, σε by output autocorrelations at horizons 1–3, the cross-sectional s.d. of output, and the s.d. of output growth at horizons 1–3; the tail parameters s and p by the IQR of output growth relative to its s.d.; and ϕ by the entrepreneurship rate. Three assigned parameters: δ=0.10, r=0.02, θ=2, with ξ=0.408 set to match the aggregate debt-to-capital ratio of 0.408. Standard errors (bootstrapped) are small because the firm sample is very large.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-main-mechanism-and-how-are-the-risk-wedge-and-credit-wedge-distinguished"&gt;Q2. What is the main mechanism, and how are the risk wedge and credit wedge distinguished?&lt;/h3&gt;
&lt;p&gt;Because labor and capital are chosen before productivity is realized and risk is undiversified, the entrepreneur weights future states by their own stochastic discount factor. The risk wedge τ (&amp;gt;1) arises from the negative covariance between marginal utility of consumption and productivity and distorts both labor and capital equally. The credit wedge ω (&amp;gt;1 when the collateral constraint binds) distorts only capital. As wealth rises, the credit wedge falls rapidly and vanishes once the firm is unconstrained, but the risk wedge declines only gradually and never disappears. The two are isolated quantitatively by setting ω=1 (to get the role of risk) or τ=1 (to get the role of credit) in the productivity-loss mapping (eq. 13).&lt;/p&gt;
&lt;h3 id="q3-why-does-risk-dominate-credit-in-the-aggregate-even-though-most-firms-are-credit-constrained"&gt;Q3. Why does risk dominate credit in the aggregate even though most firms are credit-constrained?&lt;/h3&gt;
&lt;p&gt;Aggregate outcomes depend on the distribution of wedges weighted by efficient firm size n_it (Hopenhayn 2014). Weighted by efficient size, the risk wedge ranges from 1.27 (10th pct) to 1.61 (90th pct), while the credit wedge is essentially 1 except at the very top (1.02 at the 90th pct). Unweighted, the risk wedge is only 1.12 at the 90th pct and the credit wedge is positive for more than half of firms — but those constrained firms are unproductive with small efficient size. Risk wedges are large precisely for high-ability entrepreneurs who would be large under the efficient allocation, so they drive the aggregate.&lt;/p&gt;
&lt;h3 id="q4-why-is-the-result-robust-to-the-form-of-the-collateral-constraint"&gt;Q4. Why is the result robust to the form of the collateral constraint?&lt;/h3&gt;
&lt;p&gt;The authors consider two extremes: no borrowing at all (ξ=0) and unlimited borrowing (ξ=1, no credit limit). With no borrowing, misallocation losses rise only from 10.8% to 11.7%, still mostly risk-driven (8.3% risk vs 1.4% credit). With no credit limit, risk wedges remain nearly as large as baseline and removing credit frictions has negligible effects. Intuitively, risk leads entrepreneurs to operate small and accumulate precautionary wealth, so they self-finance most desired capital and credit wedges stay small even without credit.&lt;/p&gt;
&lt;h3 id="q5-which-three-ingredients-are-essential-to-the-risk-dominates-result-and-what-happens-without-each"&gt;Q5. Which three ingredients are essential to the risk-dominates result, and what happens without each?&lt;/h3&gt;
&lt;p&gt;(1) Fat-tailed productivity shocks, (2) transitory productivity shocks, and (3) labor chosen before productivity is realized. Removing each in isolation (with re-estimation) reverses the conclusion so that credit becomes the primary driver: without fat tails, misallocation losses fall to 2.1% (credit 1.5%, risk 0.3%); without transitory shocks, losses are 12.1% (credit 10.9%, risk 0.4%); with flexible labor, losses fall to 3.3% (credit 2.4%, risk 0.1%). The flexible-labor case matters because risk then distorts only capital, whose share is smaller than labor&amp;rsquo;s, reducing income volatility and pushing firms to expand and hit the credit constraint. In all three counterfactuals, the 1st percentile of profit-share deviations ranges −0.21 to −0.43, far smaller in magnitude than the data (−1.66) or baseline model (−1.92).&lt;/p&gt;
&lt;h3 id="q6-is-the-result-driven-by-high-risk-aversion"&gt;Q6. Is the result driven by high risk aversion?&lt;/h3&gt;
&lt;p&gt;No. The baseline uses relative risk aversion θ=2. Re-estimating with θ=0.5 (low end of usual values) still yields sizable, risk-dominated losses: productivity losses 6.4%, output losses 9.2%, wage losses 16.7% — roughly three-fifths of the baseline — and again primarily driven by risk rather than credit.&lt;/p&gt;
&lt;h3 id="q7-what-untargeted-moments-does-the-model-match-model-validation"&gt;Q7. What untargeted moments does the model match (model validation)?&lt;/h3&gt;
&lt;p&gt;The model reproduces the distribution of profit-share deviations (10th pct −0.17 data vs −0.16 model; 1st pct −1.66 data vs −1.92 model), the full distribution of output growth rates, the low wage-bill/output comovement (0.58 data vs 0.55 model in the restricted sample), the profit-share/output comovement (0.46 vs 0.42; falling to 0.10 vs 0.06 when holding the labor share constant), and the persistence/volatility of capital and labor (e.g., wage-bill growth s.d. 0.36 vs 0.32). Critically, it matches the low comovement of entrepreneur consumption with profits: regressing Δc on Δπ gives a slope of 0.02 in both data and model (data based on 799 EFF observations, three-year changes).&lt;/p&gt;
&lt;h3 id="q8-what-heterogeneity-and-external-validity-does-the-paper-document"&gt;Q8. What heterogeneity and external validity does the paper document?&lt;/h3&gt;
&lt;p&gt;The motivating facts hold for Italy, France, Norway, Portugal and Slovakia, and for Spanish public firms; for young (age≤5) and old firms; for small and large firms (top decile of value added vs rest); and across the five largest sectors (manufacturing, construction, wholesale/retail, accommodation/food, professional activities). Output-growth kurtosis ranges roughly 11–18 across countries. On diversification: 12% of households are entrepreneurs; 93% of entrepreneurs own exactly one business; multi-business owners hold 71% of their business wealth in their main business; the average ownership share is 83%, and 71% own 100% of their main business.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-extensive-margin-and-unconstrained-firm-results"&gt;Q9. What are the extensive-margin and unconstrained-firm results?&lt;/h3&gt;
&lt;p&gt;Extensive margin: when the planner can also choose who becomes an entrepreneur, it cuts the entrepreneurship rate from 13.2% to 1.2%, but because marginal entrepreneurs are low-ability the gains are small — productivity, output and wage losses relative to the unconstrained planner are 10.8%, 16% and 27.8%, very close to the intensive-margin numbers. Unconstrained firms: adding a frictionless sector calibrated to match the 58.7% output share of public firms in Orbis leaves misallocation losses at 10.5% (vs 10.8% baseline), still mostly risk-driven (risk 10.1%, credit 0.1%); wage losses fall to about three-fifths of baseline because the unconstrained sector reduces the aggregate labor wedge.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-implications-for-profits-and-returns-to-wealth"&gt;Q10. What are the implications for profits and returns to wealth?&lt;/h3&gt;
&lt;p&gt;Decomposing the profit share into span-of-control, risk and credit components: risk accounts for 75% of the aggregate profit share (0.11/0.146), with the rest from span of control; credit contributes little. Risk also drives most of the profit-share dispersion (s.d. 5.5%, essentially all from risk; credit contributes only 1%). For excess returns to wealth, the mean of 2.2% is almost entirely accounted for by risk, and risk drives most of the dispersion (s.d. 5.5%). This implies dispersion in returns to private business wealth — a driver of wealth inequality — largely reflects compensation for risk rather than credit constraints.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-working-capital-robustness-check"&gt;Q11. What is the working-capital robustness check?&lt;/h3&gt;
&lt;p&gt;Adding a working-capital constraint where a fraction ϑ=0.25 of the wage bill is paid in advance (à la Mendoza 2010), evaluated at baseline parameters, gives misallocation losses of 11.1% (vs 10.8% baseline), with risk still accounting for the bulk (9.4%) and credit less important (1.3%); risk accounts for 13.4% of the 16.3% total output losses. So even when credit frictions can also distort labor, risk remains dominant.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central implication is that policies expanding firms&amp;rsquo; access to credit will have limited aggregate impact unless they also improve risk sharing. This holds within the scope of the model — undiversified private businesses with single owners, where risk exposure is endogenously chosen via scale and can be partly self-insured through wealth, labor income, and occupational switching. The authors note their framework assumes (rather than micro-founds) the lack of diversification, and suggest future work should model the moral-hazard or informational frictions preventing diversification, and broaden redistributive tax analysis to incorporate uninsurable-risk distortions (as in Di Tella et al. 2024).&lt;/p&gt;
&lt;h3 id="q13-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q13. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It contributes to the misallocation literature (Hsieh-Klenow 2009; Buera et al. 2011; Moll 2014; Midrigan-Xu 2014; Gopinath et al. 2017). Prior work on risk and investment (Tan 2018; Robinson 2021; David et al. 2022a) studies how risk distorts investment; this paper instead emphasizes how risk distorts LABOR choices, relating it to Arellano et al. (2019) and David et al. (2022b). It differs from the credit-constraint-centric tradition by showing credit matters little once undiversified risk and the three key ingredients are present. Di Tella et al. (2024), partly motivated by these findings, study optimal policy under uninsurable risk and show it is the opposite of optimal policy when misallocation stems from markups.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Risk wedge (τ)&lt;/strong&gt;: In the paper&amp;rsquo;s sense, the gap between the expected marginal product of an input and its price arising from undiversifiable business risk. It equals [1 + COV(c^{-θ}, zε)/(E c^{-θ} · E zε)]^{-1}, generally &amp;gt;1 because of the negative covariance between the entrepreneur&amp;rsquo;s marginal utility of consumption and productivity. It distorts both labor and capital, declines only gradually with wealth, and persists even for wealthy entrepreneurs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Credit wedge (ω)&lt;/strong&gt;: The distortion from a binding collateral constraint, ω=1+(1−ξ)μ/R, where μ is the multiplier on the constraint k&amp;rsquo;≤a&amp;rsquo;/(1−ξ). It exceeds one only when the constraint binds, distorts only capital, falls rapidly with wealth, and vanishes once the entrepreneur is unconstrained.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Profit share&lt;/strong&gt;: In this paper, the ratio of profits to output (value added), π_it/y_it, where profit is output net of the wage bill and the user cost of capital. Its average is 0.13; the paper studies its large transitory firm-level fluctuations as the empirical signature of uninsurable risk.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time-to-build (inputs chosen before productivity)&lt;/strong&gt;: The assumption that both capital and labor are chosen before the firm observes its productivity shock. This parsimoniously generates the imperfect high-frequency comovement between inputs and output and makes wealth affect employment as well as investment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Efficient-size-weighted wedge distribution&lt;/strong&gt;: The paper&amp;rsquo;s organizing device (following Hopenhayn 2014): aggregate productivity losses depend on the distribution of risk and credit wedges weighted by each firm&amp;rsquo;s efficient size n_it. Because high-ability firms have large efficient size and large risk wedges, risk dominates the aggregate even though most firms are credit-constrained.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Self-financing&lt;/strong&gt;: The mechanism by which entrepreneurs, operating at small scale and saving for precautionary reasons because of risk, accumulate enough wealth to finance most of their desired capital — so credit wedges stay small even in an economy with no credit, rendering the borrowing limit nearly irrelevant for aggregates.&lt;/p&gt;</description></item><item><title>The Lost Marie Curies and Foregone Economic Growth</title><link>https://macropaperwarehouse.com/papers/the-lost-marie-curies-and-foregone-economic-growth/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-lost-marie-curies-and-foregone-economic-growth/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Women accounted for only 3% of U.S. inventors in 1976 and still just 14% in 2023, a pace of convergence far slower than in law (3% to 49%) or medicine (6% to 46%) over the same period. Under the natural assumption of no innate gender differences in inventive potential, this persistent underrepresentation reveals a misallocation of talent. The paper asks how costly this misallocation is for aggregate productivity and welfare.&lt;/p&gt;
&lt;p&gt;Brouillette develops an overlapping-generations (OLG) model of semi-endogenous growth in the spirit of Jones (1995), in which individuals with heterogeneous innate inventive talent choose sequentially among three decisions: (1) whether to pursue a STEM education (the prerequisite for research), (2) whether to work in research or production, and (3) whether to have children. Three gendered barriers can deter women from their comparative advantage. First, a labor market distortion, modeled as a tax on research earnings, captures discrimination in pay and credit attribution. Second, a child penalty distortion reduces mothers&amp;rsquo; hours in research relative to fathers, amplified by the &amp;ldquo;greedy job&amp;rdquo; nature of research (a premium on long hours). Third, an exposure distortion, modeled as a Bernoulli random variable, captures the probability of ever encountering inventive career opportunities — driven empirically by the absence of female role models.&lt;/p&gt;
&lt;p&gt;The model is calibrated to the U.S. economy using two data sources: PatentsView (all USPTO patents since 1976, covering roughly 1.7 million inventors and 3.7 million patents, with gender inferred from first names) and the U.S. Decennial Census/ACS (demographic and occupational data). Across these sources, female inventors exhibit only marginally higher research productivity than men (consistent with modest positive selection from the earnings tax), while mothers in research work approximately 4.5% fewer hours per week than childless female researchers (fathers work 2.7% more). The small productivity gap and modest hours gap together imply that neither the earnings tax nor the child penalty is the dominant driver; the exposure distortion is inferred as the residual, calibrated to a benchmark female share in research of 23% (average of 19% from PatentsView and 27% from Census/ACS). The resulting distortion estimates are: labor market tax 3.3%, child penalty 7%, and exposure barrier 79%.&lt;/p&gt;
&lt;p&gt;Counterfactual elimination of all three distortions raises U.S. income per person by 14.2% in the long run, compared with only 1.5% from a 30% R&amp;amp;D subsidy in a distortion-free economy. The gain materializes slowly, with a half-life of approximately 76 years, reflecting the semi-endogenous structure (where reallocating talent shifts the level but not the long-run growth rate of living standards) and the OLG structure (where career choices are irreversible, slowing labor reallocation). Aggregate research labor increases by 49% within the first 50 years of the transition — women&amp;rsquo;s research labor more than quadruples while men&amp;rsquo;s shrinks by about 10% — but almost all of the productivity gain operates through the intensive rather than the extensive margin: the aggregate share of inventors barely rises, because exposure barriers blocked many talented women entirely rather than only marginal ones, so lifting them introduces very high-quality new researchers who crowd out less talented men. If the underrepresentation were instead attributed entirely to selection-based barriers (labor market or child penalty), long-run consumption would rise by only 3.6%, less than a quarter of the baseline 14.2%.&lt;/p&gt;
&lt;p&gt;Taking transition dynamics into account, eliminating all distortions is equivalent to permanently raising everyone&amp;rsquo;s consumption by 7.2% (lower than 14.2% because the transition is slow and future gains are discounted back at a rate exceeding the low projected U.S. population growth). Of this welfare gain, 95% comes from higher mean consumption; the remainder comes from reduced consumption inequality and utility from children. The distribution of gains is unequal across time and demographic groups: future cohorts experience an 8.6% permanent consumption increase versus only 1% for surviving cohorts. Among the current generation of inventors, women gain the equivalent of a 1.3% permanent consumption increase while men lose 1.7%, a distributional tension that complicates implementation when current costs are concentrated and future benefits diffuse.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-three-distortions-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy for the three distortions, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The three distortions are identified from three moments, each theoretically linked to a specific distortion through the model&amp;rsquo;s aggregation. The labor market distortion (earnings tax) is identified from the research productivity gender gap: positive selection under this tax implies women should be marginally more productive, and the magnitude of the observed (small) gap pins down a distortion of 3.3%. The child penalty distortion is identified from gender differences in hours worked between parent and non-parent researchers: mothers work 4.5% fewer hours than childless women while fathers work 2.7% more; after normalizing male distortions to zero, the model recovers a child penalty distortion of 7%. The exposure distortion is identified as the residual that explains remaining underrepresentation (23% female share in research) after accounting for the other two mechanisms; it is estimated at 79%. Key threats: (1) The gender productivity gap is measured from PatentsView, which uses name-based gender attribution and citation-weighted patents — both susceptible to gender bias (women are documented to receive 30% fewer citations than men with common names, and are 59% less likely to be credited with authorship on patents they contributed to), so the paper uses stock market valuation and textual similarity of patents as bias-resistant alternatives. (2) The exposure distortion is a residual and could capture other forces not in the model, including occupational preferences, gendered barriers to human capital retention, or mismeasurement of the female researcher share. (3) The model abstracts from the direction of innovation (unlike Einïo, Feng, and Jaravel 2022), so welfare effects through consumption-cost inequality across groups are not captured.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The three mechanisms operate through distinct theoretical channels, which allows moment-based identification. The labor market distortion works through selection on talent: if only highly talented women choose research despite earning below their marginal product, the female researcher pool should be right-shifted in the talent distribution, implying modestly higher measured productivity for women. The empirical counterpart is the gender gap in patent output (quality-weighted patents per career year), controlling for field fixed effects and team size. The child penalty works through hours worked: a higher opportunity cost of childbearing in research (amplified by greedy-job premiums) reduces mothers&amp;rsquo; time in research. The empirical counterpart is the gender gap in hours worked between parents and non-parents in research, from the Census/ACS. The exposure distortion works through the extensive margin of talent — it is a binary probability of ever having access to research as a career path, so it can block even the most talented women, unlike the other two distortions which induce selection. It is identified as the residual after the other two are estimated. The insight that the productivity gap is small and the hours gap is modest together rule out the first two as primary drivers, placing most explanatory weight on the exposure distortion.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-semi-endogenous-growth-framework-differ-from-an-endogenous-growth-approach-and-what-are-the-implications-for-the-results"&gt;Q3. How does the semi-endogenous growth framework differ from an endogenous growth approach, and what are the implications for the results?&lt;/h3&gt;
&lt;p&gt;In semi-endogenous growth (Jones 1995), the long-run per-capita growth rate equals n/[(sigma-1)(1-phi)], determined by population growth and idea difficulty, not by the quantity or quality of researchers. A reallocation of inventive talent therefore cannot raise the long-run growth rate but can raise the level of per-capita consumption by shifting the cumulative stock of ideas and thus the entire trajectory of living standards upward. This stands in contrast to endogenous growth models where reallocating talent can permanently raise the growth rate. The author justifies the semi-endogenous approach on two grounds: (1) despite sustained researcher-population growth in most advanced economies, the per-capita growth rate has not trended up; (2) the framework is qualitatively and quantitatively consistent with the documented fact that &amp;lsquo;ideas are getting harder to find&amp;rsquo; (Bloom et al. 2020, which estimates phi = -2.1 for the aggregate U.S. economy). The implication is that the paper finds more modest effects on productivity growth than prior endogenous-growth models, with the gain materializing entirely as a level shift with a long half-life of ~76 years. Einïo, Feng, and Jaravel (2022), using an endogenous growth model, find that barriers to female innovation reduce the growth rate by 1.4 percentage points; this paper&amp;rsquo;s semi-endogenous model finds a 14.2% level gain with no permanent growth rate effect.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-in-the-gender-gap-is-documented-empirically"&gt;Q4. What heterogeneity in the gender gap is documented empirically?&lt;/h3&gt;
&lt;p&gt;Field heterogeneity: Between the 1990 and 2020 inventor cohorts, the female share in chemistry and metallurgy rose from 13% to approximately 30%, while in fixed constructions and mechanical engineering it rose from under 5% to about 10%. Despite this, male-dominated fields accounted for about 53% of total patents granted in 2023. Importantly, when the inventive productivity gender gap is plotted against the female share across technological fields and cohorts, there is no significant relationship (the slope is -0.09 with a standard error of 0.2), implying selection-based barriers are not the primary driver of field-level disparities. Cohort heterogeneity: By cohort, the female share among new inventors rose from 7.5% (1990 cohort) to 17.6% (2020 cohort). Life-cycle heterogeneity: The inventive productivity gender gap (with women slightly ahead) is primarily a cohort effect rather than a within-career pattern; more recent cohorts show a somewhat larger productivity advantage for women at career onset, but the magnitude remains modest, which argues against gendered human capital depreciation as a leading explanation. Parental status heterogeneity: The fraction of female researchers who are mothers converged to the fraction of male researchers who are fathers over time (both around 40% by 2023, down from an 80% male vs. 40% female gap in 1960), suggesting research has become more accommodating. The child penalty in research (hours worked differential between parents and non-parents) has also narrowed over time and is smaller in research than in non-research occupations.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-conducted"&gt;Q5. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;Five sets of robustness exercises are reported. (1) Degree of increasing returns to scale (gamma): Jones (2002) estimates gamma from 0.05 to 0.33; Peters (2021) estimates 0.6. Across this range, the long-run consumption gain from eliminating all distortions ranges from about 2% to almost 27% for gamma going from 0.05 to 0.6. (2) Talent signal shape parameter (theta_s): With theta_s raised to 2 from the baseline 1.26 (implying greater scarcity of superstar inventors, so fewer marginal researchers are displaced), the long-run gain falls to 8.7% from 14.2%. (3) Demographic parameters (retirement rate d and entry rate b): Setting d to match expected working lives of 20 and 40 years (versus baseline 30) shifts the transition half-life by roughly 6-8 years, leaving long-run income unchanged but moving welfare gains slightly (7.6% or 6.9% vs. baseline 7.2%). (4) Knowledge spillover parameter (phi): Values of 0.5 and -6.2 (lower bound of Bloom et al.) are tested with sigma adjusted to hold gamma constant; long-run income gains remain at 14.2%, while the half-life varies modestly and welfare gains shift by at most 24 basis points. (5) Patent quality metrics: Three alternative measures of patent quality are used — stock market valuation (Kogan et al. 2017), textual &amp;lsquo;importance&amp;rsquo; (Kelly et al. 2021), forward citations, and unweighted counts. Results are consistent across measures, with the bias-resistant metrics (stock market valuation and textual importance) ruling out citation-based bias as a confound.&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-relate-to-and-differ-from-einïo-feng-and-jaravel-2022"&gt;Q6. How does this paper relate to and differ from Einïo, Feng, and Jaravel (2022)?&lt;/h3&gt;
&lt;p&gt;Einïo et al. (2022) is the closest antecedent. That paper develops a two-sector endogenous growth model with heterogeneous consumer tastes and unequal access to innovation across sociodemographic groups including gender, finding that barriers to female innovation are responsible for an 18.2% difference in the cost of living between women and men and reduce the economic growth rate by 1.4 percentage points. Brouillette&amp;rsquo;s paper uses a semi-endogenous growth framework and arrives at a 14.2% long-run level gain in income per person and a 7.2% consumption-equivalent welfare gain, with no permanent effect on the growth rate. Beyond the growth framework, the paper extends the analysis to include labor market discrimination and a child penalty for female researchers, which Einïo et al. do not model. However, Brouillette&amp;rsquo;s model abstracts from the direction of innovation — the idea that women and men produce inventions differently tailored to different users&amp;rsquo; needs — which Einïo et al. show is quantitatively important for cost-of-living inequality. The two papers are therefore treated as providing complementary insights.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-role-model-externality-extension-and-how-does-it-change-the-results"&gt;Q7. What is the role-model externality extension, and how does it change the results?&lt;/h3&gt;
&lt;p&gt;In the baseline model, the exposure distortion is a fixed parameter representing the probability of ever encountering inventive career opportunities. In the extension, this probability is multiplied by a technology friction that depends on the fraction of same-gender and opposite-gender inventors in prior generations, with elasticities calibrated from Bell et al. (2018): own-gender elasticity 0.24 for girls, cross-gender elasticity approximately 0 (statistically insignificant in the underlying regression). This creates a positive externality: current inventors increase exposure probabilities for future cohorts of the same gender, but they are not compensated for this spillover, constituting a market failure. In the extended model, some of what was previously captured as the exposure distortion is now attributed to the technological friction from role model scarcity, and the residual exposure distortion is smaller. The counterfactual elimination of all distortions yields a more modest long-run income gain of 10.6% and a consumption-equivalent welfare gain of 3.8% (compared to 14.2% and 7.2% in the baseline). The role model externality also opens a rationale for temporarily gender-differentiated wage subsidies for female researchers as transitional optimal policy: a welfare-maximizing planner might accept a slightly worse talent allocation today in order to accelerate the expansion of the female role model base, reaching the efficient allocation sooner.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s central policy implication is that interventions targeting exposure to innovation for girls earlier in the pipeline — before entry into the labor market — offer far larger aggregate productivity returns than either conventional R&amp;amp;D subsidies or policies aimed at reducing workplace discrimination or the child penalty in isolation. A 30% R&amp;amp;D subsidy yields only 1.5% long-run income per capita growth versus 14.2% from full elimination of female research barriers. Within those barriers, the exposure distortion alone accounts for the bulk of the gain: if the underrepresentation were entirely due to the labor market or child penalty distortions (selection-based mechanisms), long-run gains would be only 3.6%. Scope conditions and caveats: (1) The framework is calibrated to the U.S. and to patent-based inventors plus Census-classified researchers, so generalization to other settings requires re-estimation of distortions. (2) The semi-endogenous structure implies that gains are level effects, not growth rate effects, and the half-life of ~76 years means that most gains accrue to future rather than current generations. (3) Distributional effects are asymmetric: the current generation of male inventors suffers a 1.7% consumption loss, while future cohorts broadly gain 8.6%; this temporal and demographic incidence complicates implementation. (4) The model abstracts from the direction of innovation, so welfare effects through differential cost-of-living impacts on men and women are not captured. (5) The role model externality extension suggests that affirmative action policies for female researchers may be warranted on efficiency grounds, but the exact form of optimal transitional policy is not fully characterized.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-greedy-job-mechanism-and-how-is-it-quantified"&gt;Q9. What is the &amp;lsquo;greedy job&amp;rsquo; mechanism and how is it quantified?&lt;/h3&gt;
&lt;p&gt;The &amp;lsquo;greedy job&amp;rsquo; concept (Goldin 2021) refers to occupations where extended, inflexible hours are compensated at a premium, making it suboptimal for couples to share labor supply equally and thus imposing a larger effective cost of parenthood on whoever reduces hours (in practice, more often women). In the model, an individual researcher&amp;rsquo;s effective labor supply is proportional to alpha^(1+delta) when they have children (where alpha = 0.93 is the fraction of time parents spend working and delta &amp;gt; 0 governs the additional return to hours in research). This magnifies the talent threshold required for a parent to prefer research over production. The parameter delta is estimated empirically by regressing log hourly wages on log hours worked, an indicator for research occupation, and their interaction (plus controls for age, experience, education, occupation, state, race, marital status, year, gender, and occupation-by-gender fixed effects), using the Census/ACS with over 11.8 million observations. The estimated delta for researchers is 0.004, statistically significant but modest — implying research is a &amp;lsquo;modestly greedy job,&amp;rsquo; less so than law (0.011) or medicine (0.006). This small value of delta constrains the child penalty distortion&amp;rsquo;s aggregate impact and helps explain why the exposure distortion dominates empirically.&lt;/p&gt;
&lt;h3 id="q10-how-is-research-productivity-measured-and-what-biases-are-addressed"&gt;Q10. How is research productivity measured, and what biases are addressed?&lt;/h3&gt;
&lt;p&gt;Research productivity is measured as average quality-weighted patents granted per year over an inventor&amp;rsquo;s career, with experience fixed effects removed before averaging across years. Three patent quality metrics are used: (1) stock market valuation (Kogan et al. 2017), inferred from abnormal stock returns around patent grant announcements — chosen for its resistance to gender bias because it reflects market assessments rather than subjective citation choices; (2) &amp;lsquo;importance&amp;rsquo; (Kelly et al. 2021), measured from textual similarity between patent pairs, rewarding novelty relative to prior patents and influence on subsequent ones, and also robust to citation bias because it would require precise paraphrase rather than mere omission; (3) forward citation counts, acknowledged as potentially biased (Jensen et al. 2018 show women with common names receive 30% fewer citations, while women with rare names receive 20% more); (4) unweighted patent counts. All metrics are adjusted for 3-digit CPC class fixed effects and co-inventorship team size. The results are consistent across all four measures, with women slightly ahead in all cases, suggesting that citation bias does not qualitatively alter the productivity comparison. A further concern is attribution bias: Ross et al. (2022) show women are 59% less likely to be credited with authorship on patents they contributed to, meaning PatentsView may undercount the true female inventor population.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-paper-say-about-the-stem-education-gender-gap-specifically"&gt;Q11. What does the paper say about the STEM education gender gap specifically?&lt;/h3&gt;
&lt;p&gt;Women account for approximately 35% of employed STEM graduates aged 25 to 45 in the Census/ACS data (and less than 20% of engineering graduates). However, this STEM gap alone explains only 7% of the patenting gender gap (Hunt et al. 2013, using the 2003 NSCG which recorded patenting in the prior five years); a substantial 78% of the gap stems from differences in patenting behavior among STEM graduates themselves. Furthermore, since the early 2000s, female researchers have been more likely than male researchers to hold a college degree, ruling out educational attainment differences as the primary driver. The model addresses STEM underrepresentation not through a gendered STEM education cost but through the exposure distortion, on the grounds that: (1) exposure to role models is well-documented as influencing girls&amp;rsquo; decisions to pursue STEM (Carrell et al. 2010; Breda et al. 2023; Bell et al. 2018); and (2) if a higher STEM cost were the primary barrier, the model would predict women to be substantially more productive than men (strong positive selection), which the data does not support.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Semi-endogenous growth&lt;/strong&gt;: A growth framework in which the long-run per-capita growth rate is determined by population growth and the difficulty of finding new ideas (the knowledge spillover parameter phi), not by the quantity or quality of researchers. Reallocating inventive talent shifts the level of living standards permanently but cannot alter the long-run growth rate; &amp;lsquo;ideas are getting harder to find&amp;rsquo; (phi &amp;lt; 0 in the paper&amp;rsquo;s calibration, phi = -2.1) is an integral feature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exposure distortion&lt;/strong&gt;: A Bernoulli random variable with mean (1 - tau_E_gk) governing whether an individual of gender g and cohort k ever encounters inventive career opportunities, regardless of their talent. In the baseline model it captures the aggregate probability of not having relevant role models or other enabling conditions during formative years; it is estimated at 79% for women (meaning only 21% of women are exposed to research as a potential career path). Unlike selection-based distortions, it blocks access to the innovation system even for the most talented women.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor market distortion&lt;/strong&gt;: A proportional tax tau_L on the research earnings of female inventors, representing discrimination in compensation, credit attribution, promotions, and rent-sharing from intellectual property. It induces positive selection: under this tax, only sufficiently talented women prefer research over production, making the average female researcher marginally more productive than the average male researcher. Estimated at 3.3%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Child penalty distortion&lt;/strong&gt;: A proportional reduction tau_C in the effective research hours of mothers, capturing the disproportionate burden of childcare and household responsibilities on women&amp;rsquo;s research careers. Combined with the &amp;lsquo;greedy work&amp;rsquo; parameter delta (the premium on long hours in research), it raises the talent threshold above which a woman who wants children will still choose a research career. Estimated at 7%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Greedy job&lt;/strong&gt;: An occupation, in the sense of Goldin (2021), where working long and inflexible hours is rewarded at a premium over and above what a simple proportional-hours model would predict. In the model, captured by the parameter delta &amp;gt; 0 in the research labor supply function. Estimated at delta = 0.004 for researchers (modest relative to lawyers at 0.011 or doctors at 0.006), implying that research is a modestly greedy job, amplifying the child penalty but not dominating the exposure distortion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive vs. extensive margin of research labor&lt;/strong&gt;: The extensive margin refers to the number (fraction) of people who choose research careers; the intensive margin refers to the average quality (talent-weighted hours) of researchers. The paper&amp;rsquo;s key finding is that the 14.2% long-run income gain from eliminating gender barriers is achieved almost entirely on the intensive margin: the aggregate share of inventors barely rises, but average researcher quality increases substantially because exposure barriers had been blocking the most talented women entirely.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption-equivalent welfare variation&lt;/strong&gt;: The permanent proportional adjustment lambda to every person&amp;rsquo;s consumption in the distorted economy that would make utilitarian social welfare equal to that in the undistorted economy. A lambda of 1.072 (7.2% gain) means permanently raising everyone&amp;rsquo;s consumption by 7.2% would compensate for remaining in the distorted equilibrium rather than transitioning to the undistorted one. It is lower than the 14.2% long-run income gain because the slow transition and the discounting of future population growth reduce the present value of future gains.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inventive productivity gender gap&lt;/strong&gt;: The difference in average quality-weighted patents per year between female and male inventors, after controlling for technological field fixed effects, experience, and co-inventorship team size. Measured across multiple patent quality metrics (stock market valuation, textual importance, forward citations, unweighted counts). In the paper&amp;rsquo;s data, the gap is positive but small — women are slightly more productive — which is the key empirical moment used to identify the (small) labor market distortion and to rule out large selection-based barriers as the primary driver of underrepresentation.&lt;/p&gt;</description></item><item><title>The macroeconomics of automation</title><link>https://macropaperwarehouse.com/papers/the-macroeconomics-of-automation/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-macroeconomics-of-automation/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks a foundational question: can the economy-wide degree of automation be measured coherently from standard macroeconomic data, without relying on technology-specific proxies such as robot counts or AI investment surveys? Existing micro-level proxies are fragmented across technologies and difficult to aggregate, leaving it unclear how automation evolves at the macro level or how it relates to capital deepening, factor shares, and productivity growth. The authors, Hideki Nakamura, Masakatsu Nakamura, and Shota Moriwaki, address this by developing a task-based general equilibrium framework in which the aggregate degree of automation emerges endogenously and is fully identified from observable macroeconomic aggregates.&lt;/p&gt;
&lt;p&gt;The theoretical architecture begins with a continuum of tasks, each exhibiting Leontief technology at the task level. Within each task, capital and labor are perfectly substitutable, but firms choose the least-cost input given factor prices. Tasks are ordered by the relative efficiency of capital to labor; as the wage-to-capital-service-price ratio rises with capital deepening, capital performs an expanding range of tasks. Aggregating task-level Leontief decisions over a firm generates a global (envelope) production function. The paper&amp;rsquo;s first main theorem shows that under a mild regularity condition on task efficiency orderings, this aggregation delivers a standard neoclassical production function. Its second set of results identifies the precise efficiency structure under which the aggregate function takes the CES form: that structure corresponds to a Pareto cumulative distribution of input efficiencies. This Pareto structure yields a clean closed-form relationship: the degree of automation is determined entirely by the capital-labor ratio (in efficiency units) and the elasticity of substitution. When the elasticity exceeds one, the degree of automation equals the capital income share; when the elasticity falls below one, it equals the labor income share. Neutral technical progress leaves the degree of automation unchanged at a given capital-labor ratio; capital-augmenting progress raises it; labor-augmenting progress lowers it.&lt;/p&gt;
&lt;p&gt;The empirical application uses panel data from the 2023 Japan Industrial Productivity (JIP) database covering 52 manufacturing industries from 1994 to 2020 (N = 1,404 industry-year observations; two industries excluded for data quality). The CES production function is estimated via GMM using first-differenced factor-share equations derived from the normalized CES system (de La Grandville 1989 normalization), with five sets of instrumental variables drawn from lagged factor prices, information stock and its price, trade openness, workforce age composition, and part-time employment shares.&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. Under the assumption of neutral technical progress, the elasticity of substitution sigma is significantly above one but close to one, ranging from 1.049 to 1.102 across the five IV sets (all significant at least at the 10 percent level). Under the assumption of capital-augmenting technical progress (gK &amp;gt; 0, gL = 0), sigma ranges from 1.035 to 1.068, again robustly greater than one. Capital-augmenting technical progress is statistically significant across all specifications; labor-augmenting technical progress cannot be confirmed in any specification. The average estimated degree of automation across the 52 industries over the full sample period is 0.417 (standard deviation 0.171, minimum 0.138, maximum 0.811). The average rises steadily from 0.407 in 1994 to 0.426 in 2020, temporarily declining around the 2008 financial crisis before recovering. Substantial heterogeneity persists across industries throughout the sample. The distribution shifts rightward over time but retains a fat left tail, with the mode just above 0.3 and several industries exceeding 0.7.&lt;/p&gt;
&lt;p&gt;The two-level CES extension decomposes aggregate capital into industrial robots and other capital, exploiting a purpose-built robot capital stock constructed via the RAS and perpetual inventory methods (initial year 1985). Industrial robots account for only 0.44 percent of aggregate capital stock on average. The two-level estimation yields higher elasticities (sigma-a between 1.191 and 1.346 across IV sets for the composite-labor margin; sigma-b between 1.049 and 1.096 for the robots-other-capital margin). The degree of automation for the composite rises from 0.398 to 0.430 over the sample, a more pronounced increase than the standard CES estimate, reflecting robots&amp;rsquo; amplifying role in automation.&lt;/p&gt;
&lt;p&gt;The paper benchmarks three automation measures against an internal consistency criterion: the squared distance between the automation degree inferred from the capital-labor ratio and that inferred from output per worker, given the same CES structure. The Pareto-based measure (the paper&amp;rsquo;s preferred measure) achieves a distance of 0.0000319, far below the Cobb-Douglas alternative (0.002484) and the continuity-preserving alternative (0.00999), validating the Pareto efficiency-distribution assumption. The Cobb-Douglas alternative yields a mean automation of 0.500 rising from 0.454 to 0.529; the continuity alternative rises more sharply from 0.208 to 0.589 but is discontinuous and sometimes falls outside the unit interval.&lt;/p&gt;
&lt;p&gt;For policy and theory, the paper&amp;rsquo;s framework implies that Japan&amp;rsquo;s sustained capital accumulation during its prolonged stagnation after 1990 translated into rising automation even without commensurate TFP growth, connecting automation dynamics to the &amp;ldquo;productivity paradox.&amp;rdquo; The model also shows that automation can rise alongside an increasing labor income share when sigma is below one, caution against interpreting a stable or rising labor share as evidence against ongoing automation. The degree of automation provides a unified lens connecting capital deepening, factor shares, and productivity in a single theory-consistent measure.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-and-what-observables-are-used-to-infer-the-degree-of-automation"&gt;Q1. What is the core identification strategy and what observables are used to infer the degree of automation?&lt;/h3&gt;
&lt;p&gt;The degree of automation is identified from the first-order conditions of the CES production function. Under the Pareto efficiency-distribution assumption, the CES structure implies a one-to-one mapping from the aggregate capital-labor ratio (in efficiency units), the share parameter s, and the elasticity of substitution rho to the degree of automation (Theorem 4, Eq. 25 and 31). In practice, the authors estimate the CES production function via GMM on first-differenced factor-share equations, recover rho and gK, and plug those into the formula for the degree of automation. No direct observation of tasks, robots (in the standard CES step), or technology-specific adoption decisions is required.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-threats-to-identification-and-how-do-the-authors-address-them"&gt;Q2. What are the main threats to identification and how do the authors address them?&lt;/h3&gt;
&lt;p&gt;The main threats are endogeneity of the output-to-labor and output-to-capital ratios (both simultaneously determined with factor prices) and measurement error in the capital-labor ratio (arising from industry classification changes and the RAS procedure used to construct robot data). The authors address endogeneity via GMM estimation using five distinct IV sets that include lagged factor prices, information stock and its price, trade openness, and workforce composition variables. They report that elasticity estimates are stable across all five IV sets and across alternative sample windows (including a longer 1973-2011 sample from pre-SNA-revision data), and conclude that measurement error is unlikely to drive the results. The overidentification test is not rejected for any IV set in the baseline CES specification (and for most in the two-level specification).&lt;/p&gt;
&lt;h3 id="q3-what-theoretical-result-connects-the-degree-of-automation-to-factor-income-shares"&gt;Q3. What theoretical result connects the degree of automation to factor income shares?&lt;/h3&gt;
&lt;p&gt;Corollary 1 establishes that under the Pareto efficiency structure (Eq. 22) with competitive factor markets, the degree of automation equals the capital income share when sigma &amp;gt; 1, and equals the labor income share when sigma &amp;lt; 1. This makes the degree of automation directly readable from income-share data in the theoretically preferred case (sigma &amp;gt; 1 for Japan). The empirical results are consistent with this: the average degree of automation across manufacturing industries is close to the average capital income share over the sample, providing a cross-check for Corollary 1.&lt;/p&gt;
&lt;h3 id="q4-why-does-the-paper-use-a-leontief-production-function-at-the-task-level-while-obtaining-a-ces-function-at-the-aggregate-level"&gt;Q4. Why does the paper use a Leontief production function at the task level while obtaining a CES function at the aggregate level?&lt;/h3&gt;
&lt;p&gt;The Leontief specification at the task level reflects the idea of a bottleneck in production: within a single narrowly-defined task, only capital or labor is used (once a task is automated, capital fully replaces labor in that task). Perfect substitutability between capital and labor operates at the extensive margin (which tasks are automated) rather than within a task. The aggregate (envelope) function, formed by varying the automation cutoff as the capital-labor ratio changes, generates any elasticity of substitution from zero to infinity. The Pareto efficiency-distribution assumption pins down the specific case of a CES aggregate.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-two-level-ces-extension-work-and-what-does-it-add"&gt;Q5. How does the two-level CES extension work, and what does it add?&lt;/h3&gt;
&lt;p&gt;The two-level CES nests industrial robots and other capital into a capital composite at the inner level (robots vs. other capital, with elasticity sigma-b), then combines that composite with labor at the outer level (composite vs. labor, with elasticity sigma-a). Robot data for 52 industries are constructed via the RAS and perpetual inventory methods with an initial year of 1985. Because robots account for only 0.44 percent of aggregate capital on average, they have a small direct weight, but the two-level decomposition isolates their specific contribution to the automation margin. The two-level CES estimates sigma-a between 1.191 and 1.346 (higher than the standard CES estimates), and finds that the test of equality between sigma-a and sigma-b is rejected for three of five IV sets, suggesting the two elasticities genuinely differ. The average degree of automation rises more steeply under the two-level estimate (0.398 to 0.430) than under the standard CES estimate (0.407 to 0.426), indicating that explicitly accounting for robots reveals a more pronounced automation trend.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-papers-internal-consistency-criterion-and-how-does-it-rank-alternative-automation-measures"&gt;Q6. What is the paper&amp;rsquo;s internal consistency criterion, and how does it rank alternative automation measures?&lt;/h3&gt;
&lt;p&gt;Internal consistency is defined as the mean squared gap between the degree of automation inferred from the capital-labor ratio (Eq. 37, the paper&amp;rsquo;s preferred measure) and the degree of automation implied by observed output per worker given the same CES structure (Eq. 41). A smaller gap means the measure is more coherent with the CES framework from which it is derived. The Pareto-based measure achieves a distance of 0.0000319, more than seventy times smaller than the Cobb-Douglas alternative (0.002484) and over three hundred times smaller than the continuity-preserving alternative (0.00999). The authors therefore select the Pareto-based measure as most internally consistent with CES production.&lt;/p&gt;
&lt;h3 id="q7-what-is-documented-about-heterogeneity-in-automation-across-industries"&gt;Q7. What is documented about heterogeneity in automation across industries?&lt;/h3&gt;
&lt;p&gt;The degree of automation varies substantially across the 52 manufacturing industries, with a standard deviation of 0.171 and a range from 0.138 to 0.811 in the standard CES estimation. The kernel density in 1994 has a fat left tail with a mode just above 0.3, and several industries already exceed 0.7. The distribution shifts rightward by 2020 but remains dispersed. The authors split industries into those with an increasing capital income share (34 industries) and those with a decreasing share (18 industries) and test whether the elasticity of substitution differs between groups; they find no statistically significant difference for any IV set, implying the CES structure is uniform across industries even though automation levels differ.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-connect-automation-to-tfp-and-the-productivity-paradox"&gt;Q8. How does the paper connect automation to TFP and the productivity paradox?&lt;/h3&gt;
&lt;p&gt;The theoretical framework shows that automation via task reallocation shifts the production function in a northeast direction in (k, y) space but does not shift it upward in a way that registers as TFP growth. Formally, increasing automation does not appear to impact TFP growth (citing Nakamura and Nakamura, 2008). The empirical finding that the degree of automation rose from 0.407 to 0.426 during Japan&amp;rsquo;s prolonged stagnation (1994-2020), a period of slow output-per-worker growth, is consistent with this: capital accumulation drove automation forward even though measured TFP growth was subdued. The paper thus links automation dynamics to Japan&amp;rsquo;s productivity paradox and implies that standard TFP accounting may understate the technological transformation underway.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-relationship-between-the-elasticity-of-substitution-and-the-direction-of-factor-share-changes-under-automation"&gt;Q9. What is the relationship between the elasticity of substitution and the direction of factor share changes under automation?&lt;/h3&gt;
&lt;p&gt;The CES framework implies that when sigma &amp;gt; 1 (capital and labor more substitutable), capital accumulation raises the capital income share and lowers the labor share; the degree of automation equals the capital income share. When sigma &amp;lt; 1, capital accumulation raises the wage-to-rental ratio by more, increasing the labor income share; the degree of automation equals the labor income share. In both cases automation rises with capital deepening. A key implication is that observing a stable or rising labor income share does not rule out rising automation when sigma is below one or close to one. The authors&amp;rsquo; estimate of sigma slightly above one for Japanese manufacturing implies a slightly rising capital share, consistent with the panel-estimated trend (b-hat = 0.00102, t-value = 6.84).&lt;/p&gt;
&lt;h3 id="q10-what-are-the-robustness-checks-and-how-stable-are-the-estimates"&gt;Q10. What are the robustness checks and how stable are the estimates?&lt;/h3&gt;
&lt;p&gt;Robustness checks include: (1) five distinct IV sets spanning different combinations of lagged wages, capital rental prices, information stock, trade openness, and workforce composition; (2) estimation under both neutral and capital-augmenting technical progress assumptions; (3) estimation using a longer sample (1973-2011 using pre-SNA-revision data), which yields a sigma still significantly above one and close to one, with slightly larger capital-augmenting technical progress reflecting higher growth in that period; (4) estimation of the full CES production function equation simultaneously with the two FOC equations (Appendix E.2), yielding similar elasticity estimates; (5) a structural change test splitting industries by capital-share trend, finding no significant difference in elasticity between subgroups. Unit root tests (Harris-Tzavalis and augmented Dickey-Fuller) confirm stationarity of all key variables except the part-time ratio, which also passes the ADF test.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-caveats-and-acknowledged-limitations"&gt;Q11. What are the caveats and acknowledged limitations?&lt;/h3&gt;
&lt;p&gt;The authors acknowledge several limitations. First, three conditions cannot be simultaneously satisfied: a CES aggregate, the degree of automation lying in the unit interval, and continuity of the automation measure at unit elasticity (sigma = 1). The preferred measure prioritizes the unit-interval restriction and sacrifices continuity at sigma = 1, making direct comparisons across the sigma &amp;lt; 1 and sigma &amp;gt; 1 cases problematic (an alternative continuous measure is derived in Appendix C but may fall outside the unit interval). Second, the framework abstracts from the creation of new tasks; changes in the total number of tasks over time would affect the automation measure. Third, the paper does not decompose automation by skill level; the observed differences between skilled and unskilled labor in automation suggest a need for nested CES structures in future work. Fourth, the two-level CES nesting (robots within capital composite) is dictated by data availability; alternative nestings, such as grouping robots and labor at the first level, are not separately identifiable.&lt;/p&gt;
&lt;h3 id="q12-how-does-this-paper-differ-from-and-improve-upon-the-prior-literature"&gt;Q12. How does this paper differ from and improve upon the prior literature?&lt;/h3&gt;
&lt;p&gt;The paper improves on micro-proxy approaches (robot counts, AI investment, task-exposure indices from Acemoglu-Restrepo 2020, Adachi 2025, etc.) by providing an aggregate, theory-consistent measure that does not require technology-specific data. It extends prior CES microfoundation work (Jones 2005 Pareto-Cobb-Douglas result, Growiec 2008 Weibull-CES results) by deriving the Pareto efficiency structure that yields CES specifically from task-level automation decisions. It improves on the authors&amp;rsquo; own prior work (Nakamura and Nakamura 2008, Nakamura 2009, 2010) by providing a complete theoretical justification for input efficiencies, a full treatment of the elasticity of substitution, and an empirical implementation. Relative to Artuc et al. (2023) and Adachi (2025), which use Frechet distributions for task productivity, this paper uses a deterministic framework with Pareto-distributed input efficiencies and emphasizes aggregate-level identification rather than cross-occupational substitution.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-policy-implications"&gt;Q13. What are the policy implications?&lt;/h3&gt;
&lt;p&gt;The paper does not make direct policy prescriptions, but its framework has several implications. First, policymakers tracking automation can use standard national accounts data (capital stock, labor input, output, factor shares) rather than waiting for technology-specific surveys, enabling faster and more comprehensive monitoring. Second, the result that automation can advance during periods of slow TFP growth suggests that technology policy focused solely on productivity metrics may underestimate the pace of labor displacement. Third, the finding that Japan&amp;rsquo;s capital accumulation drove automation even through prolonged stagnation implies that capital subsidies or policies encouraging investment could accelerate automation independent of TFP. Fourth, the model&amp;rsquo;s prediction that automation rises alongside increasing labor shares under low substitutability (sigma &amp;lt; 1) warns against complacency: labor-income gains and technology-driven labor displacement can coexist. Fifth, the need for future work on skill heterogeneity and task creation suggests that the framework can be extended to inform distributional policies.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Degree of automation&lt;/strong&gt;: In this paper, the share of the unit task continuum performed by capital rather than labor, denoted a_t, ranging from 0 to 1. It is determined endogenously in equilibrium by relative factor prices and increases with the capital-labor ratio. It is distinct from any technology-specific proxy and emerges as a function of aggregate macroeconomic observables.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Task-based production framework&lt;/strong&gt;: A model in which output requires completing a continuum of tasks, each exhibiting Leontief technology at the task level (capital and labor are perfectly substitutable within a task, but the firm either fully automates a task or uses labor exclusively). Tasks are ordered by the relative efficiency of capital to labor, and firms choose the automation cutoff that minimizes cost given factor prices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pareto efficiency distribution&lt;/strong&gt;: The specific parametric form of aggregate capital- and labor-input efficiency functions (Eq. 22) under which the task-level aggregation yields a CES production function at the macro level. The relationship between the degree of automation and aggregate input efficiencies follows a Pareto cumulative distribution, which also delivers the highest internal consistency among automation measures tested.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Internal consistency criterion&lt;/strong&gt;: A criterion for selecting among automation measures, defined as the mean squared gap between the automation degree inferred from the capital-labor relationship and the automation degree implied by the output-per-worker relationship, within the same CES structure (Eq. 42). A smaller gap indicates that the measure is more coherent with the CES production framework from which it is derived.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital-augmenting technical progress&lt;/strong&gt;: An exogenous shift in the efficiency of capital inputs (A_K,t) that raises the effective capital-labor ratio and therefore the degree of automation at any given physical capital-labor ratio. Distinguished from labor-augmenting and neutral technical progress. In the empirical estimation, capital-augmenting technical progress is statistically significant across all specifications, while labor-augmenting technical progress cannot be confirmed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Two-level CES production function&lt;/strong&gt;: An extension of the standard CES that nests industrial robots and other capital into a capital composite at the inner level (with substitution elasticity sigma-b), then combines the composite with labor at the outer level (with elasticity sigma-a). Allows separate identification of the automation role of robots versus other capital, yielding a more pronounced increase in the degree of automation than the standard CES when robots are explicitly accounted for.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Automation frontier&lt;/strong&gt;: The marginal task at which the cost of capital use exactly equals the cost of labor use, i.e., the task a_t at which lambda(a_t)/theta(a_t) = w_t/R_t. Tasks with indices below this frontier are automated; tasks above are performed by labor. As the wage-to-rental ratio rises, the frontier expands (more tasks become automated), capturing the central mechanism by which capital deepening drives automation.&lt;/p&gt;</description></item><item><title>Train to Opportunity: the Effect of Infrastructure on Intergenerational Mobility</title><link>https://macropaperwarehouse.com/papers/train-to-opportunity-the-effect-of-infrastructure-on-intergenerational-mobility/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/train-to-opportunity-the-effect-of-infrastructure-on-intergenerational-mobility/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether proximity to transport infrastructure can sever the occupational tie between parents and children — a question with direct bearing on the debate over place-based versus people-based policies. The authors exploit the nineteenth-century expansion of the railroad network across England and Wales, a setting where the First and Second Industrial Revolutions were remaking the occupational structure at the same time that the railroad was knitting together local labor markets and enabling geographic mobility.&lt;/p&gt;
&lt;p&gt;The empirical strategy centers on a novel dataset of close to 980,848 father-son pairs constructed from the full digitized population censuses of England and Wales in 1851, 1881, and 1911 (I-CeM project). Individuals are tracked across consecutive censuses using the Abramitzky-Mill-Perez (2019) linking procedure, which achieves match rates of 43–50% for men aged 40–52. Crucially, each individual is geolocated to the street level by matching census addresses to the GB1900 gazetteer, allowing railroad access to be measured as the straight-line distance from the childhood residence to the nearest train station — a finer measure than the district-level presence indicators used in prior work. Sons&amp;rsquo; occupations are observed at ages 40–52; fathers&amp;rsquo; occupations are measured 30 years earlier when sons were aged 10–22. Occupational mobility uses two complementary scales: HISCO categories (farming, laborer, services, sales, clerical, managerial, professional) and the continuous HISCAM social-interaction-distance ranking (scores 28–99, mean 50, SD 10).&lt;/p&gt;
&lt;p&gt;The key endogeneity problem is that railroad companies targeted low-density, cheap land, and that wealthy landowners and local politicians influenced station placement. To isolate exogenous variation, the authors construct a dynamic least-cost path (DLCP) network connecting 53 major towns identified by their 1801 populations (top 10% of the population distribution, threshold 9,172 inhabitants). The DLCP assigns slope costs to 50x50 meter grid cells and finds the minimum-cost path between every town pair. Lines are ranked by betweenness centrality to separate &amp;ldquo;early&amp;rdquo; 1851 lines from &amp;ldquo;late&amp;rdquo; 1881 lines, giving a time-varying instrument. Proximity to the nearest DLCP line is used as the instrument for proximity to the nearest actual train station, with standard errors clustered at the parish level. Controls include county and census-year fixed effects, distance to the nearest 1801 major town and its population, distance to Roman roads, ancient ports, and navigable waterways, plus household characteristics (number of servants as a wealth proxy, household size, and father&amp;rsquo;s foreign birth).&lt;/p&gt;
&lt;p&gt;Main results (preferred IV specification with full controls): sons who grew up one standard deviation — approximately 5 km, or about one hour&amp;rsquo;s walk — closer to a train station were 11 percentage points more likely to work in an occupation category different from their father&amp;rsquo;s. They were 5 percentage points more likely to be upwardly mobile, defined as a son&amp;rsquo;s HISCAM score exceeding his father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s distribution. The downward mobility estimate is 3 percentage points — positive but smaller in magnitude — indicating that railroad access raises occupational churn asymmetrically, predominantly upward. First-stage F-statistics exceed the Staiger-Stock threshold comfortably (135–414 across specifications). OLS estimates are uniformly smaller than IV estimates, consistent with historical evidence that the railroad targeted areas with weaker growth trajectories.&lt;/p&gt;
&lt;p&gt;The occupational transitions underlying these results run strongly out of farming and into professional, clerical, sales, and services categories, regardless of the father&amp;rsquo;s own occupation (Table IV). Sons growing up closer to the railroad were 19 percentage points less likely to work in a declining occupation and 16 percentage points more likely to work in a growing occupation. The distributional pattern shows an inverted-U relationship with father&amp;rsquo;s occupational decile for occupation-category switching and rank divergence, with the greatest gains concentrated among sons of middle-ranking fathers. For upward mobility specifically, the benefits diminish monotonically as father&amp;rsquo;s rank rises — sons from blue-collar backgrounds gained more (upward mobility coefficient 0.064) than sons from white-collar backgrounds (0.031).&lt;/p&gt;
&lt;p&gt;The authors decompose the total railroad effect on intergenerational mobility into three channels using a structural decomposition applied to a sample of 342,715 brothers: (1) changes in local labor-market opportunities, estimated as the effect on mobility for stayers; (2) changes in the returns to spatial mobility, estimated via a within-family comparison of brothers who moved versus stayed; and (3) changes in the rate of spatial mobility itself. Better railroad access raised the probability of moving away from the birth county by 15 percentage points. However, the estimated return to spatial mobility — the extra boost from actually moving — was reduced by railroad access (negative interaction between proximity and mover status), meaning the railroad decreased the relative advantage of leaving. The decomposition (Table C.6) shows that changes in local opportunities account for the great majority of the total mobility effect. Parish-level evidence confirms the local opportunity mechanism: better-connected parishes saw population growth, more industrial chimneys, more entrepreneurs, higher shares of skilled and literate workers, higher Gini coefficients, and higher median occupational ranks — consistent with agglomeration, industrialization, and skill-biased structural change.&lt;/p&gt;
&lt;p&gt;The policy implication is that transport infrastructure investment can reduce intergenerational persistence in occupational status, primarily by restructuring the local labor market rather than by enabling workers to exit. The caveat is that these gains were unevenly distributed — middle- and lower-ranking families benefited most, and the railroad simultaneously raised local inequality alongside local mobility.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-and-what-are-the-main-threats-it-addresses"&gt;Q1. What is the core identification strategy and what are the main threats it addresses?&lt;/h3&gt;
&lt;p&gt;The authors use a &amp;lsquo;dynamic least-cost path&amp;rsquo; (DLCP) instrument. They connect 53 major English and Welsh towns (defined as the top 10% of the 1801 population distribution, with at least 9,172 inhabitants) via least-cost routes computed over a 50×50 meter terrain grid that assigns slope-based costs to each cell. The instrument is proximity from the childhood residence to the nearest line in this DLCP network. The logic is that individuals incidentally located near the geographic route between major historical towns are more likely to be near an actual railroad — but the DLCP route is based purely on terrain costs, not on local demand, local resources, or the political lobbying that shaped where stations were actually placed. The strategy addresses: (a) reverse causality from high-growth areas attracting railroad placement; (b) sorting of ambitious or wealthy households toward connected parishes; (c) railroad companies&amp;rsquo; demand-driven routing choices. The exclusion restriction could be violated if location along least-cost paths between 1801 major towns is directly correlated with intergenerational mobility for reasons other than the railroad. The paper addresses this by controlling for distance to the nearest 1801 major town and its population (proximity to nodes), proximity to Roman roads, ancient ports, and navigable waterways (pre-existing trade routes), and household wealth proxies.&lt;/p&gt;
&lt;h3 id="q2-how-is-the-instrument-made-dynamic-and-why-does-this-matter"&gt;Q2. How is the instrument made dynamic, and why does this matter?&lt;/h3&gt;
&lt;p&gt;The authors divide the hypothetical network into &amp;rsquo;early&amp;rsquo; (1851) and &amp;rsquo;late&amp;rsquo; (1881) lines by ranking lines in decreasing order of betweenness centrality — the number of times a line connects major towns via shortest paths — until the total cost of the 1851 observed network is exhausted. This dynamic structure means the instrument varies across both space and census cohorts (sons measured in 1851-1881 versus 1881-1911). Without the dynamic feature, the instrument could conflate the effects of lines that were built early (and thus had decades to affect local economies) with lines built later. The temporal variation bolsters the plausibility of the exclusion restriction and is shown to be robust in alternative specifications using static least-cost paths and slope-free least-cost paths.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-four-dependent-variables-and-how-is-intergenerational-mobility-defined"&gt;Q3. What are the four dependent variables and how is intergenerational mobility defined?&lt;/h3&gt;
&lt;p&gt;The paper uses four measures: (1) an indicator equal to one if the son works in a different HISCO occupation category than his father; (2) the absolute value of the difference in HISCAM scores between son and father; (3) &amp;lsquo;upward mobility,&amp;rsquo; an indicator equal to one if the son&amp;rsquo;s HISCAM score exceeds his father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s score distribution; (4) &amp;lsquo;downward mobility,&amp;rsquo; the symmetric indicator for a decline greater than one standard deviation. Sons&amp;rsquo; occupations are observed when sons are 40–52 years old; fathers&amp;rsquo; occupations are measured 30 years earlier when sons were 10–22. The HISCAM scale is held constant over the period (national GB scale, 1800–1938) so that rankings reflect fixed social stratification positions rather than period-specific prestige. The paper also uses time-varying HISCAM, HISCLASS, Woollard, and Armstrong classifications as robustness checks.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-first-stage-performance-of-the-instrument"&gt;Q4. What is the first-stage performance of the instrument?&lt;/h3&gt;
&lt;p&gt;The first-stage relationship between proximity to the nearest DLCP line and proximity to the nearest actual train station is positive and statistically significant across all specifications. The Sanderson-Windmeijer F-statistic is 414 in the specification without controls, 136 with county and year fixed effects and full controls, and remains well above the conventional threshold of 10. The first-stage coefficient drops from 0.640 to 0.339 when full controls are added, indicating that a portion of the geographic correlation between the DLCP and the actual network reflects the pre-existing economic importance of towns and travel routes — which is precisely what the controls absorb.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q5. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The paper decomposes the total IV effect on intergenerational mobility using a three-part decomposition: (1) Changes in local opportunities, measured as the effect of proximity on mobility for sons who stayed in their birth county (stayers); (2) Changes in the returns to spatial mobility, estimated by comparing brothers who moved with brothers who stayed (using family fixed effects), and interacting this comparison with railroad proximity; (3) Changes in the rate of spatial mobility itself, estimated from the effect of proximity on the probability of county-to-county migration. Table C.6 shows that local opportunities account for the dominant share of the total effect. The railroad raised the migration probability by 15 percentage points (Table VI), so spatial mobility channels exist — but the railroad decreased the relative advantage of actually moving (negative interaction term in Table V), meaning the local opportunity channel more than offsets the spatial channel. Supporting evidence from parish-level regressions (Table VII) shows that better-connected parishes experienced significantly higher population growth, more industrial chimneys, more entrepreneurs per 100 square meters, higher shares of skilled and literate workers, higher Gini coefficients, and higher median occupational ranks — consistent with agglomeration and skill-biased industrialization.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-is-documented-by-fathers-occupation-and-position-in-the-distribution"&gt;Q6. What heterogeneity is documented by father&amp;rsquo;s occupation and position in the distribution?&lt;/h3&gt;
&lt;p&gt;The effects are heterogeneous by the father&amp;rsquo;s occupational position. Figure 6 shows an inverted-U pattern for occupation-category switching and absolute rank divergence: sons of middle-ranking fathers benefit most from railroad access. For upward mobility (Figure 6c), the benefits diminish monotonically from the lower end of the father&amp;rsquo;s distribution — sons of low-ranking fathers are most likely to move up. Sons of white-collar fathers see smaller (and sometimes statistically insignificant) upward mobility gains (0.031) compared with sons of blue-collar fathers (0.064), while the occupation-category switching benefit is also larger for blue-collar sons (0.108 vs. 0.057) (Table C.1). Separate transition matrices by HISCO category (Table IV) show that railroad access reduces the probability of farming for sons of all father types, and raises probabilities of clerical, sales, and services occupations. Effects on becoming a laborer are heterogeneous: for sons of farmers, proximity raises the probability of becoming a laborer; for sons in service occupations, it decreases it.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-run"&gt;Q7. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper performs an extensive battery. (1) Alternative connectivity measures: distance to the nearest railroad line, indicator variables for train station within 5, 10, and 15 km, and parish-level station presence. (2) Alternative mobility thresholds: 0.5, 1.5, and 2 standard deviations for upward and downward mobility; time-varying HISCAM to account for changing occupational prestige. (3) Removing railroad-specific occupations (train conductors, controllers) to check for mechanical effects. (4) Alternative specifications: second-order polynomials, parish fixed effects (10,419 parishes), and fully nonparametric covariate controls via k-means clustering (500 clusters). (5) Alternative instruments: a slope-free DLCP and a static (non-dynamic) least-cost path. (6) Geolocation robustness: using parish centroids instead of street-level addresses. (7) Linking bias: controlling for the individual probability of being linked using cubic polynomials on linkage probability and surname-frequency dummies; also checking that the railroad network explains little of the share of linked individuals at the parish level. (8) Subsamples: by census year (1851-1881 vs. 1881-1911), by county (leave-one-out), by rural/urban status, by father&amp;rsquo;s age, by son&amp;rsquo;s age, by birth order, by native/first-/second-generation immigrant status, by whether the son was born in the same county he grew up in, and by whether the father was in farming. (9) Causal response weighting: the Loken-Mogstad-Wiswall decomposition shows positive IV weights across the entire proximity distribution, consistent with a LATE interpretation. Results are stable across all checks.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-handle-the-selection-into-migration-problem-in-estimating-returns-to-spatial-mobility"&gt;Q8. How does the paper handle the selection-into-migration problem in estimating returns to spatial mobility?&lt;/h3&gt;
&lt;p&gt;The authors follow Abramitzky, Boustan, and Eriksson (2012) and use a within-family comparison of brothers — a subsample of 342,715 sons from 157,369 households who grew up in the same household but one moved county while the other stayed. Family fixed effects absorb the shared household characteristics (wealth, motivation, family networks, financial constraints) that jointly determine the propensity to migrate and the baseline mobility trajectory. The railroad-proximity interaction with mover status is instrumented using the interaction of the DLCP instrument with the mover indicator, via a control function approach. The estimated baseline return to spatial mobility (the mover premium) is positive and significant — movers have higher occupation-category divergence and shift more in both directions — but the railroad-induced change in return to mobility is negative, meaning that proximity to the railroad reduced the additional mobility benefit of actually migrating. This finding is the core of the conclusion that local opportunities, not spatial mobility, dominate.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-paper-document-about-local-labor-market-changes-induced-by-the-railroad"&gt;Q9. What does the paper document about local labor market changes induced by the railroad?&lt;/h3&gt;
&lt;p&gt;Parish-level IV regressions (Table VII) show that better proximity to the 1851 network (instrumented by the DLCP) is associated with: significantly higher population growth between 1851 and 1881; a significantly larger number of industrial chimneys (proxying factory concentration, sourced from Heblich-Trew-Zylberberg (2021)); more entrepreneurs per 100 square meters (from the British Business Census of Entrepreneurs); higher shares of high-skilled and literate workers; a higher Gini coefficient over occupational ranks; and a higher median occupational rank. Additionally, sons in better-connected parishes were 19 percentage points less likely to work in a declining occupation and 16 percentage points more likely to work in a growing occupation (Table C.3). Sons were also 3 percentage points more likely to be literate and 7 percentage points more likely to work in a non-manual occupation (Table C.5). These findings collectively point to agglomeration, industrialization, skill-biased technological change, and the creation of a new entrepreneur class as the mechanisms by which the railroad transformed local labor market structure.&lt;/p&gt;
&lt;h3 id="q10-what-prior-work-does-this-paper-relate-to-most-closely-and-what-distinguishes-it"&gt;Q10. What prior work does this paper relate to most closely, and what distinguishes it?&lt;/h3&gt;
&lt;p&gt;The paper sits at the intersection of the railroad-infrastructure and intergenerational-mobility literatures. In the infrastructure tradition, it relates closely to Donaldson (2018, AER) on railroads in India, Donaldson and Hornbeck (2016, QJE) on US market access, Bogart et al. (2022, JUE) on population and structural change in England and Wales, and Heblich-Redding-Sturm (2020, QJE) on London commuting and urban growth. The closest prior paper is Perez (2017) on nineteenth-century Argentina, who finds railroad access shifted children from agricultural into white-collar and skilled blue-collar occupations; this paper provides similar evidence for England and Wales at individual level and adds a full mechanism decomposition. In the intergenerational mobility tradition it relates to Long and Ferrie (2013, AER) and Long (2013, ERH) on census-based occupational mobility in Victorian Britain. The key methodological advantages of the current paper are: (a) use of the full (not 2%) census for all three years, yielding close to 1 million father-son pairs with match rates of 43–50% versus 15–33% in prior work; (b) street-level geolocation enabling individual-level rather than district-level measurement of railroad access; (c) the explicit three-way mechanism decomposition separating local opportunities, returns to migration, and migration rates; and (d) documenting rich heterogeneity by father&amp;rsquo;s occupational rank and occupation category.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-what-scope-conditions-limit-their-external-validity"&gt;Q11. What are the policy implications and what scope conditions limit their external validity?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s core policy message is that transport infrastructure investment can be an effective mechanism for reducing intergenerational occupational persistence — primarily by creating new local labor market opportunities rather than by enabling low-income workers to reach distant job centers. This provides historical support for place-based policies of the sort embodied in the Biden &amp;lsquo;Build Back Better&amp;rsquo; infrastructure proposals or the UK HS2 high-speed railway project (mentioned in the paper). The main scope conditions limiting generalizability are: (1) The setting is nineteenth-century England and Wales during the Industrial Revolution, when the occupational structure was shifting rapidly from farming to industry and commerce — the railroads arrived at a moment of latent demand for new labor market structures; (2) The benefits were not evenly distributed: middle-ranking families (by father&amp;rsquo;s occupational rank) gained most in absolute occupational switching and rank divergence, while the lowest-ranked families gained most specifically in upward mobility; (3) The railroad simultaneously raised local inequality alongside local mobility, suggesting infrastructure investment can be inequality-increasing in the cross-sectional distribution of wages even as it reduces intergenerational persistence; (4) The effects are highly localized — even 5 km of additional distance matters — implying that the placement of stations relative to where low-income families actually live is crucial for achieving distributional goals.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-paper-document-about-the-baseline-patterns-of-intergenerational-mobility-in-the-sample"&gt;Q12. What does the paper document about the baseline patterns of intergenerational mobility in the sample?&lt;/h3&gt;
&lt;p&gt;In the full sample of 980,848 father-son pairs covering 1851-1881 and 1881-1911, 80% of sons do not remain in the same HISCO occupation category as their father. The correlation between father&amp;rsquo;s and son&amp;rsquo;s HISCAM ranks is 0.28. Among sons, 18% experienced upward mobility (son&amp;rsquo;s HISCAM rank more than one SD higher than father&amp;rsquo;s) and 15% experienced downward mobility (more than one SD lower). About 31% of sons moved to a different county from where they grew up, settling on average 100 km away. Sons grew up on average 3.28 km from the nearest train station (SD 5.45 km). These descriptives reveal strong spatial clustering in intergenerational mobility patterns at the parish level.&lt;/p&gt;
&lt;h3 id="q13-does-the-late-interpretation-hold-and-what-does-the-weighting-function-show"&gt;Q13. Does the LATE interpretation hold and what does the weighting function show?&lt;/h3&gt;
&lt;p&gt;The authors verify the LATE interpretation via two approaches. First, following Loken-Mogstad-Wiswall (2012), they compute the causal response weighting function as the covariance between each discrete proximity indicator and the DLCP instrument, divided by the covariance between the proximity measure and the DLCP instrument. They find positive weights across the entire distribution of proximity to the nearest train station, concentrated most heavily for individuals residing 0.5 to 1.5 proximity units (approximately 2.7 to 8.1 km) from a train station — these are the individuals whose proximity is most affected by incidental location along the DLCP. The absence of negative weights indicates the IV estimate does not mix complier and never/always-taker effects in a sign-reversing way. Second, following Blandhol et al. (2022), a fully nonparametric specification using 500 k-means clusters for covariates yields estimates very close to the parametric baseline, consistent with a LATE interpretation of the linear IV estimator.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Dynamic Least-Cost Path (DLCP) Network&lt;/strong&gt;: The paper&amp;rsquo;s instrument for railroad access. A hypothetical railroad network connecting England and Wales&amp;rsquo;s 53 largest towns in 1801 via routes that minimize geographic cost (distance plus slope-based terrain costs), ignoring all demand-side factors. Lines are classified as &amp;rsquo;early&amp;rsquo; (1851) or &amp;rsquo;late&amp;rsquo; (1881) by betweenness centrality until the cost budget of the actual 1851 network is exhausted. Proximity from childhood residence to the nearest DLCP line instruments proximity to the nearest actual train station.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intergenerational Occupational Mobility&lt;/strong&gt;: In this paper, the degree to which a son&amp;rsquo;s adult occupation differs from his father&amp;rsquo;s, measured both categorically (same versus different HISCO category) and cardinally (difference in HISCAM scores). Upward (downward) mobility is specifically defined as the son&amp;rsquo;s HISCAM score exceeding (falling below) the father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s HISCAM distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HISCAM Score&lt;/strong&gt;: A continuous occupational ranking (range 28–99, mean 50, SD 10) derived from the frequency of social interactions — marriages, friendships, parent-child links — between occupations in historical data. Higher scores indicate a more advantageous position in the social stratification structure. The paper uses the national Great Britain scale, held constant for 1800–1938, to make rankings comparable across census years.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Local Opportunities Channel&lt;/strong&gt;: The mechanism by which railroad access improved intergenerational mobility through restructuring the local labor market — enabling commuting, attracting factories and entrepreneurs, spurring urbanization and industrialization, and creating new occupations requiring new skills — without requiring sons to migrate away from their birth county. Identified empirically as the effect of railroad proximity on mobility outcomes for sons who stayed in their birth county (stayers).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Returns to Spatial Mobility&lt;/strong&gt;: The additional intergenerational mobility benefit (or penalty) associated with actually migrating to another county, estimated using within-family variation among brothers — one who moved and one who stayed — to net out shared household-level determinants of mobility. The paper finds that railroad access reduced (made more negative) the returns to spatial mobility, meaning that the relative advantage of leaving shrank as local opportunities expanded.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inconsequential Place IV Approach&lt;/strong&gt;: An identification strategy (following Chandra-Thompson 2000 and Michaels 2008) in which the instrument for infrastructure access is constructed from the geographic convenience of locations lying between endpoints of a planned network, rather than from demand-side factors at those locations. The DLCP instrument in this paper is a specific implementation: individuals living between 1801 major towns incidentally receive railroad access because the low-cost route between towns passes near their residence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Occupational Tie (Father-Son)&lt;/strong&gt;: The tendency for sons to remain in the same occupation category or same position in the occupational ranking as their father. In this paper, severing the occupational tie means a son moves to a different HISCO category and/or achieves a HISCAM score meaningfully different from his father&amp;rsquo;s. The railroad&amp;rsquo;s main effect is framed as reducing this tie, with upward mobility being the dominant direction of change.&lt;/p&gt;</description></item><item><title>Uncertainty and Change: Survey Evidence of Firms' Subjective Beliefs</title><link>https://macropaperwarehouse.com/papers/uncertainty-and-change-survey-evidence-of-firms-subjective-beliefs/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/uncertainty-and-change-survey-evidence-of-firms-subjective-beliefs/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: A large literature shows that firms perceiving more uncertainty make more cautious intertemporal decisions (investment, hiring, price setting), but it is far less clear what makes firms uncertain in the first place. Macro models typically impose rational expectations and treat uncertainty as exogenous shocks to the conditional volatility of fundamentals. The paper asks how subjective uncertainty arises and evolves, and whether it is the same object as conditional volatility.&lt;/p&gt;
&lt;p&gt;Data and design: The authors build a new panel from a quantitative module they added in 2012 to the ifo Business Survey of German manufacturing firms. At the start of each quarter, top managers report (i) last quarter&amp;rsquo;s realized sales (&amp;ldquo;Umsatz&amp;rdquo;) growth, (ii) a one-quarter-ahead point forecast, and (iii) best- and worst-case scenarios. The &amp;ldquo;span&amp;rdquo; between best and worst case is their quantitative measure of subjective uncertainty; the forecast error is realized growth minus the point forecast. The baseline sample is 1,005 firms and 8,889 firm-quarter observations over 27 waves, 2013:Q2–2019:Q4 — a calm period with no German recession. A simple scenario-analysis model (Proposition 1) shows that under a quadratic loss and a location-scale shock family, span is proportional to subjective standard deviation, justifying span as an index of subjective conditional volatility. An organizing framework contrasts rational expectations (Example R: subjective uncertainty equals conditional volatility, forecasts unbiased) with learning about signal quality (Example L: managers are unsure of signal precision, so unfamiliar signals raise perceived uncertainty even when true volatility is constant, and generate forecast bias).&lt;/p&gt;
&lt;p&gt;Main findings with magnitudes: (1) Subjective uncertainty reflects experienced change, in both cross section and time series, following an asymmetric V-shape in growth (steeper negative branch, flatter positive branch, minimum near zero). Mean span is 12.4 pp, larger than mean absolute forecast error of 9.0 pp; cross-firm SD of time-averaged span is 7.4 pp and within-firm time-series SD of span is 6.3 pp. Cross-sectional V: a 1 pp lower (more negative) average growth goes with about 0.6 pp higher span; a 1 pp higher positive average growth with about 0.2 pp higher span. Time-series V (firm fixed effects removed): a 1 pp lower negative quarterly growth is followed by 0.2 pp higher span next quarter; a 1 pp higher positive growth by 0.1 pp (0.118 positive, -0.204 negative branch coefficients in Table 4). (2) Uncertainty is more than conditional volatility. Volatility explains about a quarter of cross-sectional variation in uncertainty; turbulence quartile dummies alone explain 30%, with span rising from 7 pp (lowest) to 18 pp (highest quartile). But controlling for turbulence, shrinking firms remain more uncertain (bottom-trend dummy ~2 pp) and make systematically too-conservative (toward-zero) forecasts, while large firms (&amp;gt;250 employees) report ~5 pp lower span holding trend/turbulence fixed (9 pp unconditionally). In the time series, after positive growth uncertainty rises but absolute forecast errors do not — inconsistent with rational expectations (Proposition R2), consistent with learning (Example L). Within-firm forecast-error/forecast correlation is -0.27 (overreaction); larger in magnitude (-0.31 vs -0.24) for low-excess-span firms. (3) Uncertainty is mostly idiosyncratic (time/industry fixed effects give R-squared ~1%, rising to ~5-7% with time-industry effects) yet matters for plans: a one-SD rise in span raises the probability of planned employment decrease by 2.4 pp (vs 4.2 pp for a one-SD forecast decline; baseline ~11%), raises planned price decreases by 0.9 pp and lowers planned price increases by 0.8 pp. Because employment (a quantity) and prices move the same direction, uncertainty acts like a negative demand shifter / &amp;ldquo;pessimism,&amp;rdquo; not a freezer of actions.&lt;/p&gt;
&lt;p&gt;Implications: Understanding subjective uncertainty requires going beyond rational-expectations models where uncertainty equals conditional volatility; learning is a promising alternative even for mature firms (median age 45 years). Decoupling of uncertainty from volatility matters for welfare and policy evaluation (misallocation, optimal policy under idiosyncratic risk).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-measurement-strategy-and-why-is-span-a-valid-index-of-subjective-uncertainty"&gt;Q1. What is the core measurement strategy, and why is span a valid index of subjective uncertainty?&lt;/h3&gt;
&lt;p&gt;The ifo module elicits best- and worst-case sales-growth scenarios; span (best minus worst) is the uncertainty measure, and the separate point forecast (answer 2b) is the subjective conditional mean. The authors model managers who think through a finite number n of scenarios to minimize expected quadratic loss based on distance from the closest scenario. Proposition 1 shows that if growth g = mu + sigma*epsilon belongs to a location-scale family, optimal span is linear in sigma (independent of mu), so span is proportional to subjective conditional standard deviation. Quadratic cost is a second-order approximation to general loss, making the link broad. Span is also robust/low-cognitive-load: it depends only on adjacent scenarios&amp;rsquo; first-order conditions, so it is insensitive to interior reshaping or tail-shape changes managers cannot confidently distinguish.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identification-strategy-for-distinguishing-uncertainty-from-conditional-volatility-and-what-are-the-threats"&gt;Q2. What is the identification strategy for distinguishing uncertainty from conditional volatility, and what are the threats?&lt;/h3&gt;
&lt;p&gt;Identification rests on contrasting two observable implications. Under rational expectations (Example R), a cross-sectional uncertainty V must be accompanied by a cross-sectional volatility V in mean absolute forecast errors (Proposition R1), and a time-series uncertainty V must coincide with a &amp;lsquo;conditional-volatility V&amp;rsquo; in absolute forecast errors (Proposition R2). Under learning (Example L), uncertainty can move with growth while debiased forecast-error volatility does not (Proposition L2). The authors test these by comparing span responses to forecast-error responses. The main threat is that span is only an index of subjective volatility (level not identified), so for the negative branch — where both uncertainty and volatility rise — they cannot fully rule out that higher uncertainty merely reflects higher conditional volatility. They argue against this because the implied span-to-volatility ratio (up to 4 in Table 4) would far exceed the roughly one-for-one cross-sectional relationship for most firms. For positive growth, the absence of any forecast-error response makes the rational-expectations explanation clean to reject.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-two-competing-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q3. What are the two competing mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;Mechanism 1 (Example R, rational expectations): subjective uncertainty equals true conditional volatility, driven by heteroskedastic fundamentals; forecasts are unbiased. Mechanism 2 (Example L, learning about signal precision): growth is homoskedastic but managers observe a noisy signal of unknown information content gamma; using a Normal-Gamma prior with confidence parameter nu, an unfamiliar signal (far from prior mean, either sign) leads managers to infer lower precision and remain more uncertain, and generates forecast bias toward zero. Distinguishing tests: (a) cross section — shrinking firms are more uncertain AND biased holding volatility fixed (supports learning, Proposition L1b); large firms are less uncertain but unbiased (supports a confidence/nu channel, L1c); (b) time series — after positive growth, uncertainty rises but absolute forecast errors do not (rejects R2, supports L2); (c) the within-firm negative correlation between forecast and forecast error (-0.27) indicates overreaction from overprecision (Proposition L3). The preferred reading is a hybrid: a known volatility component generating the negative branch (R) plus a symmetric learning V (L).&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented-across-firms"&gt;Q4. What heterogeneity is documented across firms?&lt;/h3&gt;
&lt;p&gt;Three dimensions. Turbulence (time-series SD of growth): strongly raises uncertainty — top vs bottom quartile span 18 vs 7 pp, ~1.5 cross-sectional SDs, dummies explain 30%. Trend growth: asymmetric V — both fast-growing and fast-shrinking firms are more uncertain, but after controlling for turbulence only the bottom (shrinking) trend quartile retains a significant ~2 pp effect, and shrinking firms also have biased (too-conservative) forecasts, whereas fast-growing firms lose significance once volatility is controlled. Size: larger firms perceive less uncertainty — large (&amp;gt;250 employees) firms ~9 pp lower span unconditionally, ~5 pp lower controlling for trend and turbulence, but show no significant difference in average forecast errors (so the size effect is a confidence/nu channel, not bias). Time-series heteroskedasticity of span also rises with turbulence and trend and is larger for smaller firms, consistent with smaller firms having lower nu. Employment effects of uncertainty are similar across size classes (if anything slightly stronger for large firms).&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Industry dummies (14 sectors) added to the cross-sectional span regression leave the turbulence/trend/size coefficients essentially unchanged and raise R-squared by only 2 pp, showing the effects are within-industry. Time and time-industry fixed effects confirm variation is overwhelmingly idiosyncratic (R-squared ~1% rising to ~5-7%). The within-firm uncertainty results are robust to requiring at least 5 span observations per firm (Table I4), as are the employment/price-plan results (Tables I6). Deseasonalization is corroborated at macro and micro level (Appendix B). Forecast-error analyses use a debiased absolute forecast error (residual from regressing forecast error on past growth and firm fixed effects) to separate volatility from bias, and a &amp;lsquo;statistical forecast error&amp;rsquo; (deviation of growth from firm mean) as an econometrician benchmark, both giving the same V/no-V patterns. Data quality is documented: ~73-86% of respondents are top management, the responder is the same person in ~98% of firms, ~80% of firms use in-house quantitative planning, and a majority rely on scenario analysis.&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q6. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It builds on survey-based &amp;lsquo;micro uncertainty&amp;rsquo; work (Guiso and Parigi 1999; Bontempi et al. 2010; Bachmann, Elstner and Sims 2013). Several papers found V-shapes between subjective uncertainty and lagged sales growth (Altig et al. 2022 Atlanta Fed SBU; Bloom et al. 2020 MOPS; Kumar, Gorodnichenko and Coibion 2023 New Zealand), but those use single cross sections or short pooled samples and cannot separate cross-sectional from time-series Vs. The contribution is decomposing the V into between- and within-firm components and constructing volatility Vs to contrast against the uncertainty Vs, showing uncertainty is more than volatility. It also connects to the behavioral/miscalibration literature (Ben-David, Graham and Harvey 2013; Barrero 2022) by linking forecast bias to the gap between subjective uncertainty and conditional volatility via endogenous perceived precision. Uniquely, it studies subjective idiosyncratic uncertainty jointly with both a quantity (employment) and prices in normal (non-recession) times; Kumar et al. (2023) found &amp;lsquo;uncertainty as pessimism&amp;rsquo; but for a macro variable (GDP).&lt;/p&gt;
&lt;h3 id="q7-what-are-the-policy-and-modeling-implications-and-their-scope-conditions"&gt;Q7. What are the policy and modeling implications, and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The decoupling of uncertainty from volatility matters for welfare and policy because the standard approach (regress absolute forecast errors on conditioning information and use the fitted value as uncertainty) measures &amp;rsquo;too little&amp;rsquo; uncertainty — it ignores uncertainty about features the econometrician sees only with hindsight. Heterogeneous-firm models of misallocation and optimal policy under idiosyncratic risk (e.g., Boar et al. 2025; Di Tella et al. 2025) should incorporate uncertainty distinct from volatility. Models of firm dynamics need either heteroskedastic innovations or sufficient nonlinearity, plus feedback from past growth to uncertainty (learning), and should treat idiosyncratic demand uncertainty as a driver of employment churn and price dispersion even in steady state. Scope conditions: the evidence is German manufacturing, 2013-2019, a calm idiosyncratic-shock-dominated period (so results speak to idiosyncratic, not aggregate, uncertainty); span identifies relative not absolute uncertainty; for idiosyncratic uncertainty to affect actions, firm decisions must depend on it (manager career concerns, closely-held ownership, or ambiguity/Knightian uncertainty defeating diversification). The authors note the decoupling principle extends to policy uncertainty (e.g., tariffs) even when realized paths are not volatile.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-uncertainty-as-a-negative-demand-shifter-result-tell-us-about-the-type-of-shocks-managers-fear"&gt;Q8. What does the &amp;lsquo;uncertainty as a negative demand shifter&amp;rsquo; result tell us about the type of shocks managers fear?&lt;/h3&gt;
&lt;p&gt;Because higher span lowers BOTH planned employment (a quantity) and planned prices in the same direction, the comovement indicates that managers primarily worry about demand shortfalls rather than cost shocks. A firm fearing a demand shortfall scales down production (sheds workers) and lowers prices; a firm fearing input-cost increases would still cut employment but RAISE prices. The observed pattern therefore points to idiosyncratic, subjective demand uncertainty as the relevant primitive, and (with financial frictions or risk/ambiguity-averse decision-makers placing more weight on low-payoff states) explains why uncertainty &amp;lsquo;acts like pessimism&amp;rsquo; rather than freezing actions.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-key-caveats-and-limitations"&gt;Q9. What are the key caveats and limitations?&lt;/h3&gt;
&lt;p&gt;Span is an index of subjective volatility, so levels and the exact span-to-volatility ratio are not point-identified, leaving residual ambiguity on the negative branch where uncertainty and volatility both rise. The sample is non-recessionary German manufacturing, so results characterize idiosyncratic (not aggregate) uncertainty; the authors explicitly note variation is essentially all idiosyncratic. The learning examples abstract from explicit dynamics (the prior is held fixed each period), serving as stark illustrations rather than a fully dynamic structural model; the data are interpreted through a hybrid of R and L. The plan outcomes are qualitative (up/down/same) and ifo does not elicit realized outcomes suitable for the authors&amp;rsquo; purposes, so the link to realized employment/prices relies on external evidence that ifo indicators forecast those variables.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Understanding High-Wage Firms: Monopoly, Monopsony, and Bargaining Power</title><link>https://macropaperwarehouse.com/papers/understanding-high-wage-firms-monopoly-monopsony-and-bargaining-power/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/understanding-high-wage-firms-monopoly-monopsony-and-bargaining-power/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Why do some firms pay persistently higher wages for observably similar workers, and what role do firms&amp;rsquo; product-market power (monopoly/markups), labor-market power (monopsony/markdowns), and workers&amp;rsquo; collective bargaining power play in shaping wages and welfare? Prior literature studies labor-market power as a driver of wages/profits but abstracts from product-market power and bargaining, while the markups literature abstracts from imperfect labor competition and bargaining. The paper unifies all three in one structural framework.&lt;/p&gt;
&lt;p&gt;Central theoretical insight: A firm&amp;rsquo;s wage equals its marginal revenue product of labor (MRPL) times a &amp;ldquo;labor wedge&amp;rdquo; (the share of MRPL workers receive). The labor wedge decomposes into three components — price-cost markups, monopsony markdowns, and bargaining power — via equation (3): Lambda = kappa*(product market rents term) + (1-kappa)*lambda. With positive bargaining power (kappa&amp;gt;0) workers capture a share of markup-generated rents, so the labor wedge rises with markups (rent-sharing); this nests pure monopsony as the kappa=0 special case.&lt;/p&gt;
&lt;p&gt;Data and setting: French administrative micro-data. Firm balance sheets (FARE, 2008-2019, DGFiP); firm-product output prices (EAP survey, 2009-2019, INSEE, manufacturing firms &amp;gt;=20 employees or sales &amp;gt;5m euros); matched employer-employee data (DADS, 1995-2018) which crucially includes hours worked. Firm wage premia estimated via a k-means/BLM grouped AKM regression (Bonhomme, Lamadon, Manresa 2019). Markups and labor wedges estimated with the production-function/production approach (De Loecker-Warzynski 2012; Yeh et al. 2022) using translog functions and an Ackerberg-Frazer-Caves control function, separating the two by noting markups distort all input demands while labor wedges distort only labor demand.&lt;/p&gt;
&lt;p&gt;Two key empirical facts a standard monopsony model cannot explain: (i) high-wage firms charge higher output prices and markups; (ii) high-wage firms pay a larger share of MRPL as wages (higher labor wedges). Both persist within narrow industries and conditional on TFP, pointing to product quality and positive bargaining power.&lt;/p&gt;
&lt;p&gt;Main quantitative findings (French manufacturing, 2016 unless noted): Median markup 1.32 (IQR 1.14-1.60). Median labor wedge 0.62 (median monopsony markdown 0.46) — the gap is due to bargaining power and markups. Workers capture about 12% of firm profits (bargaining power kappa ~ 0.12-0.14; falls to ~0.05-0.13 under IV correction). Median markdown 0.46 implies a median firm-specific labor supply elasticity of 0.85. Accounting for hours matters: median labor wedge is 0.62 with effective hours, 0.65/0.68/0.71 across specifications, rising to 0.71 when labor is measured by employment (near Yeh et al.&amp;rsquo;s 0.70-0.73 US figures) — so omitting hours upward-biases labor wedges.&lt;/p&gt;
&lt;p&gt;Quantitative GE model (oligopoly/oligopsony, nested-CES, Atkeson-Burstein/Berger et al.): A 1% productivity shock has wage passthrough 0.97-0.99 versus 0.23 for an equal quality shock (because varieties are close substitutes, sigma=5.17), though quality still generates more wage-premium dispersion. Markups and markdowns reduce welfare by 46% in consumption-equivalent terms, with markups alone accounting for over 80%; misallocation explains about 63% of the markup welfare cost. Equalizing markups raises average wages 39% and wage variance 99% and welfare 24% (output-restriction effect dominates rent-sharing, so equalizing markups raises wage dispersion). Raising bargaining power from 0.12 to 0.50 matches the wage gains of removing markups but yields only 10% welfare gain (vs 38%); full bargaining power (kappa=1) raises welfare 13%, under one-third of the planner&amp;rsquo;s 46% gain. Bargaining power offsets the uniform-tax and misallocation distortions on labor demand but cannot fix markup distortions to capital/material demand.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-for-separating-markups-from-labor-wedges-and-what-are-its-main-assumptionsthreats"&gt;Q1. What is the core identification strategy for separating markups from labor wedges, and what are its main assumptions/threats?&lt;/h3&gt;
&lt;p&gt;The author applies the production approach: estimate translog production functions per 2-digit manufacturing sector (via two-step GMM with an Ackerberg-Frazer-Caves control function for unobserved productivity) to recover firm-specific output elasticities. Markups distort the demand for ALL inputs while labor wedges distort ONLY labor demand, so choosing materials as a flexible, price-taken input lets markups be identified from the material cost share (mu = alpha_m * PY/(Pm*M)) and labor wedges from the wage-bill-to-materials ratio scaled by elasticity ratios (eq. 4). Key assumptions/threats: materials must be a flexible input firms take prices for (examined in Appendix B.7-B.8); unobserved productivity must satisfy scalar unobservability and monotonicity in material demand; unobserved output and input prices bias elasticities — addressed using observed EAP output prices (measuring output in quantities) plus the De Loecker et al. (2016) input-price control function, and additionally controlling for firm wage premia because monopsony markdowns create unobserved labor-price variation. Markup variation driven by idiosyncratic demand uncorrelated with TFP is controlled via export status, market shares, firm age, and a 3rd-order price polynomial. Gandhi-Navarro-Rivers concerns about identifying material elasticities are addressed in Appendix B.9.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-new-identification-challenge-for-estimating-bargaining-power-and-how-is-it-solved"&gt;Q2. What is the new identification challenge for estimating bargaining power, and how is it solved?&lt;/h3&gt;
&lt;p&gt;The rent-sharing literature estimates bargaining power kappa by regressing wages on quasi-rents using instruments (export demand, patent shocks) assumed orthogonal to the worker&amp;rsquo;s reservation wage. But in this model, when kappa=0 workers earn an endogenous monopsony wage (lambda*MRPL) that moves with the SAME firm-specific shocks (productivity, quality, amenities) that shift quasi-rents — so standard instruments violate the exclusion restriction. The solution: instead of the wage equation, exploit the labor-wedge equation (3), which relates labor wedges to markups and avoids unobserved monopsony wages. Conditional on markdowns, variation in product-market rents identifies kappa (when kappa=0 product-market rents do not affect the labor wedge). This shifts the core challenge from unobserved monopsony wages to unobserved amenities (mirroring IC3 in the rent-sharing literature), handled by a theory-consistent control function in which employment and the wage bill jointly proxy for amenities under a monotonicity assumption (labor supply increasing in amenities). Under multiplicative separability of wages and amenities, markdowns do not depend directly on amenities, so unobserved amenities do not bias kappa at all.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-bargaining-power-estimates-across-specifications"&gt;Q3. What are the bargaining-power estimates across specifications?&lt;/h3&gt;
&lt;p&gt;Pooled OLS gives ~0.135; adding firm fixed effects ~0.124; adding the amenity control function (columns 3-4) ~0.124-0.135, indicating amenities have little direct effect on markdowns; instrumenting product-market rents with their lags to correct correlated measurement error (columns 5-6) gives 0.130 and 0.059. Baseline kappa is taken as ~0.12 (specification 4). All 2-digit sectors have kappa below 0.3. These align with the rent-sharing literature&amp;rsquo;s typical 0.05-0.15, though external innovation-based instruments tend to find ~0.30.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-paper-measure-firm-wage-premia-and-why-not-use-standard-akm"&gt;Q4. How does the paper measure firm wage premia and why not use standard AKM?&lt;/h3&gt;
&lt;p&gt;Standard AKM firm effects assume time-invariant firm effects and rely on worker mobility; short panels yield noisy estimates with upward-biased variance. The author needs time-varying premia (to measure effective labor over time). He uses the BLM (Bonhomme, Lamadon, Manresa 2019) k-means approach: cluster firms by the similarity of their internal wage distributions (by 2-digit sector over overlapping 2-year windows), then run an AKM-style regression with firm-GROUP effects that vary by year, identified by workers switching between firm-groups — greatly increasing the number of switchers. DADS-Postes is used for clustering (broad coverage) and DADS-Panel for the wage-premium regression.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented-across-firms"&gt;Q5. What heterogeneity is documented across firms?&lt;/h3&gt;
&lt;p&gt;Firm wage premia dispersion accounts for 5.2% of wage dispersion; the 90-10 premium gap is ~30% (about 4 euros/hour, 25% of the median worker&amp;rsquo;s hourly wage), IQR 15%. Markdowns increase with firm wage premia (flat gradient) but DECREASE with firm size — larger firms have more monopsony power, consistent with oligopsony models. Firm-specific labor supply elasticities are 0.54/0.85/1.33 at the 25th/50th/75th percentiles. About 7% of firms have labor wedges above 1, and these tend to have much higher markups (rationalized by kappa&amp;gt;0). In the GE model, top-decile high-wage firms are ~15% more productive but have over 100% greater product quality than bottom-decile firms; amenities rise slightly more steeply with premia than productivity. Passthrough is substantially smaller for 90th-percentile firms (0.74 productivity, 0.18 quality) than for median/10th-percentile firms (~1.06/~0.26).&lt;/p&gt;
&lt;h3 id="q6-how-is-the-dispersion-of-wage-premia-decomposed-across-sources-of-firm-heterogeneity"&gt;Q6. How is the dispersion of wage premia decomposed across sources of firm heterogeneity?&lt;/h3&gt;
&lt;p&gt;Introducing one source at a time into the GE model and comparing variance to baseline (Table 6): varying only product quality reproduces 161.5% of baseline variance, only TFP 153.3%, and only amenities 40.8%. Product quality is the largest single contributor to wage-premium dispersion, closely followed by productivity, then amenities.&lt;/p&gt;
&lt;h3 id="q7-why-does-the-productivity-passthrough-differ-so-much-from-the-quality-passthrough"&gt;Q7. Why does the productivity passthrough differ so much from the quality passthrough?&lt;/h3&gt;
&lt;p&gt;Total passthrough is 0.97 for a 1% productivity shock vs 0.23 for an equal quality shock (~4x). The decomposition (Table 5) attributes most of the gap to the direct effect (1.07 vs 0.26): with high within-market substitutability (sigma=5.17), consumers are very price-sensitive, so productivity (which lowers price) moves sales and labor demand far more than quality. Higher sigma raises productivity passthrough but lowers quality passthrough. For sufficiently low sigma the ranking can reverse. The variable-market-power channel also matters: higher productivity raises markups, increasing rent-sharing (+0.06 via labor wedge) but also output restriction (-0.09 via markup), with output restriction dominating; firm-size effects (sectoral price -0.10, sectoral wage +0.03) further adjust passthrough. Amenity shocks have direct effect -0.26 (mirror of quality) but total -0.28, amplified because better amenities lower hiring costs and expand the firm.&lt;/p&gt;
&lt;h3 id="q8-how-does-worker-bargaining-power-affect-welfare-and-what-are-the-limits"&gt;Q8. How does worker bargaining power affect welfare, and what are the limits?&lt;/h3&gt;
&lt;p&gt;Bargaining power offsets two distortions firm market power imposes on aggregate labor demand: a uniform tax (Lambda/mu, lowering labor demand proportionally) and a misallocation tax (Theta, from dispersion in wedges). There exists a kappa-bar that exactly cancels the uniform tax, and kappa-bar falls as markups rise (high markups make bargaining more effective). With full bargaining power and common markups, the markdown-driven misallocation tax is fully neutralized. BUT bargaining only acts through labor demand; markups also distort capital and material demand, which bargaining cannot fix. Quantitatively: raising kappa from 0.12 to 0.50 matches the wage gain of removing markups but yields only 10% welfare gain (vs 38%) and far less dispersion increase; full kappa=1 raises welfare 13%, under one-third of the planner&amp;rsquo;s 46% gain. So bargaining power is a partial, not full, remedy for firm market power.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-welfare-accounting-for-markups-vs-markdowns"&gt;Q9. What is the welfare accounting for markups vs markdowns?&lt;/h3&gt;
&lt;p&gt;Comparing the decentralized economy to the social planner&amp;rsquo;s (Table 7, column 3): eliminating both markups and markdowns raises wage-premium dispersion 113%, average wages 303%, and welfare 46% (consumption-equivalent). Over 80% of the welfare gain comes from removing markups. Equalizing markups alone (column 4) gives 24% welfare, +39% wages, +99% wage variance, implying ~63% of the markup welfare cost is misallocation. Equalizing markdowns alone (column 5) has little welfare effect (2%), though a wide markdown level reduces welfare significantly (column 2).&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-and-caveats-does-the-author-flag"&gt;Q10. What robustness checks and caveats does the author flag?&lt;/h3&gt;
&lt;p&gt;Caveats: (1) Multiplication bias — mismeasured output elasticities enter both labor wedges and product-market rents multiplicatively, mechanically biasing kappa upward (Appendix B.10); IV with lags only fixes classical, not serially-correlated, measurement error. (2) Labor adjustment costs get absorbed into the labor wedge and bias kappa; firm fixed effects do not fully fix this (Appendix B.11). (3) The markdown estimation imposes that all markdown variation reflects firm size and amenities — more general than kappa=0 approaches but restrictive in this dimension. (4) The model uses collective (not individual) bargaining and abstracts from sequential-auction wage-setting (Cahuc-Postel-Vinay-Robin); robustness to hiring-wages-only following Di Addario et al. (2020) is shown (Appendix B). (5) Worker types assumed perfect substitutes; an Appendix E two-skill extension gives similar results. (6) Empirical patterns hold without TFPQ controls (Figure D.3) and by firm size (Figure D.4).&lt;/p&gt;
&lt;h3 id="q11-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q11. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Versus the labor-market-power literature (Berger et al. 2022; Lamadon et al. 2022) it adds product-market power and bargaining, showing their pure-monopsony labor wedge is a kappa=0 special case. Versus the markups/welfare literature (De Loecker et al. 2020; Edmond et al. 2023) it adds imperfect labor competition and bargaining. Versus recent integrated product+labor power models that use wage-posting and no bargaining (Kroft et al. 2024; Deb et al. 2024), it adds the rent-sharing channel where markups raise (not just lower) the labor wedge. Versus production-approach markdown estimation (Yeh et al. 2022; Mertens 2020), it shows their estimates are labor wedges (not markdowns) once kappa&amp;gt;0, and that omitting hours upward-biases them. Versus the rent-sharing literature (Card et al. 2018; Kline et al. 2019; Van Reenen 1996), it shows their instruments violate exclusion under endogenous monopsony wages and proposes the labor-wedge-equation alternative. The closest exception incorporating unions is Azkarate-Askasua and Zerecero (2025).&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Strengthening worker collective bargaining power can raise welfare mainly by offsetting markup-induced distortions to labor demand and redistributing rents, but it raises between-firm wage inequality and cannot restore full efficiency because it leaves markup distortions to capital/material untouched (full kappa closes under one-third of the planner gap). The wage effects of innovation depend on whether it improves productivity or quality and on the degree of product differentiation. Scope conditions: estimates are for French manufacturing under firm-level collective bargaining institutions (firms &amp;gt;=50 employees legally bargain annually); results rely on the production-approach assumptions (flexible/price-taken materials, scalar unobservability) and on data including hours and output prices that many countries lack — researchers should interpret labor-wedge/markup moments cautiously without hours data.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>What Drives the Recent Surge in Inflation? The Historical Decomposition Roller Coaster</title><link>https://macropaperwarehouse.com/papers/what-drives-the-recent-surge-in-inflation-the-historical-decomposition-roller-coaster/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/what-drives-the-recent-surge-in-inflation-the-historical-decomposition-roller-coaster/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;The paper addresses what drove the post-COVID inflation surge in the United States and internationally. Before answering the substantive question, the authors identify and diagnose a methodological obstacle: the standard tool used for such analysis — the historical shock decomposition in a structural VAR — can produce wildly inconsistent narratives depending on small, likelihood-inconsequential changes in the model&amp;rsquo;s parameters.&lt;/p&gt;
&lt;p&gt;The mathematical core is the VAR decomposition of observed data into a deterministic component (DC, the model&amp;rsquo;s period-zero forecast in the absence of any realized shocks) and a stochastic component (SC, the discounted cumulative sum of shock contributions). Because DC and SC sum to data, imprecision in DC is mechanically transmitted to SC, making inferences about shock contributions unreliable. The authors establish that conditional likelihood-based estimation leaves the VAR constant C poorly identified: parameter perturbations that move the likelihood only negligibly can shift DC dramatically. This &amp;ldquo;excess volatility&amp;rdquo; in DC is distinct from the better-known overfitting problem: excess volatility is about cross-draw uncertainty in DC, not its average level, and can be severe even when overfitting is mild.&lt;/p&gt;
&lt;p&gt;The illustrative case is a bivariate SVAR of US real GDP and the GDP deflator (log first differences, 1983:Q1–2022:Q4, four lags, sign restrictions, Jeffreys diffuse prior). The three draws closest to the point-wise median impulse response — draws whose impulse responses are virtually indistinguishable — produce entirely contradictory post-pandemic narratives: the first assigns more than two-thirds of the inflation rise to supply shocks, the second assigns more than two-thirds to demand shocks, and the third assigns roughly equal shares. The US GDP deflator peaked at 7.7 percent in 2022:Q2; euro area inflation peaked around 10 percent on an annual basis, with some European countries exceeding 15 percent in 2022.&lt;/p&gt;
&lt;p&gt;The excess volatility problem is shown to be pervasive: it arises regardless of identification scheme (sign restrictions, Blanchard-Quah long-run restrictions, Cholesky zero-impact restrictions), persists with standard priors (Normal-Inverse Wishart and Minnesota) that shrink AR coefficients but leave the constant diffuse, worsens with longer or more heterogeneous samples (the 1949:Q1–2022:Q4 sample produces substantially larger dispersion than the baseline), and survives in larger VAR systems (the problem is if anything more severe in a 5-variable BVAR).&lt;/p&gt;
&lt;p&gt;The preferred solution is the single-unit-root prior (Sims 1993), implemented as a dummy initial observation that constrains the VAR&amp;rsquo;s unconditional mean to the sample average. As the tightness hyperparameter δ → 0, DC converges across all posterior draws to a common value. The modal posterior value of δ, estimated data-adaptively using the approach of Giannone et al. (2015) with a Gamma prior of mode 1, is 0.0001 for US data — indicating the data strongly favor tight shrinkage. In simulations, after roughly 20 periods, all 1,000 draws of DC converge to virtually identical values regardless of data persistence or sample size.&lt;/p&gt;
&lt;p&gt;With the single-unit-root prior, the US results are unambiguous: supply shocks were important in the initial phase of the inflation surge, but demand factors became the main driver from 2021 onward, accounting for 56 percent of inflation fluctuations in 2021 and 77 percent in 2022. Two pragmatic alternatives for frequentists — demeaning the data prior to estimation, and computing point-wise median historical decompositions — both corroborate demand dominance.&lt;/p&gt;
&lt;p&gt;International evidence is estimated using the same bivariate SVAR and identification restrictions. For the euro area (industrial production and HICP inflation, 2001:M1–2023:M3), demand factors account for more than 50 percent of inflation fluctuations in 2022, but supply shocks remain significant through at least mid-2023, reflecting the region&amp;rsquo;s greater exposure to the Ukraine-war commodity supply shock. For four small open economies (Norway, Sweden, Canada, Australia; quarterly GDP growth and year-on-year CPI inflation, 1993:Q1–2023:Q2), the pattern closely resembles the US: supply shocks dominate in 2020, but demand forces become prevalent already in 2021 and are nearly dominant in some cases thereafter. The finding that demand factors were the primary driver of the inflation surge thus holds robustly across six economies with heterogeneous policy responses, supply-chain exposures, and Ukraine-war commodity price effects. The policy implication is that the aggressive monetary tightening implemented by central banks was appropriate given the demand-driven nature of the surge — though the paper is careful to note that its &amp;ldquo;demand shock&amp;rdquo; aggregates monetary, fiscal, and other demand-side disturbances, limiting precise policy prescriptions.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The baseline uses sign restrictions: a demand shock moves real GDP and the GDP deflator in the same direction on impact; a supply shock moves them in opposite directions. Restrictions are imposed only on impact, following Canova and De Nicolo (2002). The authors acknowledge that the demand shock bundles monetary, fiscal, and other demand-side disturbances, while the supply shock aggregates productivity, commodity, markup, and other supply-side factors. Blanchard-Quah (long-run zero restrictions) and Cholesky (impact zero restrictions) are used as alternative schemes to show the excess-volatility problem is identification-independent. The main threat to credible decompositions is not misidentification of shocks per se but rather imprecision in the VAR&amp;rsquo;s deterministic component, which contaminates all inferences about shock contributions regardless of the identification scheme.&lt;/p&gt;
&lt;h3 id="q2-what-exactly-is-the-excess-volatility-problem-and-why-does-it-arise"&gt;Q2. What exactly is the excess volatility problem and why does it arise?&lt;/h3&gt;
&lt;p&gt;The VAR&amp;rsquo;s deterministic component DC_t depends on the companion matrix A and the constant vector C. Conditional likelihood-based estimation identifies A well — impulse responses are relatively precisely estimated — but leaves C poorly pinned down, because many combinations of (A, C) produce nearly identical likelihood values while implying very different unconditional means and thus very different DC paths. Even parameter perturbations negligible relative to the likelihood surface can shift DC dramatically. Because the stochastic component SC_t = Data - DC_t, imprecision in DC is mechanically transmitted to SC_t and to estimated shock contributions. The problem is a property of the reduced-form model and arises before any structural identification is imposed.&lt;/p&gt;
&lt;h3 id="q3-how-is-excess-volatility-distinguished-from-the-overfitting-problem"&gt;Q3. How is excess volatility distinguished from the overfitting problem?&lt;/h3&gt;
&lt;p&gt;Overfitting (Sims 1996, 2000; Giannone et al. 2019) refers to the deterministic component attributing an implausibly large share of low-frequency data variation to itself — the DC level tracks the data in-sample but implies poor out-of-sample forecasts. Excess volatility refers to the uncertainty across posterior draws in DC, not the average level of DC. A model can exhibit mild overfitting (as in the baseline bivariate model, whose DC paths stabilize after only two or three years) while having extreme excess volatility across draws. Solving the overfitting problem — for example by using the prior for the long run (Giannone et al. 2019) — does not solve the excess volatility problem. The single-unit-root prior addresses both, but for distinct reasons.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-single-unit-root-prior-solve-the-excess-volatility-problem-technically"&gt;Q4. How does the single-unit-root prior solve the excess volatility problem technically?&lt;/h3&gt;
&lt;p&gt;The prior adds a dummy observation that imposes the stochastic constraint [I − A]Ȳ₀ − C = δu₀, where Ȳ₀ is set to the sample average and δ governs tightness. Substituting into the DC formula shows that, for a stationary ergodic system, A^t(Y₀ − Ȳ₀) → 0 as t grows, so DC_t converges across all posterior draws to Ȳ₀. The hyperparameter δ is estimated from the data using a Gamma prior with mode 1, following Giannone et al. (2015). The modal posterior value is 0.0001 with negligible posterior dispersion, indicating strong data support for near-exact shrinkage. The prior does not eliminate uncertainty in the stochastic component — draws of A and F still produce variation in shock contributions — but that remaining uncertainty is the same type as in impulse response estimation, making the two statistics mutually consistent.&lt;/p&gt;
&lt;h3 id="q5-why-do-standard-priors-normal-inverse-wishart-minnesota-fail-to-solve-the-problem"&gt;Q5. Why do standard priors (Normal-Inverse Wishart, Minnesota) fail to solve the problem?&lt;/h3&gt;
&lt;p&gt;Standard priors shrink the AR coefficient matrices and the residual covariance matrix but leave the prior on the VAR constant C diffuse. Because the excess volatility arises specifically from poorly identified values of C, these priors leave the deterministic component as uncertain as with a diffuse prior. The paper demonstrates this directly by plotting deterministic component draws under Normal-Inverse Wishart and Minnesota priors (Figure 3, rows 2) — the dispersion remains large and whimsical historical decompositions persist.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-is-documented-across-countries"&gt;Q6. What heterogeneity is documented across countries?&lt;/h3&gt;
&lt;p&gt;The euro area shows a more balanced demand-supply split than the US: demand and supply factors contribute roughly equally overall, with demand becoming prevalent in 2022 (exceeding 50 percent of inflation fluctuations) but supply shocks remaining significant through mid-2023. The authors attribute this persistence of supply shocks in the euro area to the region&amp;rsquo;s greater exposure to the Russia-Ukraine energy supply disruption. The four small open economies (Norway, Sweden, Canada, Australia) have outcomes surprisingly similar to the US: supply shocks drive inflation in 2020, demand becomes prevalent in 2021 and is nearly dominant in some cases in 2022. Overall, despite heterogeneity in fiscal stimulus, supply-chain exposure, and commodity price effects, demand factors are the primary driver across all six economies examined.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-run"&gt;Q7. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper runs five main robustness exercises. (1) Three identification schemes — sign restrictions, Blanchard-Quah, and Cholesky — all exhibit the same excess-volatility problem under diffuse priors and produce similar demand-dominance results with the single-unit-root prior. (2) Four prior specifications — diffuse, Normal-Inverse Wishart, Minnesota, single-unit-root — are compared using a proposed dispersion measure (max-minus-min across top 100 draws, averaged over time); the single-unit-root prior uniformly produces the smallest dispersion across all identification schemes. (3) Two sample periods for the US: the baseline 1983:Q1–2022:Q4 and an extended 1949:Q1–2022:Q4 sample; excess volatility is substantially larger with the longer, heterogeneous sample. (4) A 5-variable VAR (real GDP, GDP deflator, real private investment, federal funds rate, real wages), baseline sample and diffuse prior — the excess-volatility problem remains and is more severe for variables like inflation and the federal funds rate. (5) Two alternative approaches for frequentists (demeaning the data; computing point-wise median historical decompositions) both reproduce the demand-dominance finding.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-two-pragmatic-alternatives-offered-for-researchers-reluctant-to-use-priors"&gt;Q8. What are the two pragmatic alternatives offered for researchers reluctant to use priors?&lt;/h3&gt;
&lt;p&gt;First, demeaning all variables before estimation and estimating the VAR without a constant. This eliminates the first term of DC (which depends on C) and forces DC to follow A^t·Y₀, which approaches zero for stationary systems. It is a partial solution — draws with different A matrices still produce different DC paths, so dispersion is reduced but not eliminated; dispersion is smaller than under a diffuse prior but larger than under the single-unit-root prior. Second, computing the point-wise median historical decomposition: across all posterior draws, take the median contribution of each shock at each date. The resulting summary is non-additive (a residual deterministic component absorbs the gap between data and the two median stochastic components) but robust to outliers and reflective of parameter uncertainty. Bergholt et al. (2023) use this approach in prior work. The paper shows that median decompositions under all four prior specifications deliver demand-dominance conclusions similar to those from the single-unit-root prior.&lt;/p&gt;
&lt;h3 id="q9-what-dispersion-measure-do-the-authors-propose-and-what-do-the-numbers-show"&gt;Q9. What dispersion measure do the authors propose, and what do the numbers show?&lt;/h3&gt;
&lt;p&gt;The authors define D_{i,j,t} as the max-minus-min spread of shock j&amp;rsquo;s contribution to variable i at time t across the 100 draws closest to the point-wise median impulse response. M_{i,j} is the time-average of D_{i,j,t}. Applied to the contribution of demand shocks to US inflation over 2020:Q2–2022:Q4, the values are: diffuse prior — 1.07 (sign), 0.88 (Blanchard-Quah), 2.33 (Cholesky); Normal-Inverse Wishart — 1.53, 1.20, 0.91; Minnesota — 0.87, 0.71, 0.61; single-unit-root — 0.68, 0.48, 0.54. The single-unit-root prior produces the smallest dispersion uniformly across all identification schemes.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q10. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Bernanke and Blanchard (2024) use a simple wage-price dynamic model and find most of the surge resulted from shocks to prices given wages. Rubbo (2023) uses disaggregated price data and finds roughly three-quarters of the CPI rise since 2021 is demand-driven. Eickmeier and Hofmann (2022) use a large factor model and find demand predominant. Ascari et al. (2023) use a Bayesian SVAR on euro area data and find demand factors crucial from fall 2020. The present paper&amp;rsquo;s demand-dominance conclusion is broadly consistent with this literature. Its distinctive contribution is not the substantive finding but the methodological diagnosis: it shows that standard VAR-based historical decompositions are whimsical under diffuse priors, explains why, and provides credible solutions. It also contributes international evidence spanning six economies with comparable methodology.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The finding that demand factors were the primary driver of the post-COVID inflation surge supports the appropriateness of the aggressive monetary tightening implemented by the Federal Reserve and other central banks. A demand-driven inflation surge calls for a different policy response than a supply-driven one; the paper&amp;rsquo;s results vindicate the central bank interpretation that monetary tightening was warranted. However, scope conditions are important: the identified &amp;lsquo;demand shock&amp;rsquo; aggregates monetary, fiscal, and other demand-side disturbances; the paper cannot decompose the demand category further into, for example, fiscal stimulus versus pent-up household demand. Additionally, the bivariate model omits many potentially relevant variables. The policy implication applies to the broad nature of the shock (demand vs. supply) and does not prescribe specific instruments or magnitudes of policy response.&lt;/p&gt;
&lt;h3 id="q12-what-future-research-directions-are-identified"&gt;Q12. What future research directions are identified?&lt;/h3&gt;
&lt;p&gt;The authors note that the excess volatility problem is even more acute when separating permanent from transitory components of data, because imprecision in DC translates directly into imprecision in the level of the permanent component. In small samples, long-run shock contributions are also imprecisely estimated, compounding the problem. These issues make estimates of trend inflation poor and inflation regimes difficult to characterize. The authors flag this as a planned area of future research.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Deterministic component (DC_t)&lt;/strong&gt;: The period-zero forecast of the endogenous variables in the absence of any unforecastable shock realizations — the counterfactual trajectory the VAR assigns based on its parameters and initial conditions alone. Not a statistical trend, but the baseline path the model says would have prevailed had no shocks occurred.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stochastic component (SC_t)&lt;/strong&gt;: The discounted cumulative sum of all structural shock realizations from period 1 through period t. Together with the deterministic component, it sums to the observed data; it is the part of the observed series attributable to identified economic shocks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Historical shock decomposition&lt;/strong&gt;: The retrospective attribution of observed data fluctuations at each point in time to the contributions of individual identified structural shocks. Distinct from the impulse response function (which characterizes prospective shock propagation): the historical decomposition integrates shock realizations and is thus a function of the stochastic component&amp;rsquo;s draw-specific paths.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Excess volatility (of the deterministic component)&lt;/strong&gt;: The phenomenon whereby posterior draws of VAR parameters that produce nearly identical impulse response functions nevertheless imply radically different paths for the deterministic component. Caused by the likelihood surface being nearly flat with respect to the VAR constant C. Distinct from overfitting: excess volatility is cross-draw uncertainty in DC, not the average level of DC.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Single-unit-root prior (dummy initial observations prior)&lt;/strong&gt;: A prior on VAR parameters implemented by adding one artificial observation, where both current and lagged values equal (1/δ)·Ȳ₀ and the intercept equals 1/δ. As tightness parameter δ → 0, the prior constrains the VAR&amp;rsquo;s unconditional mean to equal Ȳ₀ across all posterior draws, eliminating excess volatility in DC while leaving structural shock uncertainty intact.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dispersion measure (M_{i,j})&lt;/strong&gt;: The authors&amp;rsquo; proposed metric for quantifying how whimsical a historical decomposition is: the time-average of the max-minus-min spread of shock j&amp;rsquo;s contribution to variable i across the 100 draws closest to the point-wise median impulse response. Smaller values indicate more robust, less draw-dependent decompositions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Whimsical historical decomposition&lt;/strong&gt;: The paper&amp;rsquo;s term for a shock decomposition whose narrative about the relative importance of structural drivers changes substantially across draws that are otherwise observationally equivalent in terms of impulse responses. Caused by excess volatility in the deterministic component forcing shocks to compensate for different DC paths.&lt;/p&gt;</description></item><item><title>Banks of a Feather: The Informational Advantage of Being Alike</title><link>https://macropaperwarehouse.com/papers/banks-of-a-feather-the-informational-advantage-of-being-alike/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/banks-of-a-feather-the-informational-advantage-of-being-alike/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Can banks effectively monitor their peers under asymmetric information? Effective peer monitoring matters for functioning interbank markets and, by implication, financial markets and the transmission of monetary policy. If banks monitor effectively, central banks can stay in a &amp;ldquo;night-watchman&amp;rdquo; role (Goodfriend and King 1988); if they systematically fail to identify solvent counterparties, central banks should be more active (Freixas and Jorge 2008). The paper argues that PORTFOLIO SIMILARITY between two banks is the key to their reciprocal monitoring ability: a lender uses private information about its own loan portfolio to assess the quality of a peer&amp;rsquo;s portfolio, so it is better informed the more similar the two exposures.&lt;/p&gt;
&lt;p&gt;Data and setup: Quarterly bilateral bank-to-bank and bank-to-firm exposures from the German credit register, 2009-2018, covering 2,054 lending and 2,035 borrowing banks, balanced into 2,644,640 lender-borrower-quarter combinations; 701,533 true credit relations (102,044 within the same banking network, 2,087 within the same holding company). Interbank exposure represents 21% of German banks&amp;rsquo; total borrowing and 20% of total lending; ~1.4 trillion euros average quarterly exposure by end-2018. The authors build three novel measures: (1) Portfolio quality = 1 minus the exposure-weighted average probability of default (PD) from proprietary supervisory filings (a forward-looking, private quality proxy); (2) Portfolio opacity = exposure-weighted standard deviation of PDs different banks assign to the same borrower (peers&amp;rsquo; disagreement); (3) Portfolio similarity = cosine similarity of two banks&amp;rsquo; exposure vectors across 10 industries (WZ 73 one-digit) and 9 regions (first zip digit). Estimation uses a Heckman (1977) two-step sample selection model: a Probit selection equation for the extensive margin (whether a credit relation exists) and an OLS outcome equation for the intensive margin (percentage change in bilateral exposure), with lagged credit relation as exclusion restriction, plus lender, borrower and quarter-year fixed effects. Independent variables are standardized.&lt;/p&gt;
&lt;p&gt;Main findings (signs, magnitudes, scope): Portfolio quality validation - it negatively and significantly predicts next-quarter NPL ratios up to 2 years ahead, explaining 16-17% of cross-sectional NPL variation and 71-77% with fixed effects. For the AVERAGE bank, lending does NOT respond to borrower Portfolio quality (coefficients negative, mostly insignificant), but DOES respond to the backward-looking NPL ratio: a one-SD higher borrower NPL ratio lowers the probability of receiving a loan by 118 basis points (vs. unconditional 26.53%) and reduces amounts by 133-236 bp (avg. quarterly change 1.46%). Higher borrower Portfolio opacity reduces lending (extensive -38 bp; intensive -57 to -111 bp). The key result: interacting similarity with quality reverses this for similar pairs. For HIGH-similarity pairs (3 SD above mean), a one-SD increase in borrower Portfolio quality raises matching probability by 50 bp and lending by 408 bp; a deterioration cuts lending by 348-368 bp (avg. change between similar banks 10.95%). For LOW-similarity pairs, higher Portfolio quality LOWERS lending (matching -80 bp; amount -563 bp), and lending rises after quality deteriorates (370/342 bp), which Section 6 shows is a demand effect. The NPL-ratio response vanishes for similar pairs. Portfolio similarity itself raises lending: one-SD more sectoral similarity raises intensive-margin lending ~100-259 bp, regional similarity ~84-114 bp - jointly comparable in magnitude to relationship lending, the strongest known predictor. For opaque borrowers, high-similarity lenders lend MORE (extensive +23 bp; intensive +129 to +162 bp). A variance decomposition (Lemmon et al. 2008 ANCOVA) finds common/bank-pair characteristics explain 98.0% of extensive-margin variation and 18.9% of intensive-margin variation; lender, borrower and market characteristics explain only 1.2/0.8/0.1% (extensive) and 35.6/44.2/9.1% (intensive). Implication: peer monitoring works, but only among similar banks; this raises interbank efficiency at the cost of higher systemic risk and too-interconnected-to-fail concerns.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The core estimation is a Heckman (1977) two-step sample selection model: a first-stage Probit for the extensive margin (existence of a bilateral credit relation) and a second-stage OLS for the intensive margin (log change in bilateral exposure), with the inverse Mills ratio carried into the second stage. The exclusion restriction is the lagged existence of a credit relation (Credit relation_{i,j,t-1}), which strongly predicts a current relation (first-stage t-statistic 335; t=293 in the similarity specification) because German interbank exposures are long-lived, yet carries no information on whether exposure will rise or fall next quarter. The chief threats are: (1) demand vs. supply confounding - observed lending is equilibrium, so a negative quality-lending link could reflect borrowers&amp;rsquo; demand rather than lenders&amp;rsquo; screening; addressed in Section 6. (2) Correlated portfolio quality of similar banks - a lender cutting lending in response to its OWN deteriorating portfolio could be misread as a reaction to a similar borrower&amp;rsquo;s portfolio; addressed via a matched sample in Section 7. The paper also notes both Portfolio quality and NPL series are persistent, so the predictive regressions should be read as &amp;lsquo;gentle evidence,&amp;rsquo; not strict causal proof.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-authors-separate-supply-effects-from-demand-effects"&gt;Q2. How do the authors separate supply effects from demand effects?&lt;/h3&gt;
&lt;p&gt;They adapt Degryse et al. (2019). They define an adjusted exposure change bounded in [-2,2] (Chodorow-Reich 2014; Davis-Haltiwanger 1992) that captures both margins, then regress it on lending-bank-time fixed effects (proxying supply) and borrowing-bank-class x industry x region x time fixed effects (proxying demand, assuming homogeneous demand across lenders). The estimated lender-time fixed effects, demeaned and aggregated to the borrowing-bank level, give a borrower-specific liquidity-supply shock. Regressing this on borrower Portfolio quality, NPL ratio and opacity shows supply is restricted when quality deteriorates, NPL rises, or opacity increases. This confirms the puzzling positive lending-to-low-quality result for dissimilar pairs is a DEMAND effect: low-quality borrowers, shunned by similar lenders, demand more liquidity and turn to dissimilar lenders. The authors stress this borrower-level approach supports but cannot replace the bank-pair analysis, since it cannot include pair characteristics like similarity.&lt;/p&gt;
&lt;h3 id="q3-how-do-they-rule-out-that-lenders-are-just-reacting-to-their-own-correlated-portfolio-quality"&gt;Q3. How do they rule out that lenders are just reacting to their own correlated portfolio quality?&lt;/h3&gt;
&lt;p&gt;In the full sample, the correlation of Portfolio quality between two above-average-similarity banks is 0.0499 versus only 0.0150 for below-average-similarity pairs. They build a matched subsample (nearest-neighbour matching, assigning each &amp;lsquo;similar&amp;rsquo; pair - both similarities above the 75th percentile in 2009Q1 - three &amp;lsquo;dissimilar&amp;rsquo; pairs below the 25th percentile with the closest Portfolio-quality correlation) so that within-pair quality correlation is the same for similar and dissimilar pairs, and redefine similarity as binary. If lenders only reacted to their own portfolio, the similarity x quality interaction should vanish in this sample. Instead, the interaction stays positive and mostly significant (and NPL x similarity too); weaker significance in some fixed-effect models reflects the smaller sample, since coefficient sizes are comparable to the main results.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q4. What are the main mechanisms, and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;Mechanism: information on a peer&amp;rsquo;s asset quality is private and costly to obtain; a lender proxies a peer&amp;rsquo;s portfolio quality by the average quality of the industries/regions it lends to, and can do this more cheaply when it already lends to the same industries/regions (similar portfolio). So similar lenders are better informed. Empirically distinguished by: (a) the average bank reacts to the public NPL ratio but not to private Portfolio quality, while similar pairs react strongly to Portfolio quality and barely to NPL - showing similar lenders access private information; (b) the similarity x quality and similarity x opacity interactions; (c) the supply-shock decomposition separating screening from demand; (d) the matched sample ruling out own-portfolio reactions. A competing mechanism, risk shifting (Elliott et al. 2018) - banks deliberately courting correlated counterparties to raise bailout probability - cannot be ruled out and may co-drive preferential lending between similar peers.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented"&gt;Q5. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;(1) By similarity: similar pairs (3 SD above mean) react to forward-looking Portfolio quality and lend more to higher-quality and more-opaque peers; dissimilar pairs (3 SD below mean) react only to the backward-looking NPL ratio and end up lending more to low-quality borrowers via demand. (2) By opacity: lending between similar banks is especially important for opaque borrowers, who otherwise struggle to refinance; opaque banks are shunned by dissimilar lenders and turn to similar ones, while low-quality banks are shunned by similar lenders and turn to dissimilar ones. (3) Sectoral vs. regional similarity: both matter; sectoral similarity tends to have larger intensive-margin effects (e.g., 259 vs. 94 bp in Model 3). (4) Lender&amp;rsquo;s own quality: lenders cut lending when their own Portfolio quality falls (one-SD drop reduces amounts by 215-226 bp within-bank), consistent with prior work (Acharya-Merrouche 2013).&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-and-additional-analyses-are-run"&gt;Q6. What robustness checks and additional analyses are run?&lt;/h3&gt;
&lt;p&gt;(1) Multiple fixed-effect layers: cross-section, lender/borrower fixed effects, and added quarter-year fixed effects (Models 1-4 across tables). (2) Control set: lagged Capital ratio, Liquidity ratio, ROA, Loans-to-assets, Size, relationship lending and reverse relationship lending over an 8-quarter window, difference in liquidity surplus, same-network and same-holding-company dummies. (3) Supply-vs-demand decomposition (Section 6). (4) Matched-sample analysis breaking the quality correlation (Section 7). (5) Validation of Portfolio quality via NPL-predictive regressions and a panel Granger causality test (Juodis et al. 2021; Half-Panel Jackknife Wald &amp;gt; 300; Dumitrescu-Hurlin Z &amp;lt; -50), significant 5-50 quarters ahead. (6) Two-digit WZ 73 industry classification (100 industries) in Appendix B. (7) Variance decomposition (ANCOVA, Type III sums of squares) quantifying explanatory power.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It extends peer-monitoring literature (Goodfriend-King 1988; Rochet-Tirole 1996; Flannery-Sorescu 1996; Furfine 2001) by showing that even among banks, the more similar the lender, the better its monitoring - identifying Perignon et al. (2018)&amp;rsquo;s &amp;lsquo;informed lenders&amp;rsquo; as similar-portfolio banks. Versus relationship-lending work (Affinito 2012; Braeuning-Fecht 2017; Cocco et al. 2009), it shows that with a similar portfolio NO long-standing relationship is needed to obtain quality information, and that similarity mitigates opaque banks&amp;rsquo; hampered access on top of relationships. It augments lender/borrower/market-characteristic studies by adding dyadic (common) covariates. Unlike prior work using aggregate bank-level ratios, CDS spreads, or rating-agency disagreement, it uses granular real-exposure data and proprietary supervisory PDs to measure private quality and peer-perceived opacity directly. It links to systemic-risk/contagion literature (Allen-Gale 2000; Fecht et al. 2011; Elliott et al. 2018), showing banks over-expose to similar counterparties despite indirect-contagion risk, surfacing an efficiency-vs-systemic-risk trade-off akin to focus-vs-diversification in Acharya et al. (2006).&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Peer monitoring is real but partial: only similar banks effectively screen on private, forward-looking quality, while others fall back on inferior public proxies (NPL ratios). This bears on the central-bank &amp;rsquo;night-watchman vs. active&amp;rsquo; debate - because monitoring fails for dissimilar pairs, a purely hands-off stance may be insufficient. The headline trade-off: stronger lending between similar banks raises interbank informational efficiency and monitoring, but the above-average direct exposure between similar (correlated) banks multiplies systemic risk and too-interconnected-to-fail concerns, and reflects a lack of diversification. Scope conditions: results are specific to the German banking system (2009-2018), a tiered market dominated by private, savings, and cooperative banks with mostly long-term interbank loans (45% over a year, only 15% overnight); the data lack interest rates, so the analysis covers quantities/existence of lending, not prices; effects are estimated on bank-pairs that lent at least once; and the supply-identification assumes homogeneous borrower demand across lenders.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-key-caveats-the-authors-themselves-flag"&gt;Q9. What are the key caveats the authors themselves flag?&lt;/h3&gt;
&lt;p&gt;(1) No interest-rate data, so price effects of similarity, quality and opacity are untested. (2) Portfolio quality and NPL series are persistent, so the forward-looking predictive evidence is &amp;lsquo;gentle,&amp;rsquo; not definitive. (3) The supply-shock approach gives borrower-level (not pair-level) shocks and cannot incorporate similarity. (4) Risk shifting cannot be ruled out as a co-driver of preferential lending between similar peers. (5) Portfolio quality is built using the median PD across IRB banks, excluding borrowers exposed only to Standardised-Approach banks. (6) The balanced sample includes only pairs that lent at least once, ignoring pairs that could theoretically but realistically would not lend (consistent with tiered-market evidence).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Financial Stability with Fire Sale Externalities</title><link>https://macropaperwarehouse.com/papers/financial-stability-with-fire-sale-externalities/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/financial-stability-with-fire-sale-externalities/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Asset fire sales were a defining feature of the 2007-08 crisis, and post-crisis reforms (Basel III liquidity requirements, Money Market Mutual Fund reforms) were introduced to mitigate fire sale externalities by reducing distressed debt obligations and forcing larger liquidity buffers. The paper asks whether policies that successfully mitigate fire sale externalities actually improve financial stability, since it is not obvious how banks re-optimize in response.&lt;/p&gt;
&lt;p&gt;Model setup (no empirical data — this is a theoretical paper): The authors build a three-period (t = 0,1,2) Diamond-Dybvig (1983) model of financial intermediation augmented with (i) cash-in-the-market pricing in a financial market as in Allen and Gale (1998), and (ii) limited commitment as in Ennis and Keister (2009), following Li (2017). A unit continuum of ex ante identical depositors have CRRA preferences with relative risk aversion γ &amp;gt; 1. Each depositor is impatient with known probability π. There are two assets: a short-term storage asset (1 unit yields 1 next period) and a long-term asset (1 unit at t=0 yields R &amp;gt; 1 at t=2). The bank invests fraction x in the long-term asset and 1−x short. Long-term assets can be sold at t=1 at an endogenous price p to risk-neutral investors who receive endowment ws (market liquidity) and have outside return R* &amp;gt; 0. Runs are introduced via a sunspot s ∈ {α, β} with run probability q; runs are partial (stop after fraction π is served), following Ennis and Keister. The authors assume R* = R, which implies p ≤ 1 in equilibrium. Financial fragility is measured by q-bar, the maximum run probability q for which the run strategy is an equilibrium (run condition c1 ≥ c2β).&lt;/p&gt;
&lt;p&gt;Main analytical findings: (1) Without intervention, banks over-invest in long-term assets relative to the socially efficient level because each competitive bank takes p as given and does not internalize that selling long-term assets in a run depresses p (the fire sale externality); the equilibrium price is inefficiently low. (2) The bank&amp;rsquo;s best response is in Case I (no excess liquidity, fire sale occurs) when 0 &amp;lt; q &amp;lt; q_l, and Case II (excess liquidity held) when q_l ≤ q &amp;lt; 1 (Lemma 1). There is a unique q_c at which the market-clearing price p* turns from decreasing to increasing in q (Lemma 3). (3) Comparative statics on market liquidity ws (Proposition 1): when the relevant q-bar lies in Case II (low ws), q-bar is strictly increasing in ws, so a small rise in market liquidity raises fragility; when q-bar lies in Case I (high ws), q-bar is strictly decreasing in ws. The mechanism (Lemmas 4-5) is that a higher p* raises c1 via intertemporal substitution; the c2α/c2β effect is always dominant, flipping the sign of dq-bar/dws between cases. (4) The intervention: a regulator controls (x, c1), internalizing the effect on p, while the bank still chooses (c2α, c1β, c2β) taking p as given. The regulator chooses lower x and higher c1 than the bank in Case I (Lemma 6: c1 ≤ c1R, x ≥ xR), raising the market-clearing price (Proposition 2: p* ≤ pR* in Case I). (5) Key result (Proposition 3): q-bar_R ≥ q-bar when both solutions are in Case I (intervention always raises fragility); ambiguous otherwise. When ws (or R) is high, intervention raises fragility (q-bar_R &amp;gt; q-bar); when ws or R is low, intervention involves excess liquidity and lowers fragility (q-bar_R &amp;lt; q-bar). Proposition 4 gives a sufficient condition for q-bar_R &amp;gt; q-bar via four thresholds ws1≤ws≤ws2 and ws3&amp;lt;ws&amp;lt;ws4. When ws is sufficiently high, p = pR = 1, the externality vanishes, and q-bar = q-bar_R. (6) Welfare (Proposition 5): WR(q-bar) ≤ W(q-bar) when both in Case I, and for some parameter values otherwise — intervention does not always improve welfare and can worsen it when market liquidity is large.&lt;/p&gt;
&lt;p&gt;Policy implication: Mitigating fire sale externalities does not necessarily increase stability. Because the regulator takes q as given, it ignores that its own intervention can raise q-bar. Policymakers must internalize the fragility effect and balance externality mitigation against increased fragility, especially when market liquidity is high.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-is-there-an-identification-strategy-or-empirical-data-what-are-the-threats"&gt;Q1. Is there an identification strategy or empirical data? What are the threats?&lt;/h3&gt;
&lt;p&gt;No. This is a purely theoretical paper with no data, sample period, or estimation. The quantitative content consists of analytical comparative-statics results (Lemmas 1-6, Propositions 1-5) and numerical illustrations rendered as figures (Figures 4-9) for specific parameter combinations of (ws, R, q, γ, π). There is no econometric identification; the analog of robustness is the set of modeling assumptions and the parameter regions over which results hold.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-core-economic-mechanism-and-how-does-intervention-raise-fragility"&gt;Q2. What is the core economic mechanism, and how does intervention raise fragility?&lt;/h3&gt;
&lt;p&gt;The regulator internalizes the fire sale externality by reducing the bank&amp;rsquo;s long-term holdings x and holding more short-term assets, which reduces asset supply in a crisis and raises the market value p of each long-term asset (this mitigates the externality and is the intended benefit). But two competing effects act on long-term payments c2β: the higher price raises the value of remaining long-term assets, while there are fewer long-term assets left for c2β (whose period-2 return R is fixed, so the price increase does not help c2β as it does c1β). The net effect on c2β is ambiguous. Simultaneously, reducing x lowers the relative cost of t=1 consumption, optimally pushing the regulator to raise short-term payment c1. Since the run condition is c1 ≥ c2β, raising c1 while c2β may fall makes early withdrawal more attractive, raising q-bar. When market liquidity is high, the net effect always increases fragility.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-role-of-excess-liquidity-and-how-does-it-reverse-the-result-at-low-market-liquidity"&gt;Q3. What is the role of &amp;rsquo;excess liquidity&amp;rsquo; and how does it reverse the result at low market liquidity?&lt;/h3&gt;
&lt;p&gt;Excess liquidity (Case II: πc1 &amp;lt; 1−x, holding more short-term assets than needed for the first π payments) is the bank&amp;rsquo;s/regulator&amp;rsquo;s hedge against runs. When ws is low, the anticipated fire sale price is low, so the regulator chooses to hold more excess liquidity than the bank. Excess liquidity supplies additional resources to pay c1β and further reduces asset supply (raising p), leaving more resources for c2β. This makes the net effect on c2β favorable enough that q-bar falls. Thus at low market liquidity the regulator can simultaneously mitigate the externality and reduce fragility; at high market liquidity, excess liquidity is small or zero and the fragility-increasing channel dominates.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity--regime-dependence-is-documented"&gt;Q4. What heterogeneity / regime dependence is documented?&lt;/h3&gt;
&lt;p&gt;Results depend critically on the regime (Case I = no excess liquidity / fire sale; Case II = excess liquidity; Case III = excess liquidity, no fire sale, which never arises in equilibrium). The sign of dq-bar/dws flips between Case I (decreasing) and Case II (increasing). The intervention&amp;rsquo;s effect on fragility flips with market liquidity ws and long-term return R: low ws or low R → intervention reduces fragility; high ws or high R → intervention raises fragility; very high ws → externality vanishes (p = pR = 1) and intervention is neutral (q-bar = q-bar_R). The switch from Case I to Case II is governed by thresholds q_l (bank) and q_l,R (regulator), with q_l,R &amp;lt; q_l because the regulator internalizes the price and is more inclined to hold excess liquidity.&lt;/p&gt;
&lt;h3 id="q5-what-robustness--generality-checks-are-discussed"&gt;Q5. What robustness / generality checks are discussed?&lt;/h3&gt;
&lt;p&gt;Several modeling-assumption relaxations are argued not to change results qualitatively: (i) the assumption R* = R (giving p ≤ 1) can be generalized to allow p &amp;gt; 1, which does not undermine findings in the p &amp;lt; 1 range; (ii) partial runs can be generalized to multiple waves via a richer sunspot space without changing mechanisms; (iii) depositors not observing the bank&amp;rsquo;s portfolio can be replaced by observing it only after the withdrawal decision, with identical results; (iv) the simultaneous-move game is shown equivalent to a dynamic game in which the regulator moves first, as long as depositors cannot observe regulator choices; (v) the assumption that interventions convey no information to depositors can be relaxed (justified by the complexity of post-crisis regulation, e.g., the 848-page Dodd-Frank Act) without undermining the structure.&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q6. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;It builds on the fire sale externality literature (Lorenzoni 2008; Gale and Gottardi 2015; He and Kondor 2016; Davila and Korinek 2018 on over/under-investment; Acharya et al. 2011 and Gale and Yorulmazer 2020 on distorted portfolios; Perotti and Suarez 2011, Walther 2016, Kara and Ozsoy 2019 on optimal capital/liquidity regulation). It also builds on the bank-run literature (Bryant 1980; Diamond-Dybvig 1983) and on general-equilibrium / endogenous-portfolio extensions (Allen-Gale 2004; Farhi et al. 2009; Eisenbach-Phelan 2021; Cooper-Ross 1998; Ennis-Keister 2006; Li 2017). The stated novel contribution is being the first to show that policies designed to correct fire sale externalities can worsen financial fragility, achieved by jointly endogenizing the portfolio choice, the general-equilibrium asset price, and the equilibrium probability of a run.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q7. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Macroprudential interventions that regulate short-term liabilities and portfolio choice to curb fire sale externalities can increase the equilibrium probability of runs. The scope condition is market liquidity: the harmful trade-off (mitigate externality but raise fragility, and sometimes lower welfare) arises specifically when market liquidity ws is high (and/or R high); when ws is low, the regulator&amp;rsquo;s optimal excess-liquidity holding lets intervention both mitigate the externality and reduce fragility. A central caveat is that the regulator takes q as given and so does not perceive that its policy raises q-bar; the prescriptive takeaway is that policymakers must internalize q-bar (the endogenous run probability) when designing such policies, balancing externality mitigation against fragility.&lt;/p&gt;
&lt;h3 id="q8-are-the-quantitative-results-exact-magnitudes-or-signs"&gt;Q8. Are the quantitative results exact magnitudes or signs?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s results are predominantly signs and ordinal comparisons (e.g., x ≥ xR, p* ≤ pR*, q-bar_R ≥ q-bar, monotonicity in ws and p) plus closed-form threshold expressions (q_l, p_l, p_u, the four ws thresholds in Proposition 4) given in the text and appendices. Specific numeric magnitudes appear only as illustrative figure values (e.g., the example in Figure 9 where intervention raises fragility when ws is near 0.2); the paper does not report calibrated point estimates beyond such illustrative figures.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Fire sale externality&lt;/strong&gt;: In this model, the inefficiency arising because each competitive bank takes the t=1 asset price p as given and does not internalize that its long-term holdings and crisis-time asset sales depress p, harming other banks. It leads banks to over-invest in long-term assets and sell more than the efficient amount, pushing the equilibrium price below its efficient level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cash-in-the-market pricing&lt;/strong&gt;: The price of long-term assets at t=1 is set by the limited cash (endowment ws) that risk-neutral investors bring to the market rather than by fundamental value; when banks must sell, scarce market liquidity forces the price down (p ≤ 1 under the R*=R assumption).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Financial fragility (q-bar)&lt;/strong&gt;: Measured as q-bar, the maximum run probability q for which the partial-run strategy profile is part of an equilibrium, i.e., the largest q satisfying the run condition c1 ≥ c2β. Higher q-bar means the banking system is more fragile.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Excess liquidity&lt;/strong&gt;: Short-term asset holdings beyond what is needed to pay the first π withdrawals (πc1 &amp;lt; 1−x; Case II). It is a precautionary buffer that supplies resources for crisis payments c1β, reduces asset supply, and raises the fire sale price; the regulator holds more of it than the bank when market liquidity is low.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Case I vs Case II vs Case III&lt;/strong&gt;: Regimes of the bank&amp;rsquo;s best response: Case I = no excess liquidity, fire sale occurs (small q, high ws); Case II = excess liquidity held with fire sale (large q, low ws); Case III = excess liquidity so large that no fire sale occurs — shown never to be an equilibrium because it implies c2β &amp;gt; c2α &amp;gt; c1 (no run condition).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Regulator/intervention&lt;/strong&gt;: A planner that chooses (x, c1) internalizing the effect of these choices on the asset price p, while the bank still chooses (c2α, c1β, c2β) taking p as given and the regulator cannot direct depositors&amp;rsquo; withdrawal decisions; it represents the two policy instruments of regulating short-term liabilities and portfolio choice.&lt;/p&gt;</description></item><item><title>Heterogeneity in Manufacturing Growth Risk</title><link>https://macropaperwarehouse.com/papers/heterogeneity-in-manufacturing-growth-risk/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/heterogeneity-in-manufacturing-growth-risk/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; Since the Great Recession, quantifying downside risks to economic activity (rather than only expected outcomes) has become central for policymakers and investors. A large &amp;ldquo;growth-at-risk&amp;rdquo; literature documents that tightening financial conditions sharply raise downside risks to aggregate output while leaving upside potential roughly unchanged (Adrian, Boyarchenko and Giannone, 2019). This paper argues that the aggregate focus misses important structure: aggregate fluctuations can originate from industry-specific shocks, and recessions sharply raise cross-industry dispersion in growth (Bloom, 2014). The authors ask how downside output-growth risk from tight financial conditions differs across U.S. manufacturing industries, and which industry characteristics explain that heterogeneity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and method.&lt;/strong&gt; They use monthly industrial production (IP) growth for 74 U.S. manufacturing industries at the four-digit NAICS level over January 1973–July 2020 (Federal Reserve G.17; same industry selection as Chang and Hwang, 2015), and the Chicago Fed&amp;rsquo;s National Financial Conditions Index (NFCI) as the financial-conditions gauge. The method is a two-level (multi-level) quantile regression. Level 1 (following Adrian et al., 2019) regresses the τ-th quantile of average h-month-ahead IP growth on the current NFCI and current IP growth, industry by industry, focusing on h=3. Level 2 (inspired by Petersen and Strongin, 1996) regresses the estimated level-1 NFCI quantile coefficients cross-sectionally on standardized, time-invariant industry characteristics (capital, materials, energy, production-labor and overhead-labor intensities; a correlation-based labor-hoarding measure; four-firm concentration ratio; industry size measured by value-added share; and a durability dummy). Inference uses a stationary bootstrap (1,000 replications) that propagates level-1 estimation uncertainty into level 2. Industries split into 45 durables and 29 nondurables.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main quantitative findings.&lt;/strong&gt; Deteriorating financial conditions hit downside risk far harder than the center or upside of the growth distribution. On average across industries, a one-standard-deviation positive NFCI shock lowers three-month-ahead IP growth by 0.237% at the median and 0.773% at the 5% quantile, and raises the 95% quantile by 0.042%. The average 5% NFCI coefficient is -0.77 across all industries versus -0.31 (linear) and -0.24 (median); 47 of 74 industries (63.5%) have significant 5% coefficients, only 5 (6.8%) have significant 95% coefficients. Durables are about twice as sensitive in the left tail: average 5% coefficients are -0.96 (durables) versus -0.48 (nondurables), with 75.6% of durables versus 44.8% of nondurables significant at 5%. Some industries (computer, aerospace, food, dairy) are essentially unaffected across the whole distribution. The relationship is nonlinear for 46 of 74 industries (62.2%) at the 5% quantile (77.8% of durables, 37.9% of nondurables). Galvao et al. (2018) slope-homogeneity tests reject coefficient equality across industries for lower quantiles. Subsample analysis (1973-84 / 1985-2006 / 2007-2020) shows tail effects strongest in the most recent period (average 5% coefficient -1.38 vs -0.73 and -0.49), weakest during the Great Moderation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Explaining heterogeneity / implications.&lt;/strong&gt; In the all-manufacturing second level, large industries and durable-goods producers have significantly more vulnerable downside growth, while capital-intensive, overhead-labor-intensive, and labor-hoarding industries are less vulnerable. Within durables, size, materials intensity (more vulnerable) and overhead labor intensity (less vulnerable) matter; within nondurables, energy intensity (more vulnerable) and labor hoarding (less vulnerable) matter. Implication: industry-targeted stabilization policy may be more effective than nationwide policy given the heterogeneity, and investors can build industry-rotation strategies less exposed to financial-market shocks.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-empiricalidentification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the empirical/identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The strategy is descriptive-predictive rather than causal. Level 1 estimates industry-specific quantile regressions of average h-month-ahead IP growth on the current NFCI and current IP growth (Koenker-Bassett check-function minimization via the Frisch-Newton interior-point algorithm). Level 2 regresses the estimated NFCI quantile coefficients on standardized industry characteristics via OLS. The key inferential innovation is a stationary bootstrap (Politis-Romano 1994; block length via Politis-White 2004 with Patton et al. 2009 correction, expected block ~36.76 set by the NFCI series) that jointly resamples industry IP and NFCI and feeds level-1 estimation uncertainty into level-2 confidence bands. Main threats: (i) the relationship is associational, not identified as causal — the NFCI is endogenous to the macroeconomy; (ii) generated-regressor problem in level 2 (coefficients are estimates), addressed by the bootstrap; (iii) small cross-sections (45 durables, 29 nondurables, even fewer at the three-digit level) reduce power to detect characteristic effects; (iv) time-invariant characteristics are averaged over varying available windows, abstracting from time variation.&lt;/p&gt;
&lt;h3 id="q2-how-is-nonlinearity-established-and-against-what-benchmark"&gt;Q2. How is nonlinearity established, and against what benchmark?&lt;/h3&gt;
&lt;p&gt;Quantile coefficients are compared to OLS linear coefficients (constant across quantiles) using 95% bootstrap bands generated under a null that the data-generating process is a VAR(4) for the NFCI and IP growth (the Adrian et al. 2019 approach). Quantile estimates falling outside those bands are evidence of nonlinearity. 46 of 74 industries (62.2%) have a 5% coefficient significantly different from OLS; the total manufacturing sector is also nonlinear, mirroring Adrian et al. (2019) for aggregate GDP.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Three layers. (1) Durables vs nondurables: durables roughly twice as sensitive in the left tail (avg 5% coefficient -0.96 vs -0.48). (2) Within sectors: e.g. motor vehicles, motor bodies and motor parts have significant 5% coefficients below -2; resin and fiber below -1.5; while computer, aerospace and food are insignificant/unaffected. (3) Across the distribution: strong effects at low quantiles, near-zero at high quantiles (avg 95% coefficient 0.04). Industries with large negative 5% coefficients also tend to have larger positive 95% coefficients (higher conditional volatility under tight conditions), most clearly iron, motor vehicles, fiber and resin — though upside gains are generally smaller than the downside increase.&lt;/p&gt;
&lt;h3 id="q4-which-industry-characteristics-explain-the-heterogeneity-and-in-which-direction"&gt;Q4. Which industry characteristics explain the heterogeneity, and in which direction?&lt;/h3&gt;
&lt;p&gt;All-manufacturing (74 industries): negative effects on lower-quantile NFCI coefficients (i.e. more downside vulnerability) from industry size and durability; positive effects (less vulnerability) from overhead labor intensity, labor hoarding, and capital intensity. Durables: significant negative effect of materials intensity, negative (small) effect of size, positive effect of overhead labor intensity; production labor intensity significant at some higher quantiles. Nondurables: significant negative effect of energy intensity, positive effect of labor hoarding. Energy intensity, production labor intensity and concentration ratio are NOT significant for total manufacturing or durables in the way Petersen-Strongin found for cyclicality.&lt;/p&gt;
&lt;h3 id="q5-what-economic-mechanisms-are-offered-for-each-characteristic-effect"&gt;Q5. What economic mechanisms are offered for each characteristic effect?&lt;/h3&gt;
&lt;p&gt;Size: mean reversion — an industry larger than average is more likely to see growth fall (Braun-Larrain 2005). Durability: durable production is inherently more cyclical (Petersen-Strongin 1996). Labor hoarding / overhead labor: firms retain trained (especially nonproduction) workers due to sunk hiring/training costs (Becker 1962; Oi 1962; Parsons 1986), lowering the incentive to cut production in downturns. Capital intensity: higher fixed-to-variable cost ratio reduces incentive to cut output, and tangible capital provides collateral easing financing (consistent with Braun-Larrain 2005). Materials intensity (durables): higher share of variable costs raises cyclicality; also links to the negative materials-intensity/TFP relation of Baptist-Hepburn (2013).&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;(i) Additional controls (Gilchrist-Zakrajsek variables: term spread, real federal funds rate, credit spread, excess bond premium, plus extra IP lags) — qualitatively similar, wider bands. (ii) Unobserved heterogeneity via Ando-Bai (2020) interactive-fixed-effects panel quantile model (one common factor optimal) — highly similar. (iii) Alternative NAICS disaggregation: three-digit (21 industries; capital intensity dropped for multicollinearity; only labor hoarding and durability significant) and six-digit (101 industries; more characteristics significant, including production labor intensity and concentration ratio). (iv) Longer horizons h=6 and h=12 — qualitatively similar but weaker/less significant as horizon lengthens. (v) Subsample analysis of both the growth-risk coefficients and the characteristic construction windows (1973-84, 1985-2006, 2007-2020; and start dates 1958/1973/1987) — effects relatively stable; size and labor-hoarding effects weaken in recent periods while overhead labor and durability stay significant.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-relate-to-and-differ-from-petersen-and-strongin-1996-and-adrian-et-al-2019"&gt;Q7. How does this relate to and differ from Petersen and Strongin (1996) and Adrian et al. (2019)?&lt;/h3&gt;
&lt;p&gt;It extends Adrian et al. (2019) from aggregate to industry-level growth-at-risk, documenting substantial cross-industry variation that is invisible at the aggregate level — to the authors&amp;rsquo; knowledge the first disaggregate growth-at-risk study. It extends Petersen-Strongin (1996), who used a linear cyclicality framework, by allowing a flexible/nonlinear quantile relationship specifically with financial conditions. Findings broadly echo Petersen-Strongin for downside risk (materials intensity most important in durables; labor hoarding for nondurables — their only significant nondurable effect), but deviate by NOT finding energy intensity, production labor intensity, or concentration ratio significant in durables, and by adding size and capital intensity (cf. Braun-Larrain 2005) as relevant for total manufacturing. The agreement is attributed to business and financial cycles being closely intertwined (Claessens et al. 2012).&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Because vulnerability is highly heterogeneous, industry-level stabilization policy may be more effective than nationwide policy (OECD 2003), and policies can be targeted using the signalling characteristics (size, durability, materials/energy intensity vs capital/overhead-labor intensity and labor hoarding). Investors can build industry-rotation strategies less exposed to financial shocks. Scope conditions: evidence is U.S. manufacturing only, associational not causal, conditional on the NFCI as the financial-conditions measure, strongest at the three-month horizon and in the post-2007 subsample, and characteristic effects rest on relatively small cross-sections.&lt;/p&gt;
&lt;h3 id="q9-are-there-caveats-the-authors-themselves-flag"&gt;Q9. Are there caveats the authors themselves flag?&lt;/h3&gt;
&lt;p&gt;Yes: after splitting into durables/nondurables, fewer characteristic effects are significant, which the authors attribute to smaller cross-sections rather than absence of effects; the two-level model is estimated sequentially (two-step) not simultaneously; characteristics are treated as time-invariant averages (justified by stable cross-industry rankings, though production labor intensity shows a downward trend); and upside potential, while present, is generally smaller than the increased downside risk.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Growth-at-risk / downside growth risk&lt;/strong&gt;: The lower-quantile (e.g. 5%) of the conditional distribution of future output growth given current conditions; here the 5% quantile of average three-month-ahead industry IP growth conditional on the NFCI, capturing how bad growth could plausibly get under tight financial conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Multi-level quantile regression&lt;/strong&gt;: The authors&amp;rsquo; two-step procedure: level 1 estimates industry-specific quantile regressions of future IP growth on the NFCI and current IP growth; level 2 regresses the estimated NFCI quantile coefficients cross-sectionally on industry characteristics, with a bootstrap carrying level-1 uncertainty into level-2 inference.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;NFCI (National Financial Conditions Index)&lt;/strong&gt;: Chicago Fed weekly index of U.S. money, debt, equity, and (shadow) banking conditions built from a large dynamic factor model; positive values mean tighter-than-average financial conditions, negative values looser-than-average. Averaged to monthly here.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor hoarding&lt;/strong&gt;: Retention of employees during downturns because of sunk search, hiring and training costs; measured here as the negative correlation between changes in materials usage and changes in production-worker hours (a value of -1 = no hoarding), so higher values indicate more hoarding and predict less cyclical, less vulnerable growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Overhead labor intensity&lt;/strong&gt;: Cost of nonproduction (overhead) labor relative to value added. Because nonproduction workers embody more firm-specific investment, they are more subject to labor hoarding, so overhead-labor-intensive industries have less vulnerable downside growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Durable vs nondurable goods sector&lt;/strong&gt;: Federal Reserve classification (45 durable, 29 nondurable industries here). Durable-goods production is more cyclical and, in this paper, about twice as sensitive in the left tail of the growth distribution to adverse financial conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Slope homogeneity test&lt;/strong&gt;: Galvao et al. (2018) Swamy-type and standardized Swamy-type tests for a quantile-regression fixed-effects panel, used to formally reject equality of NFCI quantile slopes across industries, especially at lower quantiles.&lt;/p&gt;</description></item><item><title>Information Transparency of Firm Financing</title><link>https://macropaperwarehouse.com/papers/information-transparency-of-firm-financing/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/information-transparency-of-firm-financing/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Noël and Sun build an information-based theory of capital structure designed to explain the diversity of observed firm financing behavior and the coexistence of distinct optimal financial contracts. The motivating puzzle is that real-world financing methods (external equity, corporate bonds/bank loans, business credit lines/cards) differ systematically in how much firm-specific information investors require — equity and rated debt are &amp;ldquo;transparent&amp;rdquo; with firm-specific terms, while credit lines have general qualification standards and common interest rates. The paper asks three questions: what drives a firm&amp;rsquo;s optimal financing choice, why do equity, transparent debt, and opaque debt coexist as optimal contracts, and what is a firm&amp;rsquo;s optimal debt-to-equity ratio.&lt;/p&gt;
&lt;p&gt;This is a pure theory paper (no data or sample period). The model has a continuum of ex-ante heterogeneous firms, each with internal funds n (support [0, ī]), productivity θ, and survival/success rate α, all i.i.d. With investment i, output is θ·min[i,ī] with probability α and 0 with probability 1−α. The model nests two information problems: (1) adverse selection over a firm&amp;rsquo;s quality (α, θ), which a costly verification technology can reveal at cost γ &amp;gt; 0; and (2) an ex-post agency problem, since a firm can hide output and auditing recovers only a fraction σ ∈ (0,1) of hidden output. Internal funds n are public. Firms choose among four options: opaque contract, separating contract, transparent contract, or self-funding. Investors are risk-neutral with outside storage return r &amp;gt; 0. Assumption 1 (αθ̲ &amp;gt; 1+r &amp;gt; σᾱθ̄) ensures all projects are worth investing and all firms prefer some external financing.&lt;/p&gt;
&lt;p&gt;Main results (proved as a unique perfect Bayesian equilibrium):&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Three contract types arise endogenously: equity (investors get a fraction of output / ownership, payout depends on θ), transparent debt (firm-specific interest rate (1+r)/α reflecting survival rate), and opaque debt (common interest rate (1+r)/αΩ). The transparent contract is implementable by either equity or transparent debt when n ≤ nT(αθ); only transparent debt when n &amp;gt; nT(αθ).&lt;/li&gt;
&lt;li&gt;The separating (signaling without costly verification) contract does NOT survive for any firm except possibly the lowest type (α̲, θ̲); even that type is strictly better off pooling on opaque debt.&lt;/li&gt;
&lt;li&gt;The unique equilibrium has θΩ = θ̲ and αΩ = E[α] (existence requires verification cost condition (26): γ/(σᾱθ̲ī) ≥ (1−σ)θ̲(ᾱ−E[α])/(1+r−σθ̲E[α])). It is either pooling on opaque debt or mixing (transparent + opaque), never pooling on transparent. There is a threshold cost γ̄ ∈ (0,∞) above which the transparent set is empty and the equilibrium becomes pooling.&lt;/li&gt;
&lt;li&gt;Firm characteristics drive choice: all firms with αθ ≤ θ̲·E[α] use opaque debt regardless of internal funds; transparent contracts require sufficiently high quality satisfying condition (27) AND intermediate internal funds. Firms with n ∈ [n1(α,θ), nT(αθ)] are indifferent between equity and transparent debt; those with n ∈ (nT(αθ), n2(α,θ)] strictly prefer transparent debt; very low or very high n firms use opaque debt.&lt;/li&gt;
&lt;li&gt;Partial capital structure irrelevance: only a strict subset of firms (those satisfying (27) with n ∈ [n1, nT(αθ)]) are indifferent between equity and transparent debt (a Modigliani-Miller equivalence within an asymmetric-information setting).&lt;/li&gt;
&lt;li&gt;Debt weakly dominates equity: debt implements the optimal contract for all firms; equity does so only for the strict subset above. The optimal debt-to-equity ratio is not a smooth function of internal funds and need not be unique (a continuum is optimal for indifferent firms). The theory reconciles the conflicting empirical evidence of Myers (2001) (equity issues minor, mostly debt, across broad U.S. firms) versus Frank and Goyal (2003) (equity significant, often exceeding investment, for publicly-traded firms).&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-model-environment-and-the-two-layers-of-information-frictions"&gt;Q1. What is the model environment and the two layers of information frictions?&lt;/h3&gt;
&lt;p&gt;A continuum of ex-ante heterogeneous firms, each with public internal funds n ∈ [0, ī] and private quality (α, θ): productivity θ and survival/success rate α. Output is θ·min[i, ī] with probability α and 0 otherwise. Friction 1 is adverse selection over (α, θ), resolvable only via a costly verification technology (cost γ &amp;gt; 0) used before contracting. Friction 2 is an ex-post agency/moral-hazard problem: a firm can hide actual output, and auditing recovers at most a fraction σ ∈ (0,1) of hidden output — so the contract must induce truthful reporting. Investors are risk-neutral with storage return r &amp;gt; 0.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-separating-signaling-contract-collapse-in-equilibrium"&gt;Q2. Why does the separating (signaling) contract collapse in equilibrium?&lt;/h3&gt;
&lt;p&gt;A separating contract must satisfy two incentive-compatibility constraints simultaneously: the financing firm&amp;rsquo;s own truthful-output-reporting constraint (identical to the transparent contract&amp;rsquo;s IC), AND a constraint that no other firm type wants to mimic it. Proposition 3 proves the first constraint makes the second impossible to uphold for all firms except possibly the lowest type (α̲, θ̲). Firms with lower expected quality but higher actual productivity (θ̃ ≥ θ) want to mimic at low funds; higher-risk firms (α̃ &amp;lt; α) want to mimic at high funds. Since any optimal separating contract is also an optimal transparent contract minus the cost γ, any firm that could separate would never use the costly transparent contract — but no firm can successfully separate. Even the lowest type prefers opaque debt (Proposition 7), so no separating contract is used in equilibrium.&lt;/p&gt;
&lt;h3 id="q3-why-is-the-opaque-contract-necessarily-debt-and-never-equity"&gt;Q3. Why is the opaque contract necessarily debt and never equity?&lt;/h3&gt;
&lt;p&gt;With opaque financing investors do not learn firm quality. A binding incentive-compatibility constraint reduces to zO = σθΩ·iO, and the participation constraint (which binds for all n &amp;lt; ī) gives payout zO = ((1+r)/αΩ)·(iO − n) — a fixed general interest rate (1+r)/αΩ on external funds. This is a debt contract. Equity is impossible because investors cannot be convinced to take ownership shares of output without firm quality being revealed to them. Opaque debt resembles a business line of credit: general qualification standards (Assumption 1) and a common interest rate reflecting E[α], independent of firm-specific information.&lt;/p&gt;
&lt;h3 id="q4-when-are-equity-and-transparent-debt-equivalent-and-what-distinguishes-the-information-each-reveals"&gt;Q4. When are equity and transparent debt equivalent, and what distinguishes the information each reveals?&lt;/h3&gt;
&lt;p&gt;For firms with n ≤ nT(αθ), both the firm&amp;rsquo;s IC constraint (2) and investors&amp;rsquo; participation constraint (3) bind. The optimal transparent contract is then implementable equivalently by equity (payout = a fraction of output, depends on θ) or transparent debt (firm-specific interest rate (1+r)/α, depends on α). This is a Modigliani-Miller-style equivalence obtained under asymmetric information. Conditional on survival, equity investors care about θ (commercial information — technology, product lines, outlook), while transparent-debt investors care about α (creditworthiness — financial condition), matching real-world distinctions between equity due diligence and credit-rating/bank scrutiny. The equivalence holds even if verifying α and θ costs differently, as long as both constraints bind.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-in-financing-behavior-does-the-model-generate-cross-section"&gt;Q5. What heterogeneity in financing behavior does the model generate (cross-section)?&lt;/h3&gt;
&lt;p&gt;Per Table 1 and Theorem 1: (a) Equity users have high quality (αθ), are lower-intermediate in internal funds (n ∈ [n1(α,θ), nT(αθ)]), reveal both α and θ, and have the highest financial leverage. (b) Transparent-debt users have high quality, intermediate funds, reveal α and θ, with firm-specific interest rate reflecting α. (c) Opaque-debt users span all quality types and all funds levels (often very low or very high funds), reveal only general information (E[α], θ̲), face a common interest rate, and have lower leverage. Better-quality but funds-constrained firms are most likely to use transparent financing; firms with αθ ≤ θ̲E[α] always use opaque debt regardless of funds, masking inferior quality by pooling.&lt;/p&gt;
&lt;h3 id="q6-what-dynamic-firm-financing-patterns-can-the-static-model-rationalize"&gt;Q6. What dynamic firm-financing patterns can the (static) model rationalize?&lt;/h3&gt;
&lt;p&gt;The authors interpret each capital-structure decision as a reaction to updated (n, α, θ). They reconcile: (1) startups using equity (high αθ, low n relative to capacity); (2) share buybacks (rising n moving a firm from the equity-indifference region into transparent-debt or opaque-debt regions); (3) small businesses starting with a credit line then adding equity/loans/bonds as n or quality rises into the transparent region; (4) firms issuing equity when prices are high (high price signals improved quality αθ, and funds raised via equity strictly increase in αθ); (5) firms using two or three financing types simultaneously, because the theory is per-project — different projects/purposes (e.g., main operations vs. routine liquidity) can optimally use transparent and opaque contracts at the same time.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-model-reconcile-the-myers-2001-vs-frank-goyal-2003-empirical-discrepancy"&gt;Q7. How does the model reconcile the Myers (2001) vs. Frank-Goyal (2003) empirical discrepancy?&lt;/h3&gt;
&lt;p&gt;Myers (2001) reports that for broad U.S. nonfarm/nonfinancial corporations, external finance is a small share (mostly under 20%) of capital formation with equity issues minor and the bulk being debt. Frank and Goyal (2003) find that for publicly-traded U.S. firms (excluding financials, regulated utilities, major-merger firms), external finance is large (often exceeding investment) and net equity issues commonly exceed net debt issues. The theory explains both: equity finance is optimal only for high-quality, intermediate-funds firms, and amounts raised increase in quality, so publicly-traded (high-quality) samples show large, equity-heavy external finance, while broader samples include many debt-only and self-funded firms, yielding smaller, debt-dominated external finance. Verification cost γ varying over time, industry, and country also generates cross-dataset behavioral differences.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-structure-of-the-optimal-debt-to-equity-ratio"&gt;Q8. What is the structure of the optimal debt-to-equity ratio?&lt;/h3&gt;
&lt;p&gt;Proposition 10: it varies with firm characteristics and is not a smooth function of internal funds, and may not be unique. In a pooling equilibrium it equals σθ̲E[α]/(1+r−σθ̲E[α]) for n ≤ nO (constant across quality) and ī/n − 1 (strictly decreasing) for n &amp;gt; nO. In a mixing equilibrium, firms not satisfying (27) follow the same formula; firms satisfying (27) traverse: the constant ratio for n &amp;lt; n1; a continuum [0, σαθ/(1+r−σαθ)] over the equity/transparent-debt indifference region n ∈ [n1, nT(αθ)]; then the constant ratio; then ī/n − 1. The non-uniqueness over the indifference region is precisely the &amp;lsquo;partial capital structure irrelevance.&amp;rsquo;&lt;/p&gt;
&lt;h3 id="q9-how-does-the-equilibrium-switch-between-mixing-and-pooling"&gt;Q9. How does the equilibrium switch between mixing and pooling?&lt;/h3&gt;
&lt;p&gt;Theorem 1(iv): all else equal, as the verification cost γ rises, the set of transparent-contract users shrinks and opaque-debt users expand. There is a threshold γ̄ ∈ (0,∞) above which no firm uses transparent financing, so the equilibrium is pooling on opaque debt; below it, the equilibrium is mixing. Existence of the unique PBE itself requires condition (26), ensuring γ relative to the tightest discipline σᾱθ̲ī is sufficiently high so that all firms with productivity θ̲ (any α) choose opaque debt, pinning down θΩ = θ̲ and αΩ = E[α].&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-differ-from-prior-optimal-contracting-and-capital-structure-literature"&gt;Q10. How does this paper differ from prior optimal-contracting and capital-structure literature?&lt;/h3&gt;
&lt;p&gt;Prior costly-state-verification models (Diamond 1984; Gale-Hellwig 1985; Williamson 1986) yield debt as optimal with homogeneous entrepreneurs; adverse-selection models (Leland-Pyle 1977; Stiglitz-Weiss 1981; Myers-Majluf 1984 and others) and agency models (Jensen-Meckling 1976; DeMarzo-Sannikov 2006; DeMarzo-Fishman 2007) treat the frictions separately. This paper&amp;rsquo;s novelty is nesting BOTH adverse selection and the agency problem in a model of heterogeneous firms (along quality AND internal funds). That combination is what makes signaling/separating contracts fail and forces costly verification (transparency) for adverse-selection resolution, and it generates the coexistence of equity, transparent debt, and opaque debt, lends theoretical support to the pecking-order hypothesis (debt weakly dominates equity), and yields partial — not full — Modigliani-Miller irrelevance. It also contributes to the literature on optimal information control (Hirshleifer 1971, 1972; Diamond 1985; Dang-Gorton-Holmström-Ordoñez 2017; Monnet-Quintin 2017) by endogenizing the information-disclosure decision within contract design.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-key-scope-conditions-and-caveats"&gt;Q11. What are the key scope conditions and caveats?&lt;/h3&gt;
&lt;p&gt;Results hold under Assumption 1 (all projects worth investing; all firms prefer external financing — so &amp;rsquo;lowest quality&amp;rsquo; is not literally any inferior business). The model is static and per-project; &amp;rsquo;low n&amp;rsquo; means low funds relative to project capacity ī, not necessarily a small or young firm. The most severe misreporting penalty (recovering fraction σ) is imposed to make incentive compatibility least costly. ī can be made to vary across projects without changing main results. The verification cost γ is the central comparative-statics parameter governing whether the equilibrium is mixing or pooling. Equilibrium existence requires condition (26) on γ. There is no empirical estimation — quantitative claims are model-derived equilibrium objects, not data estimates.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Information transparency&lt;/strong&gt;: Defined in the paper as whether investors require business information considered confidential to the firm to aid their investment decisions. Equity and transparent debt are &amp;rsquo;transparent&amp;rsquo; because the firm pays cost γ to reveal its true (α, θ); opaque debt merely reflects general information about the pool of qualifying firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Opaque debt&lt;/strong&gt;: A pooling debt contract carrying a common interest rate (1+r)/αΩ independent of firm-specific information, reflecting the lowest productivity θΩ and the expected survival rate αΩ = E[α] of all qualifying firms. Resembles a real-world business line of credit; the only contract implementable for firms needing small external funds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Transparent debt&lt;/strong&gt;: A debt contract whose firm-specific interest rate (1+r)/α reflects the firm&amp;rsquo;s verified survival rate α (creditworthiness). Resembles corporate bonds or bank loans with firm-specific rates set after credit-rating-style scrutiny.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Transparent (equity) contract&lt;/strong&gt;: The optimal transparent contract implemented as equity: investors receive a fraction of actual output (ownership), with payout depending on productivity θ. Available only to high-quality firms with lower-intermediate internal funds (n ∈ [n1, nT(αθ)]); these firms are indifferent between equity and transparent debt.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Separating contract&lt;/strong&gt;: A contract by which a firm signals its true quality (α, θ) WITHOUT paying the verification cost γ, designed so no other type mimics it. Proved not to survive in equilibrium for any firm except possibly the lowest type, which itself prefers opaque debt.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Partial capital structure irrelevance&lt;/strong&gt;: A Modigliani-Miller-style equivalence holding only for a strict subset of firms — those satisfying condition (27) with n ∈ [n1(α,θ), nT(αθ)] — who are indifferent between equity and transparent debt. Outside this subset the financing choice is determinate, so irrelevance is &amp;lsquo;partial,&amp;rsquo; not universal.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Verification cost γ&lt;/strong&gt;: The cost of the technology (e.g., a rating agency, or the firm&amp;rsquo;s own effort to convince investors) that ascertains true firm quality (α, θ) before contracting. Its level governs whether the equilibrium is mixing (low γ) or pooling on opaque debt (γ above threshold γ̄), and existence of the unique PBE requires γ sufficiently high relative to σᾱθ̲ī (condition 26).&lt;/p&gt;</description></item><item><title>Interest Rate Pegs and the Reversal Puzzle: On the Role of Anticipation</title><link>https://macropaperwarehouse.com/papers/interest-rate-pegs-and-the-reversal-puzzle-on-the-role-of-anticipation/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/interest-rate-pegs-and-the-reversal-puzzle-on-the-role-of-anticipation/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper revisits the &amp;ldquo;reversal puzzle&amp;rdquo; — the counterintuitive result, first documented by Carlstrom, Fuerst and Paustian (CFP, 2015), that in standard New Keynesian models the effect of forward guidance (technically implemented as a perfectly anticipated interest rate peg) can switch from expansionary to contractionary as the duration of the peg increases. The authors&amp;rsquo; central claim is that the appearance of the puzzle hinges on agents&amp;rsquo; degree of anticipation of the peg, and they examine three polar/intermediate cases: perfect anticipation, no anticipation, and imperfect anticipation.&lt;/p&gt;
&lt;p&gt;Model and setup: The laboratory is the medium-scale DSGE model of Carlstrom, Fuerst and Paustian (2017), which features funding constraints and market segmentation (only financial intermediaries can hold long-term public and private bonds, subject to a leverage constraint from a hold-up problem and net-worth adjustment costs; households face a loan-in-advance constraint on investment). These frictions break Wallace neutrality so that QE has real and inflationary effects. The model has standard New Keynesian features: habit consumption, monopolistic competition, Erceg-Henderson-Levin (2000) sticky prices and wages with Christiano-Eichenbaum-Evans (2005) indexation, investment adjustment costs, and a Taylor rule with interest-rate smoothing. It is estimated with Bayesian methods on eight euro-area observables over 1998Q1-2013Q4, with a subset of parameters calibrated to CFP (β=0.99, capital share α=0.33, depreciation δ=0.025, price/wage markup elasticities ε_p=ε_w=5, steady-state leverage 6). The initial impulse in all experiments is the launch of a QE programme, modeled as a single shock to an AR(2) process for the real market value of long-term bonds (purchases last 6 quarters). Without a peg, QE raises inflation (the orthodox result).&lt;/p&gt;
&lt;p&gt;Main findings: (1) Perfect anticipation (perfect-foresight solution): reversals are a robust phenomenon. As peg duration P rises, the inflation response first grows and then explodes near a critical value; in the baseline this critical value is eight quarters. For P of 9-14 quarters inflation reverses sign (deflation instead of inflation); for 15-23 quarters the sign flips back to positive; for 24-50 quarters it turns negative again. Thus output and inflation responses oscillate with P. The authors give analytical intuition via the forward solution: complex unstable eigenvalues of matrix J, written in polar form, mean powers of J enter the solution as trigonometric functions of P (de Moivre&amp;rsquo;s formula), producing the oscillation. (2) No anticipation (extended-path method, agents expect E_t[ε_{t+n}]=0 each period and are &amp;ldquo;surprised&amp;rdquo;): the reversal puzzle is absent for all durations 0-50; the initial inflation response is always positive, because powers of J no longer enter the solution. (3) Imperfect anticipation (Markov-switching model solved with Maih&amp;rsquo;s 2015 RISE toolbox): two regimes — Taylor rule (regime 1) vs. peg (regime 2, where ρ=τ_Π=τ_y=0). Agents know transition probabilities, so the frequency F2 and average duration AD2 of the peg are known; frequency is interpreted as the degree of anticipation. Generalized impulse responses (50,000 draws) for average durations of 4, 11.5, 19, 37, 50 quarters and frequencies of 10%, 15%, 20%, 30%, 40%, 50% show: at the empirically relevant frequency of 10% (post-WWII US ZLB experience, ~7 years in 73) and at 15% and 20%, no reversals occur for any average duration. Reversals appear only at implausibly high frequencies: at 30% only for AD2=4 quarters; at 40% for AD2=4, 11.5, 19 quarters; at 50% for all average durations.&lt;/p&gt;
&lt;p&gt;Implications: A Markov-switching treatment of pegs/ZLB delivers more plausible model outcomes than perfect foresight and is a promising tool for policy simulations to avoid the reversal pathology, since under realistic anticipation forward guidance is less powerful and reversals do not arise.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-exactly-is-the-reversal-puzzle-and-where-did-it-originate"&gt;Q1. What exactly is the reversal puzzle and where did it originate?&lt;/h3&gt;
&lt;p&gt;It is the counterintuitive result that the macroeconomic effect of forward guidance — implemented technically as a perfectly anticipated interest rate peg — can switch from expansionary to contractionary depending on the peg&amp;rsquo;s duration, producing sizeable deflation instead of inflation. Carlstrom, Fuerst and Paustian (2015) first analyzed and named it. Similar sign reversals are noted in Lindé-Smets-Wouters (2016) and Binning-Maih (2017).&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identificationsolution-strategy-for-each-anticipation-case-and-what-distinguishes-them"&gt;Q2. What is the identification/solution strategy for each anticipation case, and what distinguishes them?&lt;/h3&gt;
&lt;p&gt;Perfect anticipation: perfect-foresight (deterministic) solution where the peg is implemented via binary dummy shocks (ε^TR in {0,1}) set to one for P pre-announced quarters; agents know all future ε_{t+n}, so powers of the eigenvalue matrix J enter the forward solution. No anticipation: the extended-path method, running a deterministic simulation each period with the previous period as initial condition and steady state as terminal condition, imposing E_t(ε_{t+n})=0 — agents are surprised the peg continues, so powers of J drop out. Imperfect anticipation: a Markov-switching framework (Maih 2015) with non-zero transition probabilities between a Taylor-rule regime and a peg regime; the peg is a recurring stochastic event whose frequency and average duration are known to agents.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-formal-mechanism-for-the-oscillation-under-perfect-foresight"&gt;Q3. What is the formal mechanism for the oscillation under perfect foresight?&lt;/h3&gt;
&lt;p&gt;The forward-looking (explosive) variables solve as w2,t = -E_t{Σ J^{n-1} Ω22^{-1} Q2 Φ ε_{t+n}}. Some diagonal elements of J (the unstable generalized eigenvalues) are complex; in polar form z_jj = r(cos φ + i sin φ), and by de Moivre z_jj^k = r^k(cos kφ + i sin kφ) for k=0,&amp;hellip;,P-1. Because nonzero anticipated future shocks bring in powers of J, the solution involves trigonometric functions of the peg length P, so simulations approach an asymptote, switch sign, approach another asymptote, switch again — hence oscillation as P grows.&lt;/p&gt;
&lt;h3 id="q4-why-are-reversals-absent-under-no-anticipation-given-the-same-complex-eigenvalues"&gt;Q4. Why are reversals absent under no anticipation, given the same complex eigenvalues?&lt;/h3&gt;
&lt;p&gt;Complex eigenvalues are only a necessary, not sufficient, condition. Under no anticipation E_t(ε_{t+n})=0, so the solution for w2,t no longer depends on powers of J; the simulations do not &amp;lsquo;move along&amp;rsquo; the trigonometric functions, so the explosive complex eigenvalues cannot induce cyclical/explosive effects. A sufficient degree of anticipation is necessary for reversals to occur.&lt;/p&gt;
&lt;h3 id="q5-how-are-frequency-and-average-duration-of-the-peg-pinned-down-in-the-markov-switching-model"&gt;Q5. How are frequency and average duration of the peg pinned down in the Markov-switching model?&lt;/h3&gt;
&lt;p&gt;p12 is the transition probability from Taylor regime (1) to peg regime (2); p21 from 2 to 1. Average peg duration AD2 = 1/p21. Frequency F2 = AD2/(AD1+AD2) with AD1 = 1/p12. Table 2 maps the (AD2, F2) grid to the implied p12, p21. The authors check the mean-square-stability condition for each calibration before computing generalized impulse responses from 50,000 draws.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-empirically-relevant-peg-frequency-and-how-is-it-justified"&gt;Q6. What is the empirically relevant peg frequency and how is it justified?&lt;/h3&gt;
&lt;p&gt;About 10%, based on the post-WWII US zero-lower-bound experience (7 years at the ZLB out of 73 years), the same value used by Dordal-i-Carreras, Coibion, Gorodnichenko and Wieland (2016). The paper stresses that even at double this value (20%) reversals are absent for all average durations considered.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-reversal-pattern-under-imperfect-anticipation-differ-from-perfect-anticipation"&gt;Q7. How does the reversal pattern under imperfect anticipation differ from perfect anticipation?&lt;/h3&gt;
&lt;p&gt;The patterns differ. Under perfect foresight the lowest sub-range of durations (0-8 quarters) shows no reversal, whereas under imperfect anticipation at frequencies of 30% and 40% a reversal occurs for the lowest average duration (4 quarters). Reversals also appear &amp;lsquo;grouped&amp;rsquo; across adjacent average durations. The regime-specific IRFs explain this: given the peg regime (regime 2), higher average durations lead to reversals at low frequencies; given the no-peg regime (regime 1), only frequencies of 30%+ permit reversals and there lower average durations reverse. The GIRF blends both regimes, so its resemblance to a regime&amp;rsquo;s IRF depends on how frequently that regime occurs.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-are-performed"&gt;Q8. What robustness checks are performed?&lt;/h3&gt;
&lt;p&gt;An extensive grid search (Appendix D) varies each structural parameter one at a time around benchmark values under perfect foresight. Reducing forward-lookingness (lower β) or raising habit, changing depreciation δ or investment adjustment cost ψi, varying the Calvo price/wage parameters (θp, θw) and indexation (ιp, ιw), and varying Taylor-rule coefficients (ρ, τπ, τy) all only change the peg duration required for the reversal to appear, not its existence. Notably, even shutting down price and wage indexation jointly (ιp=ιw=0) does not eliminate reversals in this medium-scale model, because other endogenous state variables (capital, wages, net worth) generate complex eigenvalues. More aggressive inflation stabilization (higher τπ) or longer Calvo durations (&amp;gt;0.9) require a longer peg before reversal appears.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q9. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It is complementary to CFP (2015), who showed reversals require complex eigenvalues from endogenous states and that switching from sticky-price to sticky-information removes the puzzle; this paper instead goes beyond perfect foresight to show the degree of anticipation is key. It differs from De Graeve-Ilbas-Wouters (2014), Maliar-Taylor (2019), and Bundick-Smith (2020), who rely on realistic calibration to weaken forward guidance; here the resolution comes from realistic modeling of expectations. Unlike de Groot and Mazelis (2020) — who modify the linearized solution so agents are fully aware of the peg — the Markov-switching approach treats the peg as a recurring stochastic event. Methodologically closest is Chen (2017), who compares perfect-foresight and Markov-switching implementations of the ZLB; consistent with her, the authors find Markov-switching delivers more plausible outcomes.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Because the ZLB and forward guidance must be accounted for in model simulations, and these are often modeled as interest-rate pegs, policy evaluations risk spurious reversals. The Markov-switching approach circumvents this pathology and yields qualitatively plausible outcomes. Scope conditions: the result holds for empirically relevant peg frequencies (up to ~20%, double the 10% benchmark) across average durations of 4-50 quarters; reversals can still arise but only under extreme, arguably implausible frequencies (30%+). The conclusions are derived within the CFP (2017) segmented-markets model estimated on euro-area data, with QE as the initiating impulse.&lt;/p&gt;
&lt;h3 id="q11-how-is-the-qe-programme-modeled-and-what-is-its-transmission"&gt;Q11. How is the QE programme modeled and what is its transmission?&lt;/h3&gt;
&lt;p&gt;QE is a single shock to a persistent AR(2) process for the real market value of long-term bonds held by the public, generating an inverse hump shape with purchases lasting 6 quarters before gradual return to steady state. Transmission: lower bond supply to FIs raises bond prices and lowers yield-to-maturity and the term premium; FI net worth and leverage fall but net-worth mobility is limited by adjustment costs, so FIs raise demand for (perfect-substitute) investment bonds, raising their price, relaxing households&amp;rsquo; loan-in-advance constraint, boosting investment, output, and inflation; monetary policy then raises the policy rate under the Taylor rule.&lt;/p&gt;
&lt;h3 id="q12-are-there-caveats-about-the-no-anticipation-case-as-a-solution"&gt;Q12. Are there caveats about the no-anticipation case as a &amp;lsquo;solution&amp;rsquo;?&lt;/h3&gt;
&lt;p&gt;Yes. The authors state the no-anticipation case is obviously not a suitable solution to the puzzle — it is an unrealistic polar case (agents are surprised every period). Both polar cases (perfect and no anticipation) are unrealistic, which motivates the imperfect-anticipation Markov-switching analysis as the realistic middle ground.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Reversal puzzle&lt;/strong&gt;: The counterintuitive switching of forward guidance&amp;rsquo;s effect from expansionary to contractionary (deflation rather than inflation) as the duration of a perfectly anticipated interest rate peg increases; in this paper, the inflation response oscillates in sign across peg durations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Degree of anticipation&lt;/strong&gt;: The extent to which agents expect a future interest rate peg. The paper&amp;rsquo;s central organizing concept: in the stochastic case it is operationalized by the frequency of the peg regime, since a higher frequency makes agents consider a peg more likely.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Interest rate peg&lt;/strong&gt;: A regime in which the central bank abandons the Taylor rule and holds the nominal short-term rate fixed for a period — the technical implementation of forward guidance and the ZLB in this analysis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Imperfect anticipation (Markov-switching implementation)&lt;/strong&gt;: A scenario where agents attach non-zero transition probabilities to entering and exiting a recurring peg regime, so individual peg episodes are stochastic in occurrence and duration but their frequency and average duration are known.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Frequency of the peg (F2)&lt;/strong&gt;: The long-run share of time the economy spends in the peg regime, F2 = AD2/(AD1+AD2); interpreted as the degree of anticipation, with ~10% taken as the empirically relevant post-WWII US ZLB value.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Complex eigenvalues / forward solution&lt;/strong&gt;: Unstable generalized eigenvalues of the solution matrix J that are complex-valued; their polar-form powers introduce trigonometric functions of peg length P into the forward solution — a necessary but not sufficient condition for reversals, which require sufficient anticipation to activate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wallace neutrality breakdown&lt;/strong&gt;: The property, induced by FI funding constraints and bond-market segmentation in the CFP (2017) model, that asset purchases (QE) affect real activity and inflation rather than being neutral as in the standard New Keynesian model.&lt;/p&gt;</description></item><item><title>Macroprudential Policy in the Euro Area</title><link>https://macropaperwarehouse.com/papers/macroprudential-policy-in-the-euro-area/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/macroprudential-policy-in-the-euro-area/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation. There is now broad consensus that monetary authorities should hold a financial-stability mandate and that macroprudential policy should be part of it, yet evidence on the macroeconomic effectiveness of these policies and their interaction with monetary policy remains thin and inconclusive. The paper addresses this gap for the euro area, a case of special interest because of its international structure and because, within the short life of the euro, member states experienced major episodes of financial instability (the great financial crisis, GFC, and the sovereign debt crisis). The contribution is twofold: (1) build a novel aggregate index of the euro-area macroprudential policy stance and document its stylized facts since 1999; (2) be the first to identify, within a structural econometric framework, both unanticipated (surprise) and anticipated (news) exogenous macroprudential policy shocks and trace their macroeconomic effects.&lt;/p&gt;
&lt;p&gt;Data and method. The authors use MaPPED (Macro-Prudential Policies Evaluation Database), built by ECB staff and national central banks. For euro-area countries it records 1205 policy actions between 1995 and 2019 across 11 instrument types (capital buffers, lending standards, maturity mismatch tools, limits on credit growth, exposure limits, liquidity rules, loan loss provisions, minimum capital requirements and risk weights, leverage ratio, and &amp;lsquo;other measures&amp;rsquo;). Actions are signed (+ tightening, − loosening, 0 ambiguous) and weighted following Meuleman and Vander Vennet (2020): activation 1, change in level 0.25, change in scope 0.10, maintaining level/scope 0.05; deactivation resets the cumulative index to zero. This yields around 470 instrument-level indices, summed within each country and then aggregated across countries using GDP-share weights to form the EAMPP index. The empirical model is a seven-variable Bayesian SVAR at quarterly frequency over 1999:Q1–2019:Q2, estimated in levels with 4 lags and a Minnesota prior using the hyperparameters of Kurmann and Otrok (2013). Variables: the narrative EAMPP (which excludes countercyclical/financial-cycle-reactive policies so it is exogenous in the Romer-Romer sense), total credit to the private non-financial sector, real GDP, core CPI, inflation expectations (ZEW 6-month survey), VSTOXX, and a monetary policy rate (EONIA 1999–2009, Wu-Xia shadow rate thereafter). The surprise shock is identified by a Cholesky ordering with EAMPP first; the news shock is identified via the Barsky-Sims (2011) forecast-error-variance maximization (horizon k=0 to k=24), orthogonal to the surprise shock and not affecting EAMPP contemporaneously.&lt;/p&gt;
&lt;p&gt;Main findings. Stylized facts: EAMPP shows a positive starting value (policies predating the euro), a small positive trend up to the GFC, a loosening on average at the start of the GFC in 2009, then a clear upward (tightening) trend over the following seven years driven by sovereign-debt-crisis concerns and Basel III/CRR-CRDIV; the level in 2016 is almost twice as tightening as pre-crisis. The largest quarterly EAMPP change occurred in 2013:Q3 (CRR/CRDIV announcements). Policy announcements averaged about 13 per quarter in 1999–2015 versus about 2 per quarter in 2016–2019. Macroprudential and monetary policy moved oppositely; their correlation is about −0.90, negative and significant. SVAR results: a tightening surprise shock persistently raises the policy index, lowers total credit (on impact, accentuating over the medium term), reduces output in a way negatively correlated with credit (lowering credit pro-cyclicality), and lowers VSTOXX over the medium term after an initial rise. The effect on core CPI is negligible and on inflation expectations insignificant, so no price-stability trade-off; the monetary policy rate declines (accommodative complement). The news shock produces a gradual, persistent tightening, reduces credit, lowers credit pro-cyclicality, has muted effect on VSTOXX, and an insignificant price effect; the policy rate first rises then turns negative over the medium term. FEV decomposition: the two shocks combine to explain about half of credit variability after 24 quarters; neither shock exceeds 12% of core-CPI forecast variance and combined they never exceed 15% of prices. News shocks explain about 20% of credit forecast variance within the first quarter. Granger-causality and serial-correlation tests support exogeneity of both shocks.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Two shocks driving non-systematic macroprudential variation are identified within a seven-variable Bayesian SVAR (1999:Q1–2019:Q2, 4 lags, Minnesota prior). The surprise (unanticipated) shock is identified by a Cholesky decomposition with EAMPP ordered first, so it can affect EAMPP contemporaneously. The news (anticipated) shock uses the Barsky-Sims (2011) forecast-error-variance maximization: it is the orthonormal column that maximizes the cumulated forecast error variance of EAMPP over horizons k=0 to k=24, subject to not affecting EAMPP contemporaneously and being orthogonal to the surprise shock. A key prior step is constructing a narrative EAMPP that drops all policies with a countercyclical design (those reacting to the financial cycle), making the remaining index exogenous in the Romer-Romer (2010) sense. The main threats are: foresight/anticipation contaminating shock identification (addressed by using announcement rather than enforcement dates and by identifying news shocks); reverse causality and contemporaneous effects that plague recursive/GMM panel approaches; and informational insufficiency (whether the series are genuine shocks), which the authors test via Granger causality against forward-looking credit-standard surveys and serial-correlation tests.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The mechanism is that a tightening macroprudential stance curbs total credit to the private non-financial sector, which is the most robust predictor of financial crises, thereby moderating systemic risk and the build-up of excess credit during booms. Crucially, output responds in a way negatively correlated with credit, so the policy lowers the pro-cyclicality of credit (the key financial-stability gain). Surprise and news shocks are distinguished by their dynamics and by the FEV decomposition: news shocks dominate at short horizons (agents react quickly to signals, ~20% of credit forecast variance in the first quarter), while surprise shocks build gradually to a comparable share at medium-to-long horizons. The monetary-policy interaction is read off the policy-rate response: it moves accommodatively (declines) after a surprise tightening, complementing macroprudential policy without a price trade-off.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-or-differences-across-shock-types-are-documented"&gt;Q3. What heterogeneity or differences across shock types are documented?&lt;/h3&gt;
&lt;p&gt;The two shock types differ. The surprise shock causes an immediate credit drop that accentuates over the medium term and an accommodative (declining) monetary policy rate; VSTOXX first rises then falls below baseline. The news shock causes a gradual, persistent policy tightening, a credit decline that moderates before dropping again over the medium term, a muted VSTOXX response, and a monetary policy rate that first increases (complementing the tightening and reflecting a small initial price rise) then turns negative over the medium term. Core prices show a small initial increase under the news shock before declining, whereas the surprise shock barely affects core CPI. Both shocks ultimately lower credit pro-cyclicality and have insignificant effects on price stability.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Several. (1) Alternative macroprudential target variables replacing total credit: a systemic-risk index (CISS) — results barely change; bank credit — results similar, with a more pronounced decline in bank credit; household credit — results similar but the household-credit decline is stronger, while under the surprise shock the credit decline becomes insignificant and output rises initially. (2) Replacing VSTOXX with VDAX (German analogue) — qualitatively the same. (3) Longer FEV truncation horizons k=30 and k=40 — quantitatively and qualitatively similar. (4) Including policies with missing announcement dates (182 of 1205 actions) in the empirical analysis — results barely change. (5) Granger-causality tests: the identified shocks are regressed on up to 3 principal components (explaining ~98.4% of variance) of seven forward-looking loan-officer credit-standard surveys; the null of no Granger causality cannot be rejected at any reasonable level (p-values range roughly 0.37–0.99). (6) Serial-correlation test regressing each shock on its own two lags: p-values 0.47 (surprise) and 0.77 (news), so no serial correlation.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It relates to (a) empirical work on macroprudential effectiveness and its monetary-policy interaction (Cerutti et al., Alam et al., Akinci and Olmstead-Rumsey, Kuttner and Shim, Budnik and Kleibl, etc.), most of which uses cross-country panels with GMM and cannot make clean causal claims; and (b) the SVAR/news-shock identification literature robust to foresight (Barsky and Sims 2011; Leeper et al. 2013; Kurmann and Otrok 2013; Ben Zeev et al. 2019). The two prior SVAR studies extracting exogenous macroprudential variation are Kim and Mehrotra (2017, four Asia-Pacific countries) and Klingelhofer and Sun (2019, China), both using recursive Cholesky orderings. Like Klingelhofer and Sun, the authors find macroprudential shocks explain a meaningful share of credit but little of prices. Unlike those studies, they find a strong macroprudential-monetary link (EAMPP-policy-rate correlation about −0.90, versus roughly +0.25 for Asia-Pacific in Bruno et al. 2017), and they are the first to identify both surprise and news macroprudential shocks. The narrative exclusion of cyclically-reactive policies follows Romer and Romer (2010), Richter et al. (2019), and Rojas et al. (2020).&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Macroprudential policy in the euro area effectively safeguards financial stability over the medium term by reducing credit growth, credit pro-cyclicality, and systemic risk, without a significant trade-off against price stability (the ECB&amp;rsquo;s primary target). Because more than one objective cannot be met with one instrument, monetary policy complements macroprudential policy: it can move accommodatively to offset output/credit declines, yielding an effective overall policy mix. Scope conditions: the conclusions are specific to the euro area over 1999:Q1–2019:Q2, a sample dominated by the GFC and sovereign debt crisis and by deflationary pressures (which is why the strong, negative macroprudential-monetary correlation may not generalize, e.g., to Asia-Pacific where the correlation is positive); the narrative EAMPP only captures proactive, long-run-financial-stability-motivated policies; and price-stability effects, while insignificant overall, carry wide estimate uncertainty.&lt;/p&gt;
&lt;h3 id="q7-why-does-the-paper-use-announcement-dates-rather-than-enforcement-dates"&gt;Q7. Why does the paper use announcement dates rather than enforcement dates?&lt;/h3&gt;
&lt;p&gt;Because foresight problems arise from inside and outside lags (Leeper et al. 2013): about 54% of euro-area policy tools in MaPPED experience a delay between announcement and implementation. Using the enforcement date would contaminate the identification of an &amp;lsquo;unanticipated&amp;rsquo; shock, since agents would already know about the policy from its announcement, making the shock no longer exogenous. The authors assume agents react from the announcement moment.&lt;/p&gt;
&lt;h3 id="q8-are-there-notable-caveats-about-the-index-and-impulse-responses"&gt;Q8. Are there notable caveats about the index and impulse responses?&lt;/h3&gt;
&lt;p&gt;The first EAMPP value is not zero because 185 of 1205 policy actions were implemented before 1995, and MaPPED does not provide announcement dates for 182 of 1205 actions (assumed equal to enforcement dates only for the stylized-facts section; removed in the empirical analysis). GDP-share weights use the 2008–2015 average; time-varying weights have very limited impact since GDP shares are stable. Impulse responses report median with 16th and 84th posterior percentiles. The EONIA-shadow-rate splice is justified by a 0.98 correlation between the two over 2004:Q4–2008:Q4.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Merger guidelines for the labor market</title><link>https://macropaperwarehouse.com/papers/merger-guidelines-for-the-labor-market/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/merger-guidelines-for-the-labor-market/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation. Antitrust review of mergers has historically focused almost entirely on harm to consumers (product-market monopoly), ignoring harm to workers (labor-market monopsony). Following the July 2021 White House executive order and the DOJ&amp;rsquo;s monopsony-based challenge to the Penguin Random House (PRH)/Simon &amp;amp; Schuster (SS) publishing merger, the agencies are now putting buyer power at the center of policy. The paper asks: how should Herfindahl-based merger-review thresholds, designed for product markets, perform if applied to local labor markets, and what efficiency gains would a merger need to leave workers unharmed?&lt;/p&gt;
&lt;p&gt;Model and data. The authors extend Berger, Herkenhoff, and Mongey (2022, &amp;ldquo;BHM&amp;rdquo;) to allow multi-plant (post-merger) ownership. The model has a representative household supplying labor through a nested-CES system (within-market substitutability governed by eta, across-market by theta, with eta &amp;gt; theta &amp;gt; 0), firms competing in quantities (Cournot/oligopsony), heterogeneous firm productivity, decreasing returns to scale, and capital. Firms set wages as a variable markdown on the marginal revenue product of labor; the markdown depends on the firm&amp;rsquo;s local payroll share. Markets are defined as 3-digit NAICS by commuting zone. Calibration is taken directly from BHM using confidential US Census data (LBD). Key estimated values: theta = 0.42 and eta = 10.85 (the elasticity-substitution parameters; the paper also reports theta = 0.45 in one passage), productivity dispersion sigma_z, returns to scale alpha, etc. The average market has 113 firms, an HHI of 0.11 (about nine equal firms), the average firm share is ~0.02, and the employment-weighted average markdown is 0.72 (workers paid 72% of marginal revenue product), equivalent to a labor-supply elasticity of 2.57.&lt;/p&gt;
&lt;p&gt;Theory. Proposition 1 shows that, absent efficiency gains, a within-market merger equalizes the two merged plants&amp;rsquo; markdowns at the level implied by their combined share, depresses both merging plants&amp;rsquo; wages, lowers the market wage index and employment, and reduces total worker pay. Non-merging firms&amp;rsquo; shares rise and they expand, so the actual rise in concentration is smaller than a &amp;ldquo;naive&amp;rdquo; calculation (adding pre-merger shares) would predict. Under the monopsony limit (infinitely many firms, or eta = theta), mergers have no effect.&lt;/p&gt;
&lt;p&gt;Main quantitative findings. (1) Model validation: replicating Arnold (2020), the model generates a change in log employment of -9.0 (vs Arnold -14.4, about three-fifths), log earnings -0.7 (vs -0.8), log payroll -10.5 (vs -12.1); earnings fall -4.4% in high-concentration vs -1.1% in medium-concentration markets (Arnold: -3.1% and -0.8%); the naive-concentration regression coefficient is 0.893 (Arnold 0.834), both below one. (2) PRH/SS simulation (PRH 37% share, SS 12%): with no efficiency gains the merger cuts author wages by 5%; the Required Efficiency Gain (REG) for worker-surplus neutrality is 17%. A merger of the two largest publishers gives -10% wages and a 30% REG; the two smallest Big Five give a 13% REG. (3) Applying product-market thresholds to labor markets via a 200,000-market simulation: under the stricter 1982 guidelines (block if post-merger HHI &amp;gt; 1800 and Delta-HHI &amp;gt; 100), the average REG of permitted mergers is 4.68%; under the looser 2010 guidelines (HHI &amp;gt; 2500, Delta-HHI &amp;gt; 200) it is 5.96%. Thus at the standard assumed 5% efficiency gain, 1982-permitted mergers raise the wage index (+0.04%) while 2010-permitted mergers lower it (-0.14%) and harm workers. (4) The Gross Downward Wage Pressure Index (GDWPI) equals (1/theta - 1/eta) times the other plant&amp;rsquo;s payroll share. Among mergers with GDWPI &amp;gt; 5% at both plants, more than 80% require a REG of at least 5.8% (20th-percentile REG = 5.8%, median 6.4%); among GDWPI &amp;gt; 10% at both plants, more than 80% generate a welfare loss under an assumed 5% efficiency gain.&lt;/p&gt;
&lt;p&gt;Implications. Product-market thresholds are too lenient for labor markets because labor is harder to substitute than products (low theta). The framework lets regulators trade off Type I error tolerance and efficiency-gain priors to set concentration thresholds.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identificationestimation-strategy-for-the-key-parameters-and-what-are-the-threats-to-it"&gt;Q1. What is the identification/estimation strategy for the key parameters, and what are the threats to it?&lt;/h3&gt;
&lt;p&gt;The model is not separately estimated; calibration is inherited wholesale from BHM (2022). The crucial labor-supply substitution parameters theta (across-market) and eta (within-market) are estimated in BHM from tradeable firms&amp;rsquo; market-share-dependent employment responses to corporate tax changes, identifying how much firms with different market shares move employment when after-tax returns change. Productivity dispersion sigma_z matches the payroll-weighted HHI, alpha matches labor&amp;rsquo;s share, gamma the capital share, Z mean firm size, and phi mean worker earnings. Main threats: (i) theta and eta are estimated from tradeable (largely manufacturing) firms and held fixed economy-wide, while the authors acknowledge no economy-wide substitutability estimates exist outside manufacturing; (ii) markets are defined by NAICS3-by-CZ rather than occupation (the conceptually preferred unit), because occupation codes are unavailable for the universe of workers; (iii) the whole exercise relies on the calibrated structure being the right laboratory.&lt;/p&gt;
&lt;h3 id="q2-how-is-the-model-validated-out-of-sample"&gt;Q2. How is the model validated out of sample?&lt;/h3&gt;
&lt;p&gt;By replicating Arnold (2020), who estimates causal labor-market effects of US mergers. The authors draw and merge two firms per market, impose a pre-merger employment cutoff (tilde-n = 46, about five times average firm size) so that median pre-merger employment matches Arnold&amp;rsquo;s sample (116), and run Arnold&amp;rsquo;s exact regressions on simulated data. The model reproduces the sign and roughly the magnitude of employment and wage declines, the concentration interaction (effects more than three times larger in high-concentration markets), and the sub-one naive-concentration coefficient. This is out-of-sample because none of these moments were targeted in calibration.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-central-welfare-metric-and-policy-quantity"&gt;Q3. What is the central welfare metric and policy quantity?&lt;/h3&gt;
&lt;p&gt;Worker Surplus Neutrality: a merger is worker-surplus neutral if the market-level wage index W_j is unchanged (using a household problem in which profits are NOT rebated, to mirror the product-market consumer-surplus standard). The key policy object is the Required Efficiency Gain (REG, Delta-star): the common post-merger productivity gain at both plants needed to keep W_j constant. By Proposition 1.5 the REG is always positive.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-main-mechanisms-and-what-is-downward-wage-pressure-specifically"&gt;Q4. What are the main mechanisms, and what is downward wage pressure specifically?&lt;/h3&gt;
&lt;p&gt;Market power comes from costly worker mobility within (eta) and across (theta) markets. When two plants merge, hiring at Plant 1 raises the market wage and thus the wage the merged firm must pay its inframarginal workers at Plant 2 (and vice versa). The merged firm internalizes this cross-plant cost, which acts like a per-worker &amp;rsquo;labor cannibalization tax,&amp;rsquo; lowering the marginal benefit of hiring at both plants, so it hires less and pays less. Downward wage pressure at Plant 1 equals n_2j times the derivative of w_2j with respect to n_1j; in share form DWP_1j = w_1j (1/theta - 1/eta) s_2j. The GDWPI normalizes this by the wage: GDWPI_1j = (1/theta - 1/eta) s_2j, bounded in [0, theta^-1 - eta^-1], interpretable as a wage tax rate. Larger partner share and higher within-market substitutability (eta) raise downward pressure.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented"&gt;Q5. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Effects vary strongly with concentration: earnings fall -4.4% in high-concentration markets vs -1.1% in medium-concentration markets (model). Effects depend on the merging firms&amp;rsquo; shares: assuming a 5% efficiency gain, fewer than 12.1% of mergers in which the smaller firm&amp;rsquo;s payroll share exceeds 5% yield a worker-surplus gain. REGs differ across publisher pairings in the PRH case (17% for PRH+SS, 30% for the two largest, 13% for the two smallest). The model also generates wide firm-level variation in markdowns (small firms near competitive, large firms marked down well below 0.72).&lt;/p&gt;
&lt;h3 id="q6-what-do-the-confidencethreshold-figures-show"&gt;Q6. What do the confidence/threshold figures show?&lt;/h3&gt;
&lt;p&gt;Fixing a 5% efficiency gain, the simulation reports the fraction of mergers yielding a worker-surplus gain by concentration cell. 89.5% of mergers with post-merger HHI &amp;lt; 500 and Delta-HHI &amp;lt; 50 yield gains. Under the 2010 highly-concentrated definition (HHI &amp;gt; 2500, Delta-HHI &amp;gt; 100 in the cited cell), fewer than 34.8% yield gains. A merger with small-firm share 4% and large-firm share 18% has a 69.7% chance of a worker-surplus gain at 5% efficiency, rising to 97.7% at a 10% efficiency gain. This lets a regulator pick thresholds for a desired Type I error tolerance.&lt;/p&gt;
&lt;h3 id="q7-how-sensitive-are-results-to-the-assumed-efficiency-gain"&gt;Q7. How sensitive are results to the assumed efficiency gain?&lt;/h3&gt;
&lt;p&gt;Highly. Under 1982 guidelines, permitted mergers change average W_j by -0.40% at 1% efficiency, &amp;hellip; up to +0.04% at 5% efficiency; blocked mergers fall -7.39% (1%) to -5.99% (5%). Under 2010 guidelines, permitted mergers fall -0.63% (1%) to -0.14% (5%); blocked mergers fall -10.37% (1%) to -8.61% (5%). The 5% benchmark (Farrell-Shapiro) is itself questioned: Blonigen and Pierce (2016) find roughly zero or negative merger productivity gains, implying even the 1982 thresholds may be too lenient.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-differ-from-closely-related-prior-work"&gt;Q8. How does this paper differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It extends BHM by adding multi-plant ownership and merger analysis. Relative to Nocke and Schutz (2018a,b) and Nocke and Whinston (2022), who derive product-market merger comparative statics under Bertrand competition (and, for Nocke-Whinston, CRS), this paper derives results for the LABOR market under nested-CES supply, Cournot competition, decreasing returns to scale, and endogenous household income. Relative to Naidu, Posner, Weyl (2018) and Marinescu-Hovenkamp (2019), who translate downward-wage-pressure concepts but assume symmetric firms, this paper provides a downward-wage-pressure test with firm heterogeneity across and within markets and shows it can be computed from readily available payroll shares and existing eta/theta estimates. It empirically benchmarks to Arnold (2020) and Prager-Schmitt (2021).&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Product-market HHI thresholds are too lenient when applied to labor markets: at an assumed 5% efficiency gain, 1982 thresholds (1800/100) keep permitted mergers worker-surplus neutral while 2010 thresholds (2500/200) do not. Scope conditions: (i) results hinge on the assumed efficiency gain (which empirical evidence suggests may be well below 5%); (ii) the framework treats product-market effects as &amp;lsquo;out of market&amp;rsquo; and should be combined with consumer-harm analysis; (iii) parameters are economy-wide benchmarks that may not fit a specific industry; (iv) market definition (NAICS3-by-CZ) matters, though the low estimated theta makes it consistent with a hypothetical-monopsonist test. The framework can be modified to add monopolistic pricing or variable markups (e.g., Deb et al. 2022).&lt;/p&gt;
&lt;h3 id="q10-are-there-internal-inconsistencies-a-reader-should-note"&gt;Q10. Are there internal inconsistencies a reader should note?&lt;/h3&gt;
&lt;p&gt;Yes. Table 1 reports theta = 0.42 (and 1.49 as the data moment), but the text at one point states &amp;rsquo;theta = 0.45, and eta = 10.85, giving theta^-1 - eta^-1 = 2.29.&amp;rsquo; The 2010 threshold is described in the abstract/Section 3 as Delta-HHI &amp;gt; 200 but the headline simulation result (4.68% vs 5.96%) compares &amp;lsquo;1800/100&amp;rsquo; against &amp;lsquo;2500/200&amp;rsquo;, and one passage lists the 2010 thresholds as (2500, 200) while the highly-concentrated text uses Delta-HHI of 200 for presumption and 100 in a figure cell. These are presentational; the substantive ranking (1982 stricter, 2010 more lenient) is robust.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;!-- flags: Internal parameter inconsistency: Table 1 reports theta=0.42 but text states theta=0.45 in the GDWPI bound passage (theta^-1 - eta^-1 = 2.29)., Threshold reporting: 1982 simulation uses Delta-HHI&gt;100 while Section 3 text also references Delta-HHI thresholds of 100/200; the headline comparison is 1800/100 vs 2500/200., Efficiency-gain assumption of 5% (Farrell-Shapiro) is load-bearing for the 'workers harmed under 2010 guidelines' conclusion; paper itself notes empirical evidence (Blonigen-Pierce 2016) of near-zero gains. --&gt;</description></item><item><title>Monetary Policy, Firm Heterogeneity, and the Distribution of Investment Rates</title><link>https://macropaperwarehouse.com/papers/monetary-policy-firm-heterogeneity-and-the-distribution-of-investment-rates/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/monetary-policy-firm-heterogeneity-and-the-distribution-of-investment-rates/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; Investment is a sizable and the most volatile component of aggregate GDP, so understanding the investment channel of monetary policy matters for policymakers. Prior work has overwhelmingly studied the effect of monetary policy on the &lt;em&gt;average&lt;/em&gt; investment rate. But an estimated average effect can reflect either a uniform rightward shift of the entire distribution (all firms invest a bit more) or a change in the &lt;em&gt;shape&lt;/em&gt; of the distribution (a few firms invest a lot more). The paper asks: how does monetary policy reshape the cross-sectional distribution of firm investment rates, and what does that reveal about the frictions driving (heterogeneous) transmission?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and empirical strategy.&lt;/strong&gt; Quarterly firm-level data from Compustat, sample 1986Q1–2018Q4, U.S. nonfinancial firms (financial firms, foreign firms, and firms with incomplete/questionable data excluded). Firm age is merged from WorldScope and Jay Ritter&amp;rsquo;s database. Accounting capital stocks are converted to real economic capital via a Perpetual Inventory Method (building on Bachmann and Bayer 2014). The investment rate is real capital expenditures (CAPX) net of sales of property/plant/equipment (SPPE), deflated and divided by the lagged real capital stock. The firm-level data are aggregated into quarterly investment-rate distributions and moments. Identification uses monetary policy shocks from the Gertler and Karadi (2015) Proxy SVAR (re-extracted with updated VAR data and high-frequency instruments). Estimation is via two-step quantile/bin local projections (eq. 1), with quarter dummies for seasonality and Newey-West standard errors. Shocks are scaled to reduce the 1-year Treasury yield by 25 basis points (100bp in some distribution figures for readability). As a validity check, an expansionary shock produces hump-shaped increases in investment (peak 1.4%) and GDP (peak 0.35%).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings (three facts).&lt;/strong&gt; Fact 1: An expansionary shock changes the shape of the distribution — fewer zero and small investment rates and more large ones. The 75th percentile responds significantly more than the 25th (the interquartile range rises significantly); the share of firms in bins [0,2) and [2,4) falls significantly while higher positive bins rise, most sizably in bin [28,infinity); negative investment rates are not meaningfully affected. The spike rate (share with investment rate &amp;gt;10%) rises and the inaction rate (|i|&amp;lt;0.5%) falls. Fact 2: These shape changes are more pronounced and statistically significant among young firms (defined as less than 15 years old) than old firms; spike rates rise more and inaction rates fall more for young firms. These effects persist even among firms unlikely to be financially constrained (low leverage, high liquidity, or dividend payers), arguing against a purely financial explanation. Fact 3: A decomposition (eq. 3) into extensive vs. intensive margins shows the extensive margin accounts for around 60% (intensive 40%) of the effect on the average investment rate, and around 60% (intensive 40%) of the &lt;em&gt;heterogeneous&lt;/em&gt; average effect across age groups.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model and mechanism.&lt;/strong&gt; The authors build a general-equilibrium New Keynesian heterogeneous-firm model with fixed and convex capital adjustment costs, maintenance investment, and firm entry/exit (life cycles), in the spirit of Khan and Thomas (2008) and Winberry (2021). Calibrated to U.S. data (quarterly, beta=0.99), it replicates all three facts. Fixed costs generate lumpy investment and an extensive-margin channel: an interest-rate cut raises the discounted benefit of investing, inducing some firms to switch from inaction to a sizeable investment. Young firms are on average farther from their optimal capital (higher marginal product of capital under decreasing returns), so they are induced to invest more easily — generating heterogeneity &lt;em&gt;without any financial friction&lt;/em&gt;. This implies observational equivalence with the financial accelerator, but with opposite cyclicality: fixed costs imply &lt;em&gt;procyclical&lt;/em&gt; policy effectiveness, whereas financial acceleration implies &lt;em&gt;countercyclical&lt;/em&gt; effectiveness.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Aggregate/policy implications.&lt;/strong&gt; Monetary policy is most effective when many firms are &amp;ldquo;close to paying the fixed cost.&amp;rdquo; The decline in business dynamism / firm aging since the 1980s has made monetary policy about 12% less effective at stimulating investment; policy is also less effective in recessions than booms (about 22% more effective in a large boom than a deep recession).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The authors use exogenous monetary policy shocks from the Gertler and Karadi (2015) Proxy SVAR, re-extracted after updating both the VAR time-series data and the high-frequency (high-frequency surprise) instruments. These shocks are fed into two-step local projections: in the first step they construct time series of distributional objects (quantiles, interquartile range, the share of firms in each investment-rate bin, the spike rate, the inaction rate); in the second step (eq. 1) they regress the h-period change in each object on the shock, with calendar-quarter dummies to absorb seasonality and Newey-West standard errors for heteroskedasticity and autocorrelation. The validity check is that the shocks produce plausible hump-shaped aggregate responses (investment peak 1.4%, GDP peak 0.35%). The key threats are the standard ones for high-frequency-identified monetary shocks (the shock series being a valid instrument / external to the outcome) and the aggregation step; the paper does not run firm-level panel regressions with firm fixed effects here but instead works on aggregated distributional time series, so threats relate to the time-series identification of the GK shocks rather than firm-level confounding.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;Two margins: the intensive margin (firms changing the size of investment conditional on adjusting) and the extensive margin (firms changing whether to invest at all). Empirically they are separated via the decomposition in equation (3), which classifies observations into spikes (i&amp;gt;10%) and normal (i&amp;lt;=10%) and writes the average rate as the spike fraction times the conditional spike rate plus the complementary term. The extensive-margin component isolates the change in the average rate coming only from changes in the spike rate; the intensive component isolates changes in conditional investment rates. Two covariance terms are dropped as negligible. The shape change in the distribution (fewer small, more very-large investments, negatives unaffected), plus the rising spike rate and falling inaction rate, are the empirical fingerprints of the extensive margin. The decomposition attributes about 60% of the average effect to the extensive margin.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Heterogeneity by firm age (young = less than 15 years old, old = 15+). Young firms show larger and more statistically significant shape changes (bigger drop in bin [0,2), bigger rise in bin [28,infinity)), larger spike-rate increases, and larger inaction-rate declines. The disproportionate right-tail (upper-quantile) response holds in both groups but is much more pronounced for young firms. The extensive margin explains roughly 60% of the young-vs-old gap in average effects. Appendix C reports similar but quantitatively weaker results when comparing small vs. large firms instead of young vs. old. The heterogeneous age effect survives within groups unlikely to be financially constrained (low leverage, high liquidity, dividend payers) and is also present among likely-constrained firms.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-model-decompose-the-heterogeneous-extensive-margin-effect-and-what-is-the-heterogeneous-size-effect"&gt;Q4. How does the model decompose the heterogeneous extensive-margin effect, and what is the &amp;lsquo;heterogeneous size effect&amp;rsquo;?&lt;/h3&gt;
&lt;p&gt;Using eq. (22), the heterogeneous extensive-margin effect splits into (i) a &amp;lsquo;heterogeneous hazard rate increase&amp;rsquo; — an interest-rate cut raises young firms&amp;rsquo; hazard (adjustment probability) more than old firms&amp;rsquo;, because young firms have a higher marginal product of capital and are farther from optimal size, so the discounted benefit of investing rises more for them; and (ii) a &amp;lsquo;heterogeneous size effect&amp;rsquo; — among new adjusters, young firms choose higher conditional investment rates than old firms, so there would be a heterogeneous average effect even if hazard rates rose identically. Both are quantitatively important.&lt;/p&gt;
&lt;h3 id="q5-what-role-do-the-different-adjustment-costs-play-and-how-is-the-model-calibrated"&gt;Q5. What role do the different adjustment costs play, and how is the model calibrated?&lt;/h3&gt;
&lt;p&gt;The model has fixed adjustment costs (random, uniform on [0, xi-bar]), convex adjustment costs (parameter phi), and maintenance investment (parameter chi). In isolation, the fixed cost generates 55% of the heterogeneous average effect and the convex cost only 29%, with the remaining 16% from their interaction (the heterogeneous size effect needs both: hazard changes require fixed costs, differing conditional rates require convex costs). Five parameters (sigma_z=0.07, k0=2.27, xi-bar=0.90, phi=2.20, chi=0.34) are fitted to five moments: standard deviation of investment rates (data 0.20 / model 0.18), average investment rate (0.12/0.13), autocorrelation of investment rates (0.38/0.38), relative size of entrants (0.29/0.29), and relative spike rate of old firms (0.40/0.40). Fixed parameters include beta=0.99, psi=0.58, theta=0.21, nu=0.64, delta=1.93% (giving a 7.7% annual aggregate investment rate), rho_z=0.95, pi_exit=1.625%, phi(Rotemberg)=90, gamma=10, Taylor inflation coefficient phi_pi=1.5, smoothing rho_r=0.75, external capital adjustment cost kappa=11.&lt;/p&gt;
&lt;h3 id="q6-what-untargeted-moments-validate-the-model"&gt;Q6. What untargeted moments validate the model?&lt;/h3&gt;
&lt;p&gt;The model reproduces (i) firm life-cycle profiles — average investment rate highest for newborns and falling with age, decomposed into frequency of adjustment (extensive) and conditional investment rate (intensive), both higher for young firms; (ii) plausible aggregate monetary-policy responses; and (iii) the interest-rate elasticity of aggregate investment. All three investment frictions are needed for the life-cycle profiles: fixed costs generate adjustment frequencies below one, convex costs keep young firms&amp;rsquo; conditional investment rates plausible (no instant jump to optimal size), and maintenance investment makes hazard rates decline with age.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-run"&gt;Q7. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Robustness to alternative quantile choices (Figure A.1); alternative spike thresholds of 8% and 12% (Figure A.8); using the spike rate vs. hazard rate to identify extensive-margin adjustments in the model (Figure A.12, very similar results); replication of heterogeneous spike/inaction effects within groups unlikely to be financially constrained (Figure A.6) and within likely-constrained firms (Figure A.7); small-vs-large firm comparison (Appendix C); and comparison of extensive-margin contributions across different shocks (aggregate TFP, wage-markup) in Appendix E.4, showing the extensive-margin contribution can differ substantially when a shock directly affects adjustment costs.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q8. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It builds on the empirical investment-channel literature (Christiano et al. 2005; Gertler and Gilchrist 1994; Ottonello and Winberry 2020; Jeenas 2023; Cloyne et al. 2023) which focused on aggregate or average investment rates; its novelty is documenting effects on the &lt;em&gt;entire distribution&lt;/em&gt; and its moments. Against Cloyne et al. (2023), who interpret stronger young-firm responsiveness through the financial accelerator, this paper shows a non-financial friction (fixed adjustment costs) generates the same age heterogeneity — an observational-equivalence point — though it stresses its findings are &amp;lsquo;consistent with&amp;rsquo; and &amp;rsquo;not necessarily at odds with&amp;rsquo; the financial accelerator (the intensive margin, stronger among young firms, may reflect financial acceleration). On the lumpy-investment theory side it extends Khan and Thomas (2008), Winberry (2021), Koby and Wolf (2020), Reiter et al. (2013, 2020), Fang (2023) by adding firm life cycles. Relative to contemporaneous work by Lee (2023), which examines spike rates of small vs. large firms, this paper studies young vs. old firms and the entire distribution; relative to Gourio and Kashyap (2007), who study unconditional spike-rate cyclicality, this paper studies responses to monetary shocks.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Monetary policy stimulates aggregate investment mainly because a few firms switch from inaction to sizeable investment (extensive margin), not because many firms invest a little more. Effectiveness is state-dependent: it is higher when many firms are &amp;lsquo;close to paying the fixed cost&amp;rsquo; — i.e., in booms and in high-business-dynamism economies with many young, growing firms. Scope conditions/quantification: the post-1980s decline in business dynamism / firm aging has made policy about 12% less effective; the impact effect on aggregate investment is 1.44% in baseline, 1.61% (about 11.5% larger) under a high-dynamism calibration (13% entrant share, as in 1984) and 1.32% (about 8.5% smaller) under low dynamism (3.375% entrant share); policy is about 22% more effective in a large boom than a deep recession. Critically, the cyclicality direction differs from the financial accelerator: fixed costs imply &lt;em&gt;procyclical&lt;/em&gt; effectiveness, financial acceleration implies &lt;em&gt;countercyclical&lt;/em&gt; — a distinction that matters for policy and aligns with evidence (Tenreyro and Thwaites 2016) that policy is weaker in recessions. A key caveat from general equilibrium: a higher young-firm share does not automatically raise effectiveness, because higher investment demand raises the price of capital and crowds out investment; state dependence only arises when the price elasticity of aggregate investment is sufficiently low (as in their model).&lt;/p&gt;
&lt;h3 id="q10-what-are-the-main-caveats-and-open-questions"&gt;Q10. What are the main caveats and open questions?&lt;/h3&gt;
&lt;p&gt;The extensive-margin channel cannot rationalize the entire young-old responsiveness gap — the intensive margin is also quantitatively relevant and may reflect financial acceleration. The roughly-60% extensive-margin share of the heterogeneous effect cannot be rationalized by the classical Bernanke-Gertler-Gilchrist (1999) financial accelerator, which operates on the intensive margin. The spike rate is used as an empirical proxy for the model&amp;rsquo;s unobservable hazard rate. The paper leaves open why young firms grow slowly, how the relevant frictions respond to economic policy, and how policy effects are shaped by these frictions, pointing to non-financial constraints like productivity/demand uncertainty (Jovanovic 1982; Chen et al. 2023) as further avenues.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Nonresponse Bias in Household Inflation Expectations Surveys</title><link>https://macropaperwarehouse.com/papers/nonresponse-bias-in-household-inflation-expectations-surveys/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/nonresponse-bias-in-household-inflation-expectations-surveys/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Inflation expectations measured from household surveys are central inputs to monetary policy, but roughly half of respondents to the RBNZ Household Inflation Expectations survey decline to answer the quantitative inflation-expectations question. Because these item non-responses are not random across demographic groups, aggregate and subgroup measures derived only from those who answer can be systematically biased. The paper quantifies that non-response bias and proposes a simple, operational method to correct aggregate and subgroup inflation-expectation indices and disagreement measures.&lt;/p&gt;
&lt;p&gt;Data and strategy: Micro-data from the RBNZ Household Inflation Expectations survey, quarterly, achieving about 1,000 household responses per wave, covering 1998Q2 to 2022Q4 with 89,834 individual responses treated as repeated cross-sections. The focal question asks the expected annual rate of inflation/deflation over the next 12 months. The survey switched from telephone to online mode starting 2018Q3. Outliers are removed using a 1.5xIQR rule (excluding 4,535 observations in the baseline). The empirical approach has three steps: (1) Probit models of the probability of responding on demographics (gender, age, region, ethnicity, income, employment) plus macro controls (lagged inflation and its square, a year trend, seasonal dummies, an online-mode dummy); (2) a Heckman sample selection model (selection equation = the baseline Probit extended with online-mode interactions; outcome equation = inflation-expectation bias regression) with four exclusion restrictions dropped from the outcome equation (region, employment, year trend, lagged inflation squared); (3) a regression-on-quarter-dummies index that adds the inverse Mills ratio to deliver bias-adjusted average and dispersion series. Estimates use survey weights, extending Heckman estimators to weighted form.&lt;/p&gt;
&lt;p&gt;Main quantitative findings: Item non-responses average about 44% over the full sample, falling to about 24% after the move to online mode. Non-responses artificially raise average one-year-ahead inflation expectations by about 0.3 percentage points; the average selection adjustment is -0.288 over the full sample, ranging from -0.385 (2018Q1) to -0.138 (2022Q3). Females are about 20% less likely to respond than men; older, employed, higher-income individuals respond more; Maori and Pacific Islanders respond less. Online mode raises response probability by about 33%. Response rates rise non-linearly with lagged inflation: moving from 2% to 7% raises average response probability by about 12%, while it barely changes over the 0-4% range, with the slope turning steeply positive in the 5-7% range. There is a downward trend in response of about 1% more item non-response per year. The online switch narrowed the female-male response gap from 24.4% (telephone) to 5.5% (online) and rendered most ethnicity gaps insignificant. In the bias (outcome) regressions without selection (weighted), respondents over 25 show bias more than 0.23 pp above the under-25 base; Pacific Islanders 0.34 pp, Maori 0.15 pp, Asians 0.12 pp above the base ethnic group. After the Heckman correction, gender, ethnicity, and income differences become insignificant or shrink substantially, while age effects strengthen (older respondents over-predict; under the two-step estimator, bias for those over 35 is more than double the no-selection estimate). The online dummy in the outcome equation lowers predicted expectations by more than 2.27 pp (interpreted cautiously, as it also captures large 2020Q3-onward negative biases).&lt;/p&gt;
&lt;p&gt;Implications: Survey weights correct unit non-response but not item non-response, so published aggregates overstate expectations by ~0.3 pp. The correction lowers all subgroup means, decreases cross-subgroup disagreement for gender/income/ethnicity (increases it across age), and generally decreases within-subgroup dispersion. Correcting also makes the household-vs-professional-forecaster intercept gap statistically insignificant. Policy: online survey modes and inclusive, layered communication (especially during high-inflation periods of greater public attention) can reduce measurement error.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Identification rests on a Heckman sample selection model. A Probit selection equation models the probability of answering the inflation-expectations question; its predicted probabilities yield the inverse Mills ratio, added to the outcome (bias) regression to correct for selection-as-omitted-variable bias. Identification is sharpened by exclusion restrictions: four variables (region, employment status, year trend, lagged inflation squared) enter the selection equation but are dropped from the outcome equation. The authors justify these because region and employment were found statistically insignificant in the outcome equation, and year trend and lagged inflation squared induced collinearity/variance inflation. The selection equation also includes online-mode interaction terms to better identify heterogeneity in response rates. Threats: the validity of the exclusion restrictions (the assumption that these variables affect participation but not the level of expectations bias) and the known sensitivity of the full-information ML Heckman estimator to collinearity; the authors address the latter by also reporting the two-step estimator.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;Two mechanisms drive non-response. First, demographic propensity: young, female, low-income, and minority-ethnicity (Maori, Pacific Islander, Asian) respondents are less likely to answer, documented via Probit average partial effects. Second, state dependence on the inflation environment: response rates rise non-linearly when lagged inflation moves away from the target range (steeply positive slope at 5-7%), consistent with a &amp;lsquo;rational inattention&amp;rsquo; interpretation where agents notice inflation only when it becomes salient, and with the finding that inflation uncertainty co-moves with the inflation level (Binder, 2017). The authors also test whether non-response reflects lack of understanding using a 2018Q3-2021Q4 sub-question: only 5% of respondents indicated not understanding inflation, so 81% of non-responses are not due to lack of understanding, pointing instead to factors like cultural norms/uncertainty rather than literacy.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Response heterogeneity: females respond ~20% less than males; response probability rises with age; Maori and Pacific Islanders respond markedly less; higher income and employment raise response; households with dependent children and non-freehold owners respond less; being the main grocery shopper slightly lowers response. Bias heterogeneity before correction: age, ethnicity (Pacific Islanders 0.34 pp, Maori 0.15 pp, Asian 0.12 pp), and income show differences. After Heckman correction, gender, ethnicity, and income differences become insignificant or shrink substantially, while age effects strengthen (older respondents over-predict inflation, with an upward-sloping age profile). Online mode reduces demographic gaps: the female-male response gap fell from 24.4% to 5.5%, and most ethnicity gaps became insignificant online.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;(1) Four Probit specifications with progressively richer covariates (occupation, grocery shopping, dependent children, home ownership) across sub-periods, with baseline effects stable. (2) Two Heckman estimators, two-step and ML, mostly consistent (the main divergence is gender, insignificant under two-step). (3) Comparison against random imputation, which reproduces the distorted no-selection picture. (4) Six outlier-detection rules (fixed -2/15 interval, 1.5xIQR, 3xIQR, hybrid IQR, top/bottom 5% by quarter, top/bottom 5% overall): Probit estimates are insensitive to the outlier definition. (5) A separate Probit on outlier responses shows similar demographic patterns (low-income young minority females give more outlier responses) but with differing magnitudes and trend/inflation effects, indicating outlier responses and non-responses are related but distinct. (6) An Appendix-E forward-looking Phillips curve exercise where adjusted subgroup expectations are always preferred to unadjusted.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It builds on the heterogeneity-of-expectations literature (Bruine de Bruin et al. 2010; Pfajfar and Santoro 2010; Malmendier and Nagel 2016; D&amp;rsquo;Acunto et al. 2023) documenting demographic differences in expectations, and on studies finding non-response from young/female/low-income groups (Blanchflower and MacCoille 2009; Leung 2009). Its distinctive contribution is showing that part of the observed gender/ethnicity/income differences in expectations is an artifact of non-response (selection) rather than true belief differences, and proposing an operational correction. Unlike imputation methods (e.g., the US Michigan Survey&amp;rsquo;s distribution-based imputation), the Heckman approach accounts for the socio-demographic composition of responders. Unlike methods requiring randomized incentives or special survey-design features (McGovern et al. 2018; Comerford 2023), it works on long-running repeated cross-sections lacking such features. It differs from attrition-focused work (Burgi 2023) by addressing item non-response in repeated cross-sections rather than panel attrition.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;First, because survey weights correct only unit non-response, published aggregates overstate expectations by ~0.3 pp; central banks should apply an item-non-response correction. Second, response engagement rises when inflation deviates from target, so central banks could leverage high-inflation periods of elevated public attention for broader communication beyond financial-market audiences, using layered messaging. Third, moving surveys online substantially reduces non-response bias and improves representativeness, but requires ensuring digital accessibility to avoid new selection bias. Scope conditions: the non-linear inflation-response relationship is based on few episodes of out-of-range inflation, possibly confounded by Covid/recessions, so it should be interpreted with caution; the large online-mode coefficient on expectations also captures the post-2020Q3 negative biases from sluggish expectation adjustment; and RBNZ owns the survey and could change methodology accordingly.&lt;/p&gt;
&lt;h3 id="q7-how-is-the-adjusted-index-constructed-operationally-and-why-is-it-attractive"&gt;Q7. How is the adjusted index constructed operationally, and why is it attractive?&lt;/h3&gt;
&lt;p&gt;Average expectations are obtained by regressing micro inflation-expectations on quarter dummies (WLS); adding the inverse Mills ratio from the baseline Probit as an extra regressor yields the bias-adjusted average. Subgroup indices interact subgroup dummies with time dummies; an adjusted disagreement (dispersion) measure replaces the dependent variable with squared deviations from the quarterly mean. The approach is attractive operationally because updating each quarter only requires a new inverse Mills ratio from the pre-fitted, relatively stable Probit model, so the adjustment is unlikely to undergo severe revisions.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-comparison-with-professional-forecasters-show"&gt;Q8. What does the comparison with professional forecasters show?&lt;/h3&gt;
&lt;p&gt;Regressing one-year-ahead Survey of Professional Forecasters expectations on household expectations, the unadjusted household series gives a negative, significant intercept (-0.294, confirming households&amp;rsquo; upward divergence), but using the adjusted household average makes the intercept insignificant (-0.019), suggesting the household-professional gap is partly a non-response artifact. The slope remains below one (0.759 unadjusted, 0.740 adjusted), consistent with Carroll (2003), so household expectations still do not scale one-to-one with professional forecasters.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Precautionary Saving against Correlation under Risk and Ambiguity</title><link>https://macropaperwarehouse.com/papers/precautionary-saving-against-correlation-under-risk-and-ambiguity/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/precautionary-saving-against-correlation-under-risk-and-ambiguity/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: How much to save is a central household financial decision, and uncertainty drives the &amp;ldquo;precautionary saving motive.&amp;rdquo; The precautionary-saving literature has mostly studied one-dimensional (single-attribute) risk, yet households face multidimensional risk: both wealth and health conditions matter for saving. Because wealth and health are plausibly related, the authors argue the correlation between two risky attributes should be incorporated into precautionary-saving analysis. They further note that correlation between two attributes is harder to quantify than a single attribute&amp;rsquo;s risk (less experience, fewer observations), so they also introduce ambiguity about the correlation. The paper&amp;rsquo;s purpose is to characterize how the correlation between two risky attributes (wealth and health) affects optimal savings under multivariate preferences, both when correlation is known (risk) and when it is ambiguous.&lt;/p&gt;
&lt;p&gt;Model setup: A purely theoretical two-date model (t=0, t=1). The individual has time-separable lifetime utility from a bivariate utility function u(x,y) over wealth x and health y, increasing and concave in both (u^(1,0)&amp;gt;=0, u^(0,1)&amp;gt;=0, u^(2,0)&amp;lt;=0, u^(0,2)&amp;lt;=0); the sign of the cross derivative u^(1,1) is left unrestricted. The risk-free interest rate is zero and there is no time discounting, so the analysis isolates the effect of risk on saving. At t=1 the individual faces &amp;ldquo;good&amp;rdquo; and &amp;ldquo;bad&amp;rdquo; income risks (epsilon_G, epsilon_B occurring with probabilities 1-p, p) and &amp;ldquo;good&amp;rdquo;/&amp;ldquo;bad&amp;rdquo; health risks (delta_G, delta_B with probabilities 1-q, q), all four mutually independent. Correlation between income and health risk is captured by a parameter k: the probability of simultaneous bad income and bad health is kpq. When k=1 the risks are independent (joint probability = pq); k&amp;gt;1 (k&amp;lt;1) indicates positive (negative) correlation; correlation increases in k. The individual chooses saving s to maximize lifetime utility (equation 1). &amp;ldquo;Good&amp;rdquo; vs &amp;ldquo;bad&amp;rdquo; risks are ranked by stochastic dominance (FSD, Nth-order NSD, and Ekern&amp;rsquo;s Nth-degree risk increase).&lt;/p&gt;
&lt;p&gt;Main findings (theoretical propositions, no estimated magnitudes): (1) Proposition 1 — when income risk is ranked by Nth-order and health risk by Mth-order stochastic dominance, optimal savings increase (decrease) in correlation k if (-1)^(n+m) u^(n+1,m)(x,y) &amp;gt;= (&amp;lt;=) 0 for n=1..N, m=1..M. This condition defines &amp;ldquo;mixed correlation aversion (seeking).&amp;rdquo; In the special case N=M=1, optimal savings increase in k if u^(2,1)&amp;gt;=0, i.e., the individual is &amp;ldquo;cross prudent&amp;rdquo; (decrease if cross imprudent, u^(2,1)&amp;lt;=0). Intuition: cross-prudent individuals dislike the simultaneous occurrence of bad income and bad health, which becomes more likely as k rises, so they save more. (2) Proposition 2 (ambiguous correlation, smooth ambiguity model of Klibanoff et al. 2005, 2009) — if the second-order utility phi exhibits decreasing absolute ambiguity aversion (DAAA) and u exhibits mixed correlation aversion or seeking, then ambiguous correlation raises the optimal amount of savings relative to the risky benchmark with correlation k_O = sum q_theta k_theta. The result combines a &amp;ldquo;timing of uncertainty effect&amp;rdquo; (governed by beta(s_O)&amp;gt;=1 iff phi exhibits DAAA) and the sign of a covariance term. (3) Proposition 3 extends the same result to Nth-/Mth-degree risk increases: under DAAA and (-1)^(N+M) u^(N,M)&amp;gt;=(&amp;lt;=)0 and (-1)^(N+M) u^(N+1,M)&amp;gt;=(&amp;lt;=)0, ambiguous correlation raises savings.&lt;/p&gt;
&lt;p&gt;Implications: Whether correlation raises or lowers precautionary saving depends entirely on the signs of higher-order cross derivatives of utility, and under ambiguity additionally on the absolute-ambiguity-aversion coefficient. The authors link results to experimental evidence (Attema et al. 2019 find both cross prudence and imprudence; correlation aversion in gains, seekingness in losses) and to empirical work on public health systems, which by changing the wealth-health correlation affect precautionary saving (e.g., Rosen and Wu 2004; Atella et al. 2012; Chou et al. 2003; Jappelli et al. 2007), broadly consistent with cross prudence.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-mechanism-linking-correlation-to-saving-and-how-is-it-formalized"&gt;Q1. What is the core mechanism linking correlation to saving, and how is it formalized?&lt;/h3&gt;
&lt;p&gt;Correlation between income and health risk is parameterized by a single scalar k that scales the joint probability of the simultaneous bad outcome to kpq (with k=1 = independence, k&amp;gt;1 = positive correlation, k&amp;lt;1 = negative correlation), following the representation of Doherty and Schlesinger (1990). The derivative of expected period-1 utility with respect to k reduces (Lemma 1) to pq times [E[f(eps_B,del_B)] - E[f(eps_G,del_B)] - E[f(eps_B,del_G)] + E[f(eps_G,del_G)]], so the sign of the response to correlation is governed by a cross-difference whose sign maps directly onto the signs of higher-order cross derivatives of u. As k rises, the simultaneous occurrence of two bad outcomes becomes more likely; agents who dislike that combination (mixed correlation averse / cross prudent) save more to protect against it.&lt;/p&gt;
&lt;h3 id="q2-what-exactly-is-mixed-correlation-aversion-seeking-and-how-does-it-relate-to-correlation-aversion-and-cross-prudence"&gt;Q2. What exactly is &amp;lsquo;mixed correlation aversion (seeking)&amp;rsquo; and how does it relate to correlation aversion and cross prudence?&lt;/h3&gt;
&lt;p&gt;An individual is mixed correlation averse (seeking) if (-1)^(n+m+1) u^(n,m)(x,y) &amp;gt;= (&amp;lt;=) 0 for all n=1..N, m=1..M. It is a bivariate extension of Caballe and Pomansky&amp;rsquo;s (1996) univariate mixed risk aversion, and generalizes Epstein and Tanny&amp;rsquo;s (1980) correlation aversion (which corresponds to u^(1,1)&amp;lt;=0). Cross prudence (u^(2,1)&amp;gt;=0, per Eeckhoudt et al. 2007) is the third-order version of correlation aversion. The paper&amp;rsquo;s saving conditions use mixed correlation aversion (seekingness) excluding the second-order correlation-aversion term, expressed via the derivative pattern (-1)^(n+m) u^(n+1,m) &amp;gt;= (&amp;lt;=) 0.&lt;/p&gt;
&lt;h3 id="q3-how-is-the-good-vs-bad-ranking-of-risks-made-rigorous"&gt;Q3. How is the &amp;lsquo;good&amp;rsquo; vs &amp;lsquo;bad&amp;rsquo; ranking of risks made rigorous?&lt;/h3&gt;
&lt;p&gt;Through stochastic dominance. eps_G dominates eps_B in the sense of Nth-order stochastic dominance (NSD) iff E[u(w+eps_G,h)]&amp;gt;=E[u(w+eps_B,h)] for all u with (-1)^(n+1) u^(n,0)&amp;gt;=0, n=1..N (mixed risk aversion in wealth); analogously for health via Mth-order dominance (MSD). FSD corresponds to N=M=1. The paper also uses Ekern&amp;rsquo;s (1980) Nth-degree risk increase, where the first N-1 moments coincide (e.g., a 2nd-degree increase is a Rothschild-Stiglitz mean-preserving spread; a 3rd-degree increase is an increase in downside risk per Menezes et al. 1980).&lt;/p&gt;
&lt;h3 id="q4-how-is-ambiguity-about-correlation-modeled-and-what-drives-the-ambiguity-result"&gt;Q4. How is ambiguity about correlation modeled, and what drives the ambiguity result?&lt;/h3&gt;
&lt;p&gt;The individual perceives a finite set of possible correlations {k_1&amp;lt;&amp;hellip;&amp;lt;k_Theta} with subjective second-order probabilities q_theta, and evaluates them via the recursive smooth ambiguity model of Klibanoff et al. (2005, 2009) using an increasing, concave, thrice-differentiable second-order utility phi (concavity = ambiguity aversion). Evaluating the FOC at the benchmark s_O (the optimum under the mean correlation k_O = sum q_theta k_theta) decomposes the effect into a &amp;rsquo;timing of uncertainty effect&amp;rsquo; (Osaki and Schlesinger 2014), captured by beta(s_O) which is &amp;gt;=1 iff phi exhibits decreasing absolute ambiguity aversion (DAAA), plus a covariance term Cov(phi&amp;rsquo;(v), v_s). Under mixed correlation aversion/seeking, v(s,k) and v_s(s,k) move in opposite directions in k (Lemma 3), so because phi&amp;rsquo; is decreasing the covariance is positive; combined with DAAA this yields higher savings (Proposition 2).&lt;/p&gt;
&lt;h3 id="q5-what-is-the-role-of-decreasing-absolute-ambiguity-aversion-daaa"&gt;Q5. What is the role of decreasing absolute ambiguity aversion (DAAA)?&lt;/h3&gt;
&lt;p&gt;DAAA (lambda(z) = -phi&amp;rsquo;&amp;rsquo;(z)/phi&amp;rsquo;(z) decreasing in z) is the ambiguity analogue of decreasing absolute risk aversion. The Appendix proves (following Osaki and Schlesinger 2014) that beta(s)&amp;gt;=1 iff the ambiguity precautionary premium Psi_A &amp;gt;= the ambiguity premium pi_A, which is equivalent to DAAA. DAAA ensures the timing-of-uncertainty effect pushes toward more saving. The authors caution that empirical/experimental evidence on the sign of absolute ambiguity aversion is thin; Berger and Bosetti (2020) is cited as an exception finding evidence for DAAA, and the authors say more evidence is needed.&lt;/p&gt;
&lt;h3 id="q6-how-do-the-theoretical-predictions-connect-to-experimental-and-empirical-observations"&gt;Q6. How do the theoretical predictions connect to experimental and empirical observations?&lt;/h3&gt;
&lt;p&gt;Experimentally, Attema et al. (2019) measure multivariate risk preferences (wealth and longevity as a health proxy) and observe both cross prudence and cross imprudence, and correlation aversion in the gain domain with correlation seekingness in the loss domain. So the model implies savings can rise or fall with correlation depending on the individual. Empirically, the wealth-health correlation is shaped by public health systems: a more protective system separates wealth and health risk (lowers correlation). Rosen and Wu (2004) find poor health leads to safer investment (consistent with cross prudence); Atella et al. (2012) find households invest more in risky assets when health risk is mitigated by a protective national health system; Chou et al. (2003, Taiwan) find public health insurance reduced precautionary saving (a correlation decrease); Jappelli et al. (2007, Italy) find higher precautionary saving where health care quality is lower (a correlation increase); Ayyagari and He (2017) and Christelis et al. (2020) find Medicare/Medicare Part D increased risky investment. These are described as consistent with cross prudence.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-differ-from-the-closest-prior-work"&gt;Q7. How does this paper differ from the closest prior work?&lt;/h3&gt;
&lt;p&gt;Versus Eeckhoudt and Schlesinger (2008), which studies how risky shifts in future income affect saving via higher-order stochastic dominance, this paper adds correlation between two attributes and multivariate preferences. Versus Courbage and Rey (2007), who compare a certain-health vs risky-health setting, this paper compares two settings where health is risky in both but the income-health correlation differs, using the simpler Doherty-Schlesinger (1990) correlation representation. Versus Osaki and Schlesinger (2014) and Gierlinger and Gollier (2017), who study ambiguity in future income, this paper introduces ambiguity into the correlation rather than into income itself. The mixed-correlation-aversion concept builds on Jokung (2011) and Eeckhoudt et al. (2007, 2009).&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Because public health systems alter the correlation between wealth and health (e.g., medical-expense coverage separates the two risks, lowering correlation), they affect precautionary saving. The directional prediction is conditional: under cross prudence, lower correlation (more generous public health coverage) reduces precautionary saving and a positive wealth-health correlation raises saving above the independence benchmark; under cross imprudence the signs reverse. Under ambiguity the prediction additionally requires DAAA plus the relevant cross-derivative sign pattern. The authors stress that because experimental evidence shows both cross prudence and imprudence, no unconditional policy prediction follows &amp;ndash; e.g., for cross-imprudent individuals ambiguous correlation might lower savings.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-main-caveats-and-directions-for-future-research"&gt;Q9. What are the main caveats and directions for future research?&lt;/h3&gt;
&lt;p&gt;The results are sufficiency conditions tied to signs of higher-order cross derivatives, which are hard to interpret and whose empirical signs are not firmly established (experimental evidence is insufficient). The model is a stylized two-date setup with zero interest rate, no time discounting, additive time-separable utility, interior unique optimum, and a single scalar correlation parameter. The authors note the framework extends straightforwardly to multi-period models and suggest studying settings where the value and uncertainty of correlation change over time.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Studying Generational Risk in a Large-Scale Life-Cycle Model</title><link>https://macropaperwarehouse.com/papers/studying-generational-risk-in-a-large-scale-life-cycle-model/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/studying-generational-risk-in-a-large-scale-life-cycle-model/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Hasanhodzic and Kotlikoff ask a question prior work assumed away: how large is generational risk, and can pay-go Social Security actually mitigate it? Earlier studies (Diamond, Bohn, Krueger-Kubler, etc.) presumed generational risk is large enough to merit policy and showed Social Security can in principle share it, but did not directly measure its size. This paper measures it directly, with and without Social Security, in a realistically large overlapping-generations (OLG) model.&lt;/p&gt;
&lt;p&gt;Model setup: an 80-period annual OLG model with aggregate shocks. Agents work 45 periods (retire at R=45) and live 80, have isoelastic (CRRA) preferences with risk aversion gamma=2 (gamma=5 under the extra-large shocks calibration), annual discount factor beta=0.96 (quarterly 0.99). Production is Cobb-Douglas; log TFP is trend-stationary AR(1) (quarterly rho=0.95, sigma=0.01; annualized rho=0.814, sigma=0.019). Two calibrations add a normal capital-depreciation shock. Households invest in risky capital or one-period safe bonds (zero net supply); &amp;ldquo;soft&amp;rdquo; increasing borrowing costs (Chen-Mangasarian function, slope b) shut down private risk-sharing to expose generational risk in its purest form while still delivering a realistic risk and growth premium. Policy is pay-go Social Security with a fixed payroll tax tau=15% (also tested at 1%). The model is solved to high precision via a projection method (building on Marcet 1988; Judd, Maliar, Maliar 2011) over an 81-variable state space (79 cohort cash-on-hand values plus the TFP and depreciation shocks). Generational risk measures are evaluated 300 years into the transition; cohort utility uses generations born after year 300 of a 750-year run. The U.S. data targets cover the return to national wealth and one-month Treasuries, 1947-2015, and detrended NNP/consumption, 1929-2020.&lt;/p&gt;
&lt;p&gt;Four calibrations: (1) baseline (TFP shock only, matched to output/consumption variability); (2) larger shocks (adds depreciation shock to match variability of the return to national wealth); (3) extra-large shocks (bigger depreciation shock to match U.S. equity-market return variability, a la Krueger-Kubler); (4) negative risk-free-rate baseline (steeper borrowing costs giving a roughly negative 2% safe rate, to test Blanchard 2019).&lt;/p&gt;
&lt;p&gt;Main findings (compensating-consumption differentials needed to reach long-run average lifetime utility): generational risk is 1.396% under baseline, 2.128% under larger shocks, and 15.303% under extra-large shocks (without Social Security). The authors view baseline 1.396% as small (on the order of a good-sized distortion) and prefer the baseline calibration. Social Security slightly WORSENS baseline generational risk (rising to 1.462%), but reduces it by 8% in the larger-shocks and 19% in the extra-large-shocks calibrations. So Social Security&amp;rsquo;s risk-pooling value depends on calibration. Contemporaneous risk (absolute consumption adjustment for full risk sharing among living cohorts) is tiny: 0.206% baseline, 0.933% larger shocks, 0.437% extra-large; Social Security raises it to 0.310% in baseline but lowers it under the other two.&lt;/p&gt;
&lt;p&gt;On welfare and Blanchard&amp;rsquo;s conjecture: pay-go Social Security at a 15% tax cuts long-run expected utility by 18% in baseline and larger-shocks, and by 56% in extra-large shocks, via crowding out (long-run capital falls 28% baseline, 56% extra-large). Under the negative-safe-rate calibration there is still an 18% long-run welfare loss; the average growth rate is zero in all simulations. The authors find no support for Blanchard&amp;rsquo;s (2019) claim that deficits can be Pareto-improving when safe rates run below growth: even under Blanchard-favorable conditions, crowding out swamps risk sharing (e.g., 17.83% utility loss at 15% tax, 1.17% at 1% tax). Macro shocks are second-order for policy: the capital transition under Social Security with shocks closely tracks the no-shock (deterministic) path, echoing Lucas (1987).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-exactly-is-the-papers-primary-measure-of-generational-risk"&gt;Q1. What exactly is the paper&amp;rsquo;s primary measure of generational risk?&lt;/h3&gt;
&lt;p&gt;It is the average absolute percentage adjustment to a cohort&amp;rsquo;s annual consumption needed to equate that cohort&amp;rsquo;s realized lifetime utility to the long-run cross-cohort average realized lifetime utility. Formally, for each generation born in period t they compute lambda_t = U-bar / U_t (U_t is realized lifetime utility, U-bar the average over generations born in years 301-750), then take the mean absolute deviation of lambda from 1. It captures both being born in a bad state and being hit by a bad sequence of lifetime shocks. A value near zero means birth date barely matters.&lt;/p&gt;
&lt;h3 id="q2-why-does-annualizing-to-80-periods-matter-relative-to-two-period-models"&gt;Q2. Why does annualizing to 80 periods matter relative to two-period models?&lt;/h3&gt;
&lt;p&gt;With one year per period, an agent experiences 45 annual wage shocks and 79 annual investment-return shocks that largely average out, and can self-insure by adjusting saving annually. In a two-period model a single negative TFP shock hits a worker&amp;rsquo;s entire lifetime earnings or a retiree&amp;rsquo;s whole old-age return. The authors note, however, that because TFP shocks are positively autocorrelated, amplifying multi-period shocks could in principle generate more risk, not less, so the result is not mechanical.&lt;/p&gt;
&lt;h3 id="q3-how-is-private-risk-sharing-handled-and-why-shut-it-down"&gt;Q3. How is private risk-sharing handled, and why shut it down?&lt;/h3&gt;
&lt;p&gt;In three of four calibrations the authors impose &amp;lsquo;soft&amp;rsquo; increasing borrowing costs (Chen-Mangasarian function, parameter b) calibrated so the marginal borrowing cost is 15-20 times the safe rate (b=28 baseline, 25 larger shocks, 45 for negative-safe-rate cases). This nearly closes the bond market, isolating generational risk with no private or public mitigation. The extra-large calibration omits borrowing costs because its large depreciation shock alone delivers a realistic risk premium (and to match Krueger-Kubler). Notably, adding borrowing constraints has little impact on key macro aggregates.&lt;/p&gt;
&lt;h3 id="q4-why-does-social-security-increase-generational-risk-in-the-baseline-single-tfp-shock-case"&gt;Q4. Why does Social Security INCREASE generational risk in the baseline (single-TFP-shock) case?&lt;/h3&gt;
&lt;p&gt;Five reasons given: (1) benefits depend on the prevailing wage, so autocorrelated TFP wage shocks now interact with capital-return shocks through retirement, extending nonlinear discounting past retirement; (2) crowding out lowers wages and raises risky returns, so the same percentage TFP shock is larger in absolute terms, making realized resources more variable; (3) Social Security is a random floor on old-age living standards, encouraging less risk-averse consumption and a higher propensity to consume; (4) positive TFP autocorrelation (high benefits today predict high benefits tomorrow) further raises the propensity to consume; (5) Social Security alters the stochastic distribution of the 79 cohort cash-on-hand state variables, producing complex consumption changes. This echoes Rios-Rull&amp;rsquo;s (1994) paradox that better micro insurance can amplify macro fluctuations.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-paper-test-blanchards-2019-deficits-may-be-free-conjecture-and-what-does-it-find"&gt;Q5. How does the paper test Blanchard&amp;rsquo;s (2019) &amp;lsquo;deficits may be free&amp;rsquo; conjecture and what does it find?&lt;/h3&gt;
&lt;p&gt;It uses Blanchard&amp;rsquo;s own ex-ante Pareto criterion but with 80 periods (vs his 2), realistic risk aversion, and dropping his assumption that half of wages are perfectly safe. Calibrations engineered with negative safe rates and large growth premiums (e.g. risky ~2%, safe ~negative 2%) still show Social Security reducing long-run expected utility: 17.83% loss at a 15% tax (1.17% at 1%) in the standard-premium case, falling to 12.51%/12.582% (15% tax) under even-larger growth premiums, but always negative. Crowding out dominates any risk-sharing gains. The authors find no support for the conjecture. They note Blanchard&amp;rsquo;s Pareto gains, when they arise, depend critically on his assumption that half of wages are certain, leaving workers ideally placed to insure the elderly.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-across-cohorts-is-documented"&gt;Q6. What heterogeneity across cohorts is documented?&lt;/h3&gt;
&lt;p&gt;Baseline generational risk has mean 1.396%, s.d. 1.293%, max 4.949% (no Social Security). Decomposed: generations with worst luck need roughly +5.0% positive adjustment; those with best luck need roughly negative 5.1%. Extra-large shocks produce extreme spread: max positive adjustment 66.14%, max negative 44.10%. A separate exercise (Table 8) shows the cost of uncertainty depends on birth state due to mean reversion: those born with low capital actually prefer uncertainty (negative 1.482%) because capital and wages will rise, while those born with high capital would pay 2.374% to lock in their state.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-welfare-cost-of-uncertainty-and-precautionary-saving-findings"&gt;Q7. What are the welfare-cost-of-uncertainty and precautionary-saving findings?&lt;/h3&gt;
&lt;p&gt;Under larger shocks, the compensating variation between the stochastic steady state and a no-shocks steady state is only 1.12% (newborns would need 1.12% more consumption each year to match a never-shocked long run), despite that calibration overstating macro variability. This is small because precautionary saving raises the stochastic economy&amp;rsquo;s average capital stock 18.4% above the no-shocks steady state: the uncertain long run is &amp;lsquo;riskier, but richer.&amp;rsquo; A decomposition removing the 0.77% average age-specific consumption difference leaves a 0.34% residual (about one quarter of 1.12%) reflecting age-pattern and cohort-sequence heterogeneity.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-build-on-and-differ-from-krueger-kubler-2006"&gt;Q8. How does this paper build on and differ from Krueger-Kubler (2006)?&lt;/h3&gt;
&lt;p&gt;Five differences: (1) many more periods (80 vs 9) permit better shock-averaging and more precise autocorrelation treatment plus more self-insurance opportunities; (2) two calibrations the authors view as more realistic than KK (who chose theirs partly to favor a Pareto improvement), using borrowing costs rather than excessively large depreciation shocks to get a realistic risk premium; (3) ex-ante rather than ex-interim expected utility; (4) explicit measurement of generational risk with and without Social Security; (5) testing whether a large growth premium can sustain an intergenerational Ponzi scheme at scale. Like KK, they find a negative net long-run welfare impact of pay-go Social Security.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-model-deliberately-omit-and-why"&gt;Q9. What does the model deliberately omit, and why?&lt;/h3&gt;
&lt;p&gt;It is &amp;lsquo;intentionally bare bones to maximize the potential for generational risk&amp;rsquo;: no variable labor supply (which would help cohorts self-insure), no progressive income taxation (which redistributes from winning to losing generations), and no social insurance other than Social Security. It also omits capital-adjustment costs (which would raise asset-return volatility) because incomplete markets make firm investment policy ill-defined when differently-aged shareholders disagree; the depreciation shock is a crude proxy for adjustment-cost-driven asset-return shocks. The authors flag correlated idiosyncratic shocks (Harenberg-Ludwig) as important future work.&lt;/p&gt;
&lt;h3 id="q10-how-well-does-each-calibration-match-the-data"&gt;Q10. How well does each calibration match the data?&lt;/h3&gt;
&lt;p&gt;Baseline matches output (model 3.72% vs data 3.33%) and consumption (2.10% vs 1.75%) variability but understates the s.d. of the return to national wealth by an order of magnitude (0.14% vs 4.89%). Larger shocks reproduces the return-to-wealth s.d. (4.61-4.62% vs 4.89%) and a realistic wage/return correlation (negative 0.054) but overstates macro-aggregate variability. Extra-large shocks matches equity Sharpe ratio (model 0.333 vs target 0.286; risk premium 4.63%, return s.d. 13.92%) but overstates return-to-capital variability nearly three-fold and consumption variability sixteen-fold. The model&amp;rsquo;s overall risk premium ranges 3.55-6.03% vs 5.43% in data.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-role-of-the-bond-market-across-calibrations"&gt;Q11. What is the role of the bond market across calibrations?&lt;/h3&gt;
&lt;p&gt;The one-period bond market only operates in the extra-large shocks calibration (borrowing costs close it in the others). There, the young short bonds and the old lend: because the young&amp;rsquo;s resources are mostly human capital (less risky than, and negatively correlated with, stock returns), the young use bonds to insure the old. Workers effectively borrow to hold equity, which the authors rationalize via student loans, credit cards, mortgages alongside 401(k) equity, or implicit long-term firm contracts.&lt;/p&gt;
&lt;h3 id="q12-what-policy-implications-follow-and-what-are-their-scope-conditions"&gt;Q12. What policy implications follow, and what are their scope conditions?&lt;/h3&gt;
&lt;p&gt;If macro shocks are calibrated to realistic macro-aggregate volatility (the authors&amp;rsquo; preferred baseline), generational risk is small (about 1.4%) and pay-go Social Security slightly worsens it while imposing an 18% long-run welfare loss via crowding out; deterministic models (e.g. Auerbach-Kotlikoff 1987) then suffice to capture the long-run impact of intergenerational redistribution. Social Security&amp;rsquo;s risk-mitigation value emerges only under calibrations that overstate macro volatility (larger/extra-large shocks). The scope condition is decisive: the case for Social Security as generational insurance hinges on which calibration one finds realistic, and the authors&amp;rsquo; preferred reading implies a weak case. They also caution the conclusions may not extend to models with correlated idiosyncratic risk.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Asset Exemption in Bankruptcy, Access to and Cost of Credit</title><link>https://macropaperwarehouse.com/papers/asset-exemption-in-bankruptcy-access-to-and-cost-of-credit/</link><pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/asset-exemption-in-bankruptcy-access-to-and-cost-of-credit/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Under U.S. Chapter 7 bankruptcy, an individual entrepreneur has most unsecured debt discharged and only her non-exempt assets liquidated, producing an &amp;ldquo;insurance effect.&amp;rdquo; But this protection does not extend to assets voluntarily pledged as collateral, so a borrower can undo the insurance by posting sufficient collateral. The paper asks how asset exemption interacts with the decision to post collateral to shape access to and the cost of credit. The novel insight is that, because the opportunity cost of pledging collateral (forgoing the exempt assets one would otherwise keep in default) is lower for safe entrepreneurs than for risky ones, collateral becomes a more effective sorting device as exemption rises. Existing empirical work (Gropp et al. 1997; Berkowitz and White 2004; Berger et al. 2011) finds exemption reduces access and raises rates, but does not exploit the interaction between collateral and exemption.&lt;/p&gt;
&lt;p&gt;Model setup: A competitive credit market with risk-neutral entrepreneurs heterogeneous in success probability (safe type-H with pH, risky type-L with pL, pH &amp;gt; pL) and in pledgeable wealth w over [w, w-bar]. Each needs one unit of credit; lenders face opportunity cost r and cannot observe type. Lending contracts are triples (cost of credit RB, collateral C, access probability pi). Exemption eta shields wealth up to eta from liquidation but not wealth posted as collateral; liquidated wealth is worth only lambda &amp;lt; 1 to lenders. Competition is modeled as a three-stage game (a la Hellwig 1997) so that a subgame-perfect equilibrium exists and delivers the contract most preferred by safe types. The setup extends Besanko and Thakor (1987) by allowing any exemption between zero and infinity, adding the third (acceptance) stage, and adding wealth heterogeneity.&lt;/p&gt;
&lt;p&gt;Main theoretical results: With zero exemption, pooling is the only equilibrium and no rationing occurs. With positive exemption, the equilibrium involves separation (at least for intermediate wealth): safe entrepreneurs self-select into contracts with effective collateral and face a lower cost of credit, while risky ones post no collateral. As in Besanko and Thakor, separation entails rationing for safe entrepreneurs too wealth-constrained to meet collateral requirements. The key novelty: conditional on posting collateral, as exemption rises, access to credit rises and the cost of credit falls—collateral becomes a more powerful screening tool. The overall effect of higher exemption on aggregate rationing is ambiguous, because more safe entrepreneurs choose to separate (lowering their access probability) even as each separating safe type is rationed less; the net effect depends on the wealth distribution.&lt;/p&gt;
&lt;p&gt;Data and empirical strategy: The 2003 wave of the Survey of Small Business Finances (SSBF), 4240 firms, restricted to 1761 creditworthy firms that were financed at least once (96% always financed). Cross-state exemption variation is collapsed to a high/low dummy across nine census divisions (West North Central and West South Central coded high). Firm type is identified by whether it posts collateral (posters = type-H). An endogenous switching / inverse Mills ratio approach (Maddala 1983) handles self-selection in the cost-of-credit equation; access to credit is estimated by probit with a collateral-by-exemption interaction.&lt;/p&gt;
&lt;p&gt;Main quantitative findings: Descriptively, high-asset firms face loan rates 1.5 pp lower and rationing 3.8 pp lower. Collateral-posting firms pay 0.7 pp lower rates overall; this differential grows from 0.53% in low-exemption to 1.20% in high-exemption subsamples. The Mills-ratio coefficients are negative and significant, confirming collateral conveys private information. In the access regression, posting collateral is positively associated with rationing, but firms posting collateral are less likely to be rationed in high-exemption divisions (predicted access falls 0.6% on average from posting collateral, but rises 1.5% in high-exemption areas). Reduced-form OLS: collateral firms pay 0.30% less, with the discount rising 0.55% moving low-to-high exemption. The simultaneous structural system implies a 34-basis-point average reduction in cost of credit from guarantees, three times larger in high-exemption states (75 vs 17 bp). Heckman selection correction does not alter conclusions. All main model predictions cannot be rejected.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Identification rests on three pillars. (1) Firm type is identified by the collateral decision: the model implies only type-H (safe) firms post collateral, so posters are treated as type-H and non-posters as type-L. (2) Cross-sectional variation in asset exemption across census divisions (a high/low dummy, with West North Central and West South Central coded high) provides exogenous variation in the strength of collateral as a sorting device. (3) The cost-of-credit equation uses an endogenous switching model (Maddala 1983) identified by the non-linearity of the inverse Mills ratio, under the model-based assumption that observed loan rates are determined by the endogenous collateral decision. Threats: (a) Selection bias from restricting to creditworthy/financed firms—addressed with a Heckman selection model that leaves conclusions unchanged. (b) Coarse exemption measurement—location is only observed at the nine-census-division level rather than by state, and unlimited-exemption states must be aggregated, so the high/low dummy is a proxy; an alternative averaging procedure is reported to give the same results. (c) SSBF data are partly imputed; estimates use Rubin (1987) multiple-imputation combination rules (STATA mi estimate), which inflates variance and can reduce significance.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The central mechanism is the opportunity cost of posting collateral: in default a borrower who pledged assets loses them all, whereas without pledging she would keep the exempt part. This opportunity cost rises with exemption and is lower for safe borrowers (lower default probability), so collateral sorts types more sharply as exemption rises. Empirically this is distinguished through the collateral-by-exemption interaction: the cost-of-credit discount from posting collateral, and the access-to-credit advantage of posters, both should strengthen with exemption. The negative, significant inverse Mills ratio coefficients show the collateral choice reveals private information about type; the estimated lambda_1L,v being roughly double lambda_1H,v indicates safe firms choose contracts with lower cost-of-credit variance.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;By wealth: high-asset firms face rates 1.5 pp and rationing 3.8 pp lower. The collateral cost discount is concentrated among low-asset firms (0.9 pp) versus high-asset firms (0.04%). The collateral-rationing association also depends on wealth: among low-asset firms, rationing is 4.4% higher for collateral posters, but for high-asset firms there is no difference. By exemption: the collateral cost differential grows from 0.53% (low) to 1.20% (high). Among collateral posters, the rationed fraction falls 1.1% moving low-to-high exemption, with a larger drop for low-asset firms (-1.9%) than high-asset firms (-0.5%). In the structural cost-of-credit table, wealth reduces the cost of credit for non-posters only in high-exemption areas and for posters only outside high-exemption areas—consistent with firms undoing exemption via collateral.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Three. (1) A reduced-form OLS loan-rate regression with collateral, exemption, and their interaction confirms posters pay less (about 0.30% on average) and the discount grows 0.55% moving to high exemption; signs match predictions (beta_3 &amp;lt; 0, beta_4 &amp;lt; 0, beta_2 &amp;gt; 0). (2) A simultaneous structural two-equation system jointly determining cost of credit and guarantees yields a 34-bp average reduction in cost from guarantees, three times larger in high-exemption states (75 vs 17 bp). (3) A Heckman-style selection model accounting for the application/creditworthiness/financing stages leaves all conclusions intact. The imputation-robust (mi estimate) procedure is also applied throughout.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It confirms Gropp et al. (1997), Berkowitz and White (2004), and Berger et al. (2011) that higher exemption raises both rationing and the cost of credit. Its contribution is to use the theoretical model as an identification tool for the joint, interactive effect of exemption and the collateral decision—a prediction absent in prior empirical work. The collateral-as-quality-signal interpretation aligns with Jimenez et al. (2006) for Spanish firms and with Berger et al. (2011) on ex ante asymmetric information. Theoretically, it complements Manove et al. (2001) (too little exemption induces lazy bank screening) by showing that lower creditor protection via exemption gives lenders incentive to screen with collateral. It differs from Krasa et al. (2008) and Tamayo (2015), where creditor protection is an exogenous fraction of retained assets; here that fraction is endogenous because collateral can undo exemption. The model setup extends Besanko and Thakor (1987) with arbitrary exemption levels, a third acceptance stage (Hellwig 1997), and wealth heterogeneity.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Asset exemption levels materially affect credit-market functioning. Positive exemption lowers access and raises the cost of credit on average. But raising exemption enhances collateral&amp;rsquo;s power as a sorting device, so safe entrepreneurs who signal by posting collateral gain better access and larger rate discounts as exemption rises. The net effect of higher exemption on aggregate credit rationing is ambiguous and depends on how collateralizable wealth is distributed across entrepreneurs: more safe types separate (each facing a lower access probability) even as each separating safe type is rationed less. Scope conditions: results apply to individual entrepreneurs under Chapter 7 where exemption does not protect pledged collateral; the insurance/opportunity-cost channel requires exemption to be non-zero (at zero exemption only pooling, no rationing, and collateral conveys no signal); and the empirical magnitudes are estimated for small U.S. firms financed at least once in 2001-2003.&lt;/p&gt;
&lt;h3 id="q7-what-are-notable-caveats-and-data-limitations"&gt;Q7. What are notable caveats and data limitations?&lt;/h3&gt;
&lt;p&gt;The dataset does not record the amount of collateral posted, only whether collateral was posted, so type is inferred from a binary decision. Firm location is observed only at the nine-census-division level, forcing a coarse high/low exemption dummy rather than state-level variation. The sample is restricted to firms financed at least once, raising selection concerns (addressed via Heckman). Much SSBF data are imputed. The model abstracts from positive, non-negligible transaction costs of posting collateral (only a negligible cost is assumed to select the unique separating equilibrium with CL = 0); incorporating such costs is left as an extension.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Insurance effect (of exemption and discharge)&lt;/strong&gt;: The protection an entrepreneur enjoys under Chapter 7 because most unsecured debt is discharged and only non-exempt assets are liquidated; in the paper this protection can be voluntarily undone by posting assets as collateral.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Opportunity cost of posting collateral&lt;/strong&gt;: The exempt wealth a borrower forgoes by pledging assets: in default a collateral-poster loses everything pledged, whereas a non-poster keeps the exempt part. This cost rises with the exemption level and is lower for safe (low-default-probability) entrepreneurs, making collateral an informative sorting device.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Real guarantees (G)&lt;/strong&gt;: The effective amount of wealth a lender can actually recover in default, G = max(min(w_eta, RB/lambda), C): increasing in collateral C and decreasing in exemption eta. The model is stated in terms of guarantees rather than nominal collateral.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Separating vs. pooling equilibrium&lt;/strong&gt;: Under positive exemption, safe entrepreneurs self-select into high-guarantee, lower-rate (possibly rationed) contracts while risky ones take no-collateral contracts (separation); under zero exemption all borrow under one contract with no rationing (pooling). The model selects the subgame-perfect outcome most preferred by safe types.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Type-H / type-L identification via collateral&lt;/strong&gt;: The empirical convention, derived from the model, that firms posting collateral are safe (type-H) and those not posting are risky (type-L), since in equilibrium only safe firms post collateral.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous switching / inverse Mills ratio approach&lt;/strong&gt;: The estimation method (Maddala 1983) that corrects for self-selection in the collateral decision; negative, significant Mills-ratio coefficients indicate collateral posting conveys private information lowering the cost of credit, identified by the Mills ratio&amp;rsquo;s non-linearity.&lt;/p&gt;</description></item></channel></rss>