<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>The Economic Journal | Macro Paper Warehouse</title><link>https://macropaperwarehouse.com/journal/the-economic-journal/</link><atom:link href="https://macropaperwarehouse.com/journal/the-economic-journal/index.xml" rel="self" type="application/rss+xml"/><description>The Economic Journal</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Jan 2026 00:00:00 +0000</lastBuildDate><item><title>A Macro Study of the Unequal Effects of Climate Change</title><link>https://macropaperwarehouse.com/papers/a-macro-study-of-the-unequal-effects-of-climate-change/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/a-macro-study-of-the-unequal-effects-of-climate-change/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper develops a macro heterogeneous-agent model to quantify the distributional welfare impacts of higher temperatures from climate change across income groups in the United States. The motivation is that existing macro climate-economy models either abstract from heterogeneity entirely or focus on spatial heterogeneity across regions rather than income heterogeneity within regions. The paper fills this gap by modeling how the welfare consequences of temperature change depend on both the region a household lives in and its position in the income distribution.&lt;/p&gt;
&lt;p&gt;The model is calibrated to the US using five data sources: NIPA accounts from the BEA (averaged 1997–2020), the 2015 Residential Energy and Consumption Survey (RECS), PRISM climate data (1950–2022), a proprietary product-level data set of over 1,000 heaters, air conditioners, and heat pumps scraped from ecomfort.com in fall 2023, and county-level climate projections for year 2100 under RCP 8.5 from Rasmussen et al. (2016). The US is divided into five regions (cold, cool, mild, warm, and hot) of approximately equal population based on average county temperature. The quantitative exercise compares two stationary equilibria: a contemporary equilibrium using the current temperature distribution and a climate-change equilibrium using the projected 2100 distribution under RCP 8.5 (a no-large-scale-climate-policy scenario). Welfare is measured using the consumption-housing equivalent variation (CHEV), defined as the percent increase in consumption and housing a household would require in every period in the contemporary equilibrium to be indifferent between the two equilibria.&lt;/p&gt;
&lt;p&gt;Households adapt to temperature through two channels: an intensive margin (adjusting energy use for heating and cooling given existing equipment) and an extensive margin (deciding whether to purchase a heater, air conditioner, or heat pump, each carrying a fixed cost). The production functions for heating and cooling are estimated by OLS on the product-level data set, yielding equipment exponents of 0.35 (air conditioners), 0.28 (heaters), and 0.27 (heat pumps), and energy exponents of 0.77, 0.86, and 0.85, respectively, with R-squared values of 0.97, 0.79, and 1.00. A key analytical insight from a stylized model is that the outdoor temperature acts as a &amp;ldquo;transfer from nature&amp;rdquo; to households — warmer days in cold weather and cooler days in hot weather reduce the energy households must purchase, augmenting real income. Because this transfer is a larger share of income for lower-income households, its changes are distributionally regressive when the transfer falls (hotter regions warming further) and progressive when it rises (colder regions warming).&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. Among middle- and high-income households, climate change generates progressive welfare gains in colder regions — ranging from +0.71 percent of consumption-and-housing for households in the third income decile in the cool region to near-zero for the highest income households — and regressive welfare losses in hotter regions, ranging from −1.85 percent for third-decile households in the warm region to near-zero for high-income households. These patterns are driven by the intensive margin (changes in transfers from nature). For low-income households, the pattern reverses: low-income households in colder regions suffer welfare losses (the dominant effect is that climate change forces them to purchase their first air conditioner), while some low-income households in hotter regions experience welfare gains (they can forgo purchasing a heater). Climate change raises the Gini coefficient on lifetime welfare by 1.02, 1.01, and 0.50 percent in the cold, cool, and mild regions, and reduces it by 0.09 and 0.21 percent in the warm and hot regions. Aggregate welfare effects from the heterogeneous-agent model substantially exceed what a representative-agent model would imply: for example, in the mild region, climate change reduces aggregate welfare by 0.65 percent in the baseline but only 0.17 percent in the representative-agent version.&lt;/p&gt;
&lt;p&gt;Policy experiments reveal: (1) Fully offsetting the welfare costs of climate change for the lowest-income households would require government spending on energy assistance to more than double (a factor of 2.2 increase), with the largest increases concentrated in colder regions. (2) A universal heat-pump mandate eliminates the extensive-margin channel, producing monotonically progressive welfare gains in colder regions and monotonically regressive welfare losses in hotter regions across all income deciles. (3) Heat-pump cost parity with heaters largely increases adoption and moderates welfare costs, but low-income households in the hot region see limited improvement because they still prefer air conditioners. (4) Accounting for temperature effects on the labor productivity of outdoor workers (roughly 8 percent of the workforce, concentrated at lower incomes) amplifies welfare costs in hotter regions and moderates them in colder regions, with magnitudes tied to the share of workers affected.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper is a calibrated structural model rather than an empirical identification exercise. Identification in the sense of parameter estimation comes from two sources: (1) OLS estimation of heating and cooling production functions on cross-sectional product-level data, where manufacturers measure capacity and efficiency under standardized conditions, limiting TFP endogeneity concerns that plague aggregate production function estimation; and (2) internal calibration of remaining parameters to match a set of moments from RECS 2015 and NIPA. Threats to the structural analysis include the assumption that households treat housing and equipment as flow (rental) choices rather than durable stocks, abstracting from switching costs and adjustment costs over the transition — the paper explicitly notes this limits the analysis to long-run stationary equilibria. The small-open-economy assumption for capital removes domestic capital-market clearing as a constraint. The calibration uses 2015 RECS (not 2020) to avoid COVID-19 distortions to cooling budget shares. The paper abstracts from amenity values of outdoor temperature, mortality from temperature exposure (approximately 0.04 percent of US deaths from 1999–2020), and spatial migration responses.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-core-mechanisms-and-how-are-they-distinguished"&gt;Q2. What are the two core mechanisms and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;The two mechanisms are the intensive margin (how much energy to use given existing equipment) and the extensive margin (whether to purchase heating or cooling equipment at all). The paper distinguishes them analytically using the simple model, which isolates the intensive margin by assuming all households have equipment. The intuition from the simple model — outdoor temperature as a transfer from nature — explains why welfare effects are progressive in regions where climate change makes temperatures more moderate (transfers rise) and regressive where temperatures become more extreme (transfers fall). The extensive margin is then added in the quantitative model through fixed costs of heater, air conditioner, and heat pump equipment. The paper shows that climate change affects specialization favorability (the degree to which a temperature distribution favors concentrating on only heating or only cooling equipment), and that this extensive-margin channel is most important for lower-income households who are near a corner solution of specializing in only one type of equipment. The heat-pump-mandate counterfactual is used to isolate the intensive-margin channel: when all households use heat pumps in both equilibria, the extensive-margin decision is unchanged by climate change, and all welfare effects are driven purely by transfers from nature.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented-across-income-groups-and-regions"&gt;Q3. What heterogeneity is documented across income groups and regions?&lt;/h3&gt;
&lt;p&gt;Welfare effects vary dramatically in both sign and magnitude. Among middle- and high-income households, climate change generates progressive welfare gains in colder regions (e.g., +0.71 percent CHEV for third-decile households in the cool region, falling toward zero at the top) and regressive welfare losses in hotter regions (e.g., −1.85 percent CHEV for third-decile households in the warm region, again near-zero at the top). For low-income households, the pattern reverses: they experience welfare losses in colder regions (forced to buy first air conditioner) and welfare gains or smaller losses in hotter regions (can forgo purchasing a heater). Figure 2 in the paper shows these crossing patterns by income decile for all five regions simultaneously. The Gini coefficient changes by +1.02% (cold), +1.01% (cool), +0.50% (mild), −0.09% (warm), and −0.21% (hot). Migration incentives also differ: high-income households gain incentives to move to cooler regions (driven by transfers from nature), while low-income households gain incentives to move to warmer regions (driven by specialization changes).&lt;/p&gt;
&lt;h3 id="q4-what-is-the-transfers-from-nature-concept-and-why-does-it-produce-differential-welfare-effects"&gt;Q4. What is the &amp;rsquo;transfers from nature&amp;rsquo; concept and why does it produce differential welfare effects?&lt;/h3&gt;
&lt;p&gt;The paper formalizes the idea that outdoor temperature provides free heating or cooling that substitutes for costly purchased energy. On a cold day with outdoor temperature ζ, nature provides ζ degrees of heating for free, effectively augmenting household income by p_eh * ζ (the value of that heating at market prices). This transfer is identical in absolute terms for all households regardless of income, but it is a larger fraction of income for low-income households, so its loss or gain has greater proportional welfare impact on them. This parallels the progressivity of lump-sum transfers in public finance: losing a dollar matters more when income is lower. Consequently, when climate change moves a region to more moderate temperatures (colder regions), the resulting increase in transfers from nature is progressive — lower-income households gain proportionally more. When climate change moves a region to more extreme temperatures (hotter regions), the decrease in transfers is regressive — lower-income households lose proportionally more. The amenity value of outdoor temperature (distinct from the heating/cooling transfer) is abstracted from in the quantitative model on the grounds that, per the simple model, it does not affect the cross-income distribution of welfare changes if preferences over amenities are uncorrelated with income.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-extensive-margin-generate-the-reversal-of-welfare-effects-for-low-income-households"&gt;Q5. How does the extensive margin generate the reversal of welfare effects for low-income households?&lt;/h3&gt;
&lt;p&gt;The extensive margin works through what the paper calls &amp;lsquo;specialization favorability.&amp;rsquo; When a temperature distribution is dominated by cold days, households can optimally purchase only heater equipment, avoiding the additional fixed cost of an air conditioner; the reverse holds in hot climates. Climate change reduces the specialization favorability index in colder regions by adding more hot days, and increases it in hotter regions by reducing cold days. The welfare impact of moving between a corner solution (one type of equipment) and an interior solution (two types of equipment, or a heat pump) tends to be larger than moving between two interior solutions. In the cold region, climate change causes the majority of households in the bottom three income deciles to transition from not having air conditioning to having it (Figure 5, left panel). The fixed cost of buying an air conditioner for the first time exceeds the intensive-margin gains from more moderate temperatures, producing net welfare losses. In the hot region, many second-through-fourth decile households move from having heat in the contemporary equilibrium to not having heat in the climate-change equilibrium (Figure 5, right panel), saving the fixed cost and producing net welfare gains despite more extreme temperatures.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-model-calibrated-and-what-is-the-quality-of-fit"&gt;Q6. How is the model calibrated and what is the quality of fit?&lt;/h3&gt;
&lt;p&gt;Externally calibrated parameters include: capital income share α = 0.26 (Kiyotaki et al., 2011), depreciation rate δ = 0.066, interest rate r* = 0.04, CRRA coefficient σ = 2, bliss point temperature ζ* = 18°C, labor productivity process (ρ = 0.97, σ²_ε = 0.02, σ²_ξ = 0.66 from Kaplan, 2012), and production function exponents estimated from the ecomfort.com data. Internally calibrated parameters are jointly chosen to match: wealth-to-output ratio (3.0), housing-to-non-housing capital ratio (0.88), average heating budget share for non-heat-pump households (0.014), average cooling budget share (0.0055), energy budget share for heat-pump households (0.014), fractions of households with heating (0.95), cooling (0.86), and heat pumps (0.09), the ratio of energy budget shares between the fifth and first income quintile (0.12), the ratio of energy expenditures between high and low income (1.72), and energy assistance as a fraction of energy expenditures (0.83). Table 3 shows the model matches all targeted moments closely. External validation (untargeted moments) shows the model also replicates the associations between heating/cooling degree days and budget shares, equipment ownership, and indoor temperature choices, with similar signs and magnitudes to RECS 2015 data. One limitation is that the model overstates heat pump adoption (17% in model vs. 9% in 2015 RECS, though 14% in 2020 RECS), because it treats modern cold-weather-capable heat pumps as the default.&lt;/p&gt;
&lt;h3 id="q7-what-do-the-policy-counterfactuals-show"&gt;Q7. What do the policy counterfactuals show?&lt;/h3&gt;
&lt;p&gt;Four policy experiments are analyzed. First, scaling energy assistance proportionally to energy needs under climate change reduces assistance by 24% in cold and 20% in cool regions (where transfers from nature increase) and raises it by 9%, 36%, and 79% in mild, warm, and hot regions. Government spending increases by 25%, but the program remains smaller than 0.02% of output. This scaling partially offsets but does not eliminate the distributional distortions. Fully eliminating welfare costs for the lowest-income households would require multiplying energy assistance spending by a factor of 2.2. Second, a universal heat-pump mandate (analogous to natural gas bans like New York, Washington DC, or California&amp;rsquo;s post-2030 ban on natural gas furnaces) eliminates all extensive-margin effects because all households hold heat pumps in both equilibria. Under this mandate, climate change produces monotonically progressive welfare gains across all income groups in colder regions and monotonically regressive welfare costs in hotter regions. Third, heat-pump cost parity with heaters drives near-universal heat pump adoption and broadly moderates welfare costs relative to baseline, but the lowest-income households in the hot region see limited improvement because they still prefer air conditioners over heat pumps even at cost parity (air conditioners are cheaper and heat pumps&amp;rsquo; heating advantage is less valuable in an already-hot, increasingly-hotter climate). Fourth, the labor productivity extension (using the Richardson construction cost database adjustment factor of 1% per degree outside 40°F–85°F) implies that climate change raises low-income productivity by 2% in cold and 0.9% in cool regions and reduces it by 0.1%, 1.1%, and 2.2% in mild, warm, and hot regions. These labor-productivity changes modestly moderate welfare costs in colder regions and amplify them in hotter regions for low-income households.&lt;/p&gt;
&lt;h3 id="q8-why-does-income-heterogeneity-matter-for-aggregate-welfare-calculations"&gt;Q8. Why does income heterogeneity matter for aggregate welfare calculations?&lt;/h3&gt;
&lt;p&gt;The paper demonstrates that a representative-agent model substantially underestimates the aggregate welfare cost of climate change in all regions except the hot region. In the cold region, the aggregate CHEV is −1.03% in the baseline but the average (seventh-decile) household experiences small positive welfare effects (+0.19%), and the representative-agent model yields −0.00%. In the mild region, the aggregate is −0.65% but the representative-agent model gives −0.17%. The discrepancy arises because the welfare distribution is skewed: large losses for low-income households in colder regions are not offset by small or negative gains for high-income households, so the average is dominated by the tails. In the hot region the direction reverses: the baseline aggregate benefit (+0.24%) is driven by large gains at the bottom that the representative-agent model (−0.43%) misses entirely. This finding parallels the broader macroeconomics literature showing that income heterogeneity affects the aggregate welfare cost of business cycles, inflation, and asset pricing.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q9. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;The paper sits at the intersection of two literatures. The macro climate-economy literature (Acemoglu et al., 2012; Golosov et al., 2014; Barrage, 2020) typically uses representative-agent models that abstract from heterogeneity. The spatial heterogeneity literature (Cruz and Rossi-Hansberg, 2024; Bilal and Rossi-Hansberg, 2023; Rudik et al., 2022) studies how welfare consequences vary across regions based on their income levels and exposures but not within-region income differences. The within-region inequality literature (Dennig et al., 2015; Kornek et al., 2021; Belfori and Macera, 2022; Douenne et al., 2023) adds heterogeneous fixed income types to integrated assessment models, but does not model endogenous income and wealth distributions. Blanz (2023) is the closest precursor: it uses a standard incomplete-markets model to study food-price effects of climate change in developing countries, but does not model the temperature-equipment-energy production technology. The empirical literature (Hsiang et al., 2017; Park et al., 2018; Doremus et al., 2022) estimates reduced-form relationships between temperature and energy spending by income group, but cannot decompose intensive vs. extensive margin mechanisms or conduct structural policy counterfactuals. The key novel contributions are: (1) endogenous income and wealth heterogeneity within the Bewley-Huggett-Aiyagari tradition, (2) explicit modeling of both margins of temperature adaptation with estimated production functions, and (3) the ability to separately identify the roles of transfers from nature and specialization favorability.&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-are-conducted"&gt;Q10. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The paper reports several robustness checks. First, the main calibration uses the housing exponent γ = 0.1, but Appendix Figure D.1 shows results with γ = 0.4 (the upper bound implied by the RECS regression of energy on square footage, before controlling for quality), finding broadly similar qualitative results. Second, the 2015 RECS is used instead of the 2020 RECS due to COVID-19 distortions to cooling budget shares; the paper notes heating budget shares are similar between the two surveys while cooling shares are materially higher in 2020. Third, external validation of the model on untargeted moments (associations between HDD/CDD and heating/cooling budget shares, equipment ownership, and indoor temperatures) confirms the model&amp;rsquo;s predictive validity. Fourth, the welfare results are computed for both the main five-region model and a representative-agent version, documenting the magnitude of the aggregation bias. Fifth, the labor productivity extension bounds the relevant population (bottom 3% vs. bottom 16% of workers) to bracket the Occupational Requirements Survey estimate of 8% of workers constantly or frequently exposed outdoors.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-scope-conditions-and-limitations-of-the-main-results"&gt;Q11. What are the scope conditions and limitations of the main results?&lt;/h3&gt;
&lt;p&gt;Several important scope conditions apply. The analysis focuses exclusively on the direct effects of higher temperatures in the US; it does not cover other forms of climate damage (sea level rise, storm frequency, drought, wildfire) or effects in other countries. The model is solved for stationary equilibria, so it cannot speak to transition dynamics or the welfare costs of adjustment during the period when households are switching equipment. Housing and equipment are modeled as flow (rental) choices, abstracting from switching costs, adjustment frictions, and the interaction between homeownership and equipment decisions. The model abstracts from the amenity value of outdoor temperature (e.g., preference for pleasant weather), temperature-related mortality (about 0.04% of US deaths, 1999–2020, heavily concentrated among the unhoused population outside the model), and behavioral adaptation beyond energy and equipment choices (migration is analyzed only as a partial equilibrium incentive calculation, not as an equilibrium outcome). The capital market operates as a small open economy, so general equilibrium effects on interest rates are absent. Labor productivity effects of temperature are only explored for low-income workers in the outdoor sector, not for higher-income or indoor workers.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-migration-findings-and-their-caveats"&gt;Q12. What are the migration findings and their caveats?&lt;/h3&gt;
&lt;p&gt;The paper shows that climate change increases incentives for high-income households to migrate to cooler regions (driven by the transfers-from-nature channel — cooler regions offer larger increases in transfers) and increases incentives for low-income households to migrate to warmer regions (driven by the specialization channel — warmer regions allow forgoing heater equipment). The magnitude of the change in migratory pressure for high-income households is much smaller (order of magnitude roughly 0.15 on the paper&amp;rsquo;s scale) than for low-income households (order of magnitude roughly 3 on the same scale). The authors explicitly caveat that this is a partial equilibrium exercise: the model abstracts from the amenity value of temperature (which would reduce pressure to move to warmer regions by reducing the attractiveness of hot destinations) and from other dimensions of climate change (storm risk, fire risk) that would affect migration incentives independently.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Transfers from nature&lt;/strong&gt;: In this paper&amp;rsquo;s framework, outdoor temperature acts as a subsidy equivalent to income: on a cold day, nature provides degrees of heating for free, augmenting household real income by the value of that heating energy; on a hot day, it provides degrees of cooling. The transfer is the same in absolute terms for all households but represents a larger fraction of income for lower-income households, making changes in temperature distributionally progressive (when transfers rise) or regressive (when transfers fall).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extensive margin of temperature adaptation&lt;/strong&gt;: The binary decision of whether to purchase temperature-control equipment — a heater, air conditioner, or heat pump — each carrying a fixed cost. Households at the extensive margin may optimally forego one type of equipment entirely (complete specialization), and climate change can force them to acquire equipment they previously lacked or allow them to drop equipment they previously held.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive margin of temperature adaptation&lt;/strong&gt;: The continuous decision of how much energy to purchase to operate existing heating and cooling equipment in order to achieve a desired indoor temperature, conditional on having that equipment. Changes in the outdoor temperature distribution affect energy expenditures along this margin for all households that already own equipment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Specialization favorability index&lt;/strong&gt;: A region-level index S_n ∈ [0,1] defined as the absolute difference between total degrees of heating need and total degrees of cooling need, divided by their sum. Higher values indicate that the temperature distribution is more dominated by either heating or cooling demand, making it more efficient for households to specialize in a single type of temperature-control equipment rather than purchasing both. Climate change reduces specialization favorability in colder regions and increases it in hotter regions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption-housing equivalent variation (CHEV)&lt;/strong&gt;: The paper&amp;rsquo;s welfare metric: the percentage by which a household&amp;rsquo;s consumption and housing would need to increase in every period of the contemporary equilibrium for the household to be indifferent between remaining in the contemporary equilibrium and living in the climate-change equilibrium. Negative CHEV values indicate welfare losses from climate change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Temperature damage function D(T)&lt;/strong&gt;: A function mapping the deviation of indoor temperature from the bliss point to the fraction of full utility the household receives from housing services. D equals 1 when indoor temperature equals the bliss point (18°C in calibration) and falls below 1 as indoor temperature deviates in either direction, with the rate of decline governed by parameter χ. This function creates the motive to use energy for heating and cooling.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;RCP 8.5&lt;/strong&gt;: As used in this paper, a climate scenario from the CMIP archive representing emissions in the absence of large-scale climate policy, used to construct the 2100 temperature distribution in the climate-change equilibrium. County-level projections come from Rasmussen et al. (2016), probability-weighted across climate models.&lt;/p&gt;</description></item><item><title>Are Targeted Matching Schemes Effective in Stimulating Retirement Savings?</title><link>https://macropaperwarehouse.com/papers/are-targeted-matching-schemes-effective-in-stimulating-retirement-savings/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/are-targeted-matching-schemes-effective-in-stimulating-retirement-savings/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Governments across ten-plus countries — including Australia, the United States, Germany, and New Zealand — have introduced matching schemes to encourage low- and middle-income earners to contribute voluntarily to private pensions, motivated by the concern that progressive tax systems give these groups weaker incentives to save for retirement than high-income earners. Whether such schemes actually raise retirement savings is theoretically ambiguous: by reducing the cost of contributing they produce a substitution effect favoring more contributions, but the government payment also raises anticipated retirement income, reducing the desire to save further (a retirement income effect). The sign of the net effect depends on the distribution of contributions that would have occurred in the scheme&amp;rsquo;s absence, and it is especially unclear for those who would already have contributed above the matching ceiling.&lt;/p&gt;
&lt;p&gt;This paper tests the full set of theoretical predictions from a two-period intertemporal savings model using Australia&amp;rsquo;s Superannuation Co-contribution Scheme as a clean natural experiment. The scheme matches personal after-tax superannuation contributions up to $1,000 per year at a single, flat matching rate that varied over time — 100% in 2003-04 and 2009-10 to 2011-12, 150% in 2004-05 to 2008-09, and 50% from 2012-13 onward — and eligibility is phased out smoothly with income (no sharp income discontinuity, unlike the US Saver&amp;rsquo;s Credit), removing incentives for income manipulation. The maximum co-contribution payment was accordingly $1,000, $1,500, or $500 depending on the period. Estimation uses the ATO Longitudinal Information Files (ALife), a 10% random sample of all registered Australian tax filers linked longitudinally since 1990-91, covering 1,416,622 individual-year observations from 1999-2000 to 2016-17. The authors employ a first-differenced estimator exploiting within-individual variation in eligibility and match rates across years, conditioning on income, income squared, demographic controls, and year fixed effects.&lt;/p&gt;
&lt;p&gt;On the extensive margin, eligibility is associated with statistically significant but small increases in the probability of making any voluntary after-tax contribution: 0.6 percentage points at the 50% match rate, 0.9 percentage points at 100%, and 2.7 percentage points at 150%. Bunching at the salient $1,000 eligible maximum rises monotonically with the match rate: 0.23, 0.84, and 1.4 percentage points, respectively. Below $1,000, the probability of contributing in that range increases by 1.2, 1.6, and 2.7 percentage points — consistent with the substitution effect drawing in non-contributors and low contributors. Above $3,000, however, the probability of contributing falls significantly at all match rates: -0.66 pp (50%), -0.91 pp (100%), and -0.98 pp (150%), consistent with a retirement income windfall effect inducing high contributors to reduce their contributions toward the kink at $1,000.&lt;/p&gt;
&lt;p&gt;These opposing forces mean that average personal after-tax contributions (intensive margin) fall under all match-rate regimes: by $24.0 (50%), $24.6 (100%), and $6.49 (150%) per person-year, all significant. The attenuation of the fall at the 150% rate is consistent with substitution effects beginning to overshoot the eligible maximum and partially offsetting the income effect. When the government co-contribution payment itself is included, the combined personal-plus-government contribution rises ($40 at 100%, $126 at 150%), but these gains are partly offset by crowding out of voluntary concessional (salary sacrifice, pre-tax) contributions: eligibility is associated with 1.1 percentage point and 0.8 percentage point reductions in the proportion making voluntary concessional contributions at the 50% and 100% match rates respectively.&lt;/p&gt;
&lt;p&gt;Symmetry tests show no evidence of persistent habit formation: increases and decreases in treatment intensity produce contributions changes of roughly equal and opposite magnitudes on the extensive margin (gains +1.3 pp, losses -1.4 pp), ruling out the hypothesis that temporary eligibility establishes lasting savings behavior.&lt;/p&gt;
&lt;p&gt;Heterogeneity analysis reveals that the small average response reflects constrained liquidity. The response is largest for partnered females (+2.7 pp on the extensive margin), who have more discretionary income as secondary earners, and for those in the top permanent-income quintile (+3.6 pp), compared with bottom quintile (+0.4 pp) and second quintile (+0.7 pp). Responses increase with age and with lagged superannuation balance, with those holding balances above $100,000 responding at around 2.5 pp versus only 0.6 pp for those with balances below $25,000. There is no evidence that information is the binding constraint: respondents who use a tax consultant respond no more than those who self-file, and survey data document approximately 80% scheme awareness among superannuants.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central policy conclusion is that even a simple, transparent, and generous co-contribution scheme fails to meaningfully raise contributions of those it targets. The negative intensive margin arises because the scheme acts as a windfall for existing high contributors rather than newly inducing saving. These findings raise doubts about analogous reforms under discussion for the US Saver&amp;rsquo;s Credit.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-key-threats-to-it"&gt;Q1. What is the identification strategy and what are the key threats to it?&lt;/h3&gt;
&lt;p&gt;The primary estimator is a first-differenced OLS regression exploiting within-individual, year-on-year changes in co-contribution eligibility and match rates. Because the income thresholds shift over time and individuals&amp;rsquo; income fluctuates, the same person can move in and out of eligibility or across match-rate regimes, providing 16 distinct combinations of year-on-year changes in treatment status that identify the three match-rate coefficients. The key identification assumption is that first-differenced treatment indicators are contemporaneously uncorrelated with first-differenced idiosyncratic shocks. The main threat is income endogeneity — treatment is inversely related to income, and unobserved preferences to save may correlate with income. The authors address this by differencing out individual fixed effects and including income and income-squared as controls. They also test whether income manipulation around thresholds is occurring (it is not, unlike the US Saver&amp;rsquo;s Credit): frequency distributions of income show no bunching at the eligibility thresholds. The only income bunching observed is at the top of the lowest tax bracket (~$37,000), unrelated to scheme thresholds. As a robustness check, the authors also estimate individual fixed-effects models; results are broadly consistent, except for a theoretically inconsistent anomaly on the extensive margin for the 50% rate in the fixed-effects version, which the authors attribute to that model&amp;rsquo;s stricter exogeneity assumption being more likely violated in a life-cycle context.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-decompose-income-and-substitution-effects-and-what-is-the-empirical-test-for-each"&gt;Q2. How does the paper decompose income and substitution effects, and what is the empirical test for each?&lt;/h3&gt;
&lt;p&gt;The paper uses a two-period intertemporal model to show that the scheme creates a kinked budget constraint at the maximum eligible contribution (pmax). Those who would have contributed below pmax in the absence of the scheme face a lower cost of saving (substitution effect) and may increase contributions up to pmax. Those who would have contributed above pmax receive the co-contribution as a pure retirement income windfall, face no substitution incentive (the matching rate applies only below pmax), and respond only via a negative income effect by reducing contributions toward pmax. The empirical decomposition tests these predictions by estimating contribution probabilities in three ranges: contributions up to $1,000 (captures substitution effect), contributions between $1,001 and $3,000 (theoretically ambiguous — outflow from above $3,000 may offset inflow to $1,000), and contributions above $3,000 (captures negative income effect, as this range sits entirely above pmax). In Figure 5, the paper plots cumulative distribution function effects for each match rate across $100 increments from $0 to $10,000, showing negative effects on the CDF below $1,000 (substitution draws people above zero) and positive effects at and above $1,000 (income effect shifts mass below the maximum). The sign pattern is consistent with theory across all three match rates, and is more pronounced at higher match rates.&lt;/p&gt;
&lt;h3 id="q3-what-does-the-paper-find-about-bunching-at-the-1000-maximum-eligible-contribution"&gt;Q3. What does the paper find about bunching at the $1,000 maximum eligible contribution?&lt;/h3&gt;
&lt;p&gt;Eligibility is associated with significantly increased probability of contributing exactly $1,000, rising with the match rate: 0.23 pp at 50%, 0.84 pp at 100%, and 1.4 pp at 150%. The alternative specification distinguishing full eligibility (income below lower threshold, pmax = $1,000) from part eligibility (income in the tapered zone, pmax &amp;lt; $1,000) shows that part-eligible individuals also bunch significantly at $1,000 despite being entitled to match payments only for contributions below $1,000. This highlights the salience of the nominal maximum — people in the tapered zone treat $1,000 as the focal contribution amount rather than computing their individual optimal eligible contribution. The ATO online calculator does not report the maximum eligible contribution for part-eligible individuals, which likely reinforces this behavioral pattern.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-crowding-out-effects-on-unmatched-concessional-contributions"&gt;Q4. What are the crowding-out effects on unmatched (concessional) contributions?&lt;/h3&gt;
&lt;p&gt;The co-contribution scheme is associated with reductions in the use of voluntary concessional contributions (salary sacrifice, which are pre-tax and thus ineligible for matching). Using data from 2009-10 to 2016-17 (when salary sacrifice can be separated from compulsory employer contributions), the authors find that eligibility reduces the proportion of people making voluntary concessional contributions by 1.1 pp at the 50% match rate and 0.8 pp at the 100% match rate (both statistically significant). The data do not allow estimation at the 150% match rate because salary sacrifice records are unavailable before 2010. This crowding out compounds the scheme&amp;rsquo;s limited impact on total retirement savings: the net addition to retirement income from voluntary contributions is even smaller than the after-tax contribution estimates suggest. The mechanism attributed is the income windfall effect — for those who already made after-tax contributions in the absence of the scheme, the matching payment reduces their need for additional voluntary pre-tax saving.&lt;/p&gt;
&lt;h3 id="q5-is-there-evidence-of-asymmetry-in-scheme-effects--do-people-who-gain-eligibility-respond-differently-from-those-who-lose-it"&gt;Q5. Is there evidence of asymmetry in scheme effects — do people who gain eligibility respond differently from those who lose it?&lt;/h3&gt;
&lt;p&gt;The symmetry test in Equation (6) separates increases in treatment intensity (becoming eligible or moving to a higher match rate) from decreases (losing eligibility or moving to a lower rate). On the extensive margin, the effects are approximately symmetric: gaining intensity raises the contribution rate by 1.3 pp on average, while losing intensity reduces it by 1.4 pp. This rules out the &amp;rsquo;early targeting&amp;rsquo; hypothesis that short-term scheme exposure establishes lasting contribution habits that persist after eligibility ends. There is, however, some distributional asymmetry: bunching at $1,000 and the negative income effect above $3,000 are weaker in response to decreases in treatment intensity than to increases, suggesting some stickiness — people whose treatment falls may sustain slightly higher contributions for a period because prior co-contributions made them feel wealthier. But on the intensive margin, the reduction in average contributions is significant when treatment increases and statistically indistinguishable from zero when treatment decreases. The overall conclusion is no meaningful asymmetry that would justify life-cycle &amp;lsquo;seeding&amp;rsquo; arguments for young-age eligibility phased out later.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-in-responses-is-documented-and-what-does-it-imply-about-who-benefits"&gt;Q6. What heterogeneity in responses is documented, and what does it imply about who benefits?&lt;/h3&gt;
&lt;p&gt;Responses are largest among groups with greater discretionary income relative to their current consumption needs. Partnered females respond at 2.7 pp on the extensive margin (versus 1.2 pp for partnered males, 1.1 pp for single females, and 0.6 pp for single males). The interpretation is that partnered females are more likely to be secondary earners whose income is discretionary, reducing the liquidity cost of foregoing current consumption. The extensive margin response increases monotonically with permanent income quintile: 0.4 pp (bottom), 0.7 pp (2nd), 1.3 pp (3rd), 1.8 pp (4th), and 3.6 pp (top). Those in the top quintile are eligible only when their transitory income is temporarily low, and they appear to have both the liquid assets and the foresight to exploit the scheme. Responses increase with age, consistent with older workers facing lower liquidity constraints and having stronger retirement income motives. Lagged superannuation balance matters: those with balances above $100,000 respond at ~2.5 pp versus ~0.6 pp for those with balances below $25,000 — the scheme does not help low-balance individuals catch up. Importantly, there is no evidence that scheme uptake is constrained by information: tax-agent filers and self-filers respond at similar rates (~1.3 pp vs ~1.9 pp), and external surveys show roughly 80% public awareness. This rules out information provision as a policy lever likely to substantially raise the scheme&amp;rsquo;s impact.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-study-relate-to-and-differ-from-prior-evaluations-of-the-us-savers-credit-and-german-riester-schemes"&gt;Q7. How does this study relate to and differ from prior evaluations of the US Saver&amp;rsquo;s Credit and German Riester schemes?&lt;/h3&gt;
&lt;p&gt;Prior work on the Saver&amp;rsquo;s Credit (Duflo et al. 2007, Ramnath 2013, Heim and Lurie 2014) found small or null effects, attributed mainly to the scheme&amp;rsquo;s complexity — non-refundable tax credit with match rates of 11%, 25%, or 100% depending on income thresholds that create sharp discontinuities and strong income manipulation incentives. The Riester scheme (Corneo et al. 2009, 2010) showed zero effects on total savings, attributed to its complex co-contribution formula where the effective match rate depends on income and number of children, making the true incentive opaque. This paper&amp;rsquo;s contribution is to evaluate a scheme explicitly designed to avoid those complexities: a single flat match rate, co-contribution paid directly to the pension account, eligibility smoothly phased out with no discontinuities, and near-universal institutional coverage through mandatory superannuation. This design is analogous to the Duflo et al. (2006) H&amp;amp;R Block field experiment (which found 5–11 pp increases in contribution rates for 20–50% match rates), and the paper can be read as asking whether those larger field-experiment effects generalize to a national, ongoing program at comparable design simplicity. The answer is no: the national scheme produces responses an order of magnitude smaller than the field experiment. The paper attributes this partly to the field experiment&amp;rsquo;s &amp;lsquo;one-time-only&amp;rsquo; nature (creating urgency), potential interaction with Saver&amp;rsquo;s Credit tax refunds, and selection of H&amp;amp;R Block clients. The Australian study also goes beyond prior work by estimating distributional effects (contribution ranges), crowding out of unmatched contributions, and symmetry tests — none of which were examined in the prior national scheme evaluations.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-papers-policy-implications-and-their-scope-conditions"&gt;Q8. What are the paper&amp;rsquo;s policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary implication is that co-contribution matching schemes, even when simple, generous, and widely known, are likely to produce small effects on retirement savings of low- and middle-income earners. The mechanism is that many in the eligible population already contributed more than the scheme maximum and treat the matching payment as a windfall, reducing personal contributions. The scheme is particularly ineffective for the lowest permanent-income earners, who face binding liquidity constraints and respond least even when they are aware of the scheme. This is directly relevant to proposed US reforms of the Saver&amp;rsquo;s Credit (the Retirement Security and Savings Act considered by Congress at time of writing) that would convert it to a direct co-contribution more like Australia&amp;rsquo;s scheme — the paper&amp;rsquo;s results suggest such simplification may not yield large savings increases. A scope condition concerns institutional context: Australia has near-universal mandatory superannuation with employer contributions at 9.5% of earnings, which may reduce the marginal value of voluntary contributions. The authors acknowledge that responses might be higher in countries without mandatory employer coverage, though the finding that lower-balance individuals respond least makes this qualification weak. A second scope condition is that the scheme excludes compulsory employer contributions from the matching base, so the results speak specifically to voluntary behavior. Future research is identified on whether tightening access to public pensions (raising the pension access age) would increase voluntary contributions among low-income earners who currently rely on public pensions as their retirement backstop.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-conducted"&gt;Q9. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The authors report four main robustness exercises. First, they estimate an individual fixed-effects model alongside the first-differenced model; results are broadly consistent, with the noted exception of a theoretically inconsistent anomaly at the 50% match rate for the extensive margin in the fixed-effects version, attributed to violation of the strict exogeneity assumption. This validates the first-differenced approach as the preferred specification. Second, they extend the base model to distinguish full eligibility (income at or below the lower threshold, pmax = $1,000) from part eligibility (income in the tapered zone, pmax &amp;lt; $1,000), confirming that even partial eligibility generates bunching at the salient $1,000 level. Third, they examine distributional predictions by estimating the model for 100 incremental contribution thresholds from $0 to $10,000 (Figure 5), verifying that the CDF-effect pattern is consistent with the theoretical predictions across all three match rates. Fourth, information access is tested by interacting scheme response with whether a tax agent was used to lodge the return; the absence of any significant difference between tax-agent filers and self-filers, combined with documented high public awareness, eliminates information deficiency as an explanation for the small response.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Co-contribution matching scheme&lt;/strong&gt;: A government program that pays a specified fraction (the matching rate) of the individual&amp;rsquo;s voluntary personal pension contributions up to a maximum eligible contribution ceiling, credited directly to the individual&amp;rsquo;s retirement account — as distinct from a tax credit that may not reach the account.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Retirement income effect (windfall effect)&lt;/strong&gt;: The tendency of matching payments to reduce voluntary personal contributions among those who would have contributed above the scheme maximum in the scheme&amp;rsquo;s absence: because the government contribution supplements their retirement income regardless of their own effort, they rationally reduce personal saving to the eligible maximum.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Substitution effect (in this scheme)&lt;/strong&gt;: The scheme&amp;rsquo;s reduction in the effective cost of contributing by raising the return to each dollar contributed, inducing those who previously contributed below the eligible maximum to increase contributions toward that maximum.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bunching at the eligible maximum&lt;/strong&gt;: Mass concentration of contributions at exactly $1,000 (the scheme&amp;rsquo;s nominal maximum eligible contribution), drawing both from below (via the substitution effect) and from above (via the income/windfall effect), and reinforced by the salience of the round-number maximum even for part-eligible individuals whose true eligible maximum is below $1,000.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Permanent income (in this context)&lt;/strong&gt;: The predicted value of long-run log total personal income estimated from a Mincer-style regression including individual fixed effects, used to distinguish individuals who are structurally low-income (and face genuine liquidity constraints) from those whose transitory income is temporarily low and who are high-permanent-income individuals exploiting the scheme.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crowding out of concessional contributions&lt;/strong&gt;: The reduction in voluntary pre-tax (salary sacrifice) superannuation contributions associated with scheme eligibility, reflecting the income windfall from the matching payment reducing the need for supplementary retirement saving through the pre-tax channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Symmetry of scheme effects&lt;/strong&gt;: The property that the contribution response to gaining eligibility (or a higher match rate) is equal in magnitude and opposite in sign to the response to losing eligibility (or a lower match rate); symmetry implies no lasting habit formation from scheme exposure and rules out &amp;rsquo;early targeting&amp;rsquo; strategies aimed at establishing lifetime saving patterns.&lt;/p&gt;</description></item><item><title>Balancing Work and Care: How Workplace Factors Can Mitigate the Gendered Impacts of Caregiving</title><link>https://macropaperwarehouse.com/papers/balancing-work-and-care-how-workplace-factors-can-mitigate-the-gendered-impacts-of-caregiving/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/balancing-work-and-care-how-workplace-factors-can-mitigate-the-gendered-impacts-of-caregiving/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper examines how workplace environments shape the economic consequences that fall on mothers — but not fathers — when a child is diagnosed with cancer. The motivation is a gap in the caregiving-and-labor-markets literature: while the earnings penalties from childbirth are well-documented, less is known about caregiving shocks that arrive later in childhood, or about whether and how the firm, occupation, or industry a parent works in moderates those penalties.&lt;/p&gt;
&lt;p&gt;The empirical setting is Australia. The authors use the ABS Person Level Integrated Data Asset (PLIDA), a longitudinal administrative database linking tax records (ATO, 2005–2022), Medicare health records, and 2011 Census occupation and hours data. A distinctive feature is matched employer-employee identifiers, enabling construction of workplace characteristics at the firm, occupation, and industry levels. The sample comprises 3,258 families in which a child (age 4–18, average age 12.98) began chemotherapy between 2012 and 2023 and both parents were employed two years before treatment. Pre-diagnosis average earnings are $37,639 for mothers and $79,702 for fathers (CPI-adjusted to 2012).&lt;/p&gt;
&lt;p&gt;The identification strategy is a dynamic difference-in-differences (DiD) model following Fadlon and Nielsen (2019, 2021). The treatment group consists of parents whose children started chemotherapy between 2012 and 2017; the control group consists of parents whose children will receive the same diagnosis later, between 2018 and 2023, with placebo treatment assigned six years before actual treatment. Individual fixed effects absorb time-invariant heterogeneity; year fixed effects absorb common trends. Childhood cancer — specifically chemotherapy-requiring cancer — is treated as a largely random shock with no pre-trend in earnings or employment between treated and control families before diagnosis.&lt;/p&gt;
&lt;p&gt;Main findings on the average effects: Maternal earnings fall by $5,608 in the year chemotherapy begins (14.9% of baseline earnings). The earnings decline persists for at least three years even as measured caregiving intensity (child healthcare service use) returns to baseline by year 3, leaving earnings approximately 9.7% below baseline in year 3 (−$3,645). The primary mechanism is a reduction in hours worked rather than outright job exit: employment falls by 4.9 percentage points in year 0, peaking at a decline of 5.6 percentage points two years post-treatment, a modest reduction relative to the earnings loss. Job-to-job transitions are not significantly elevated. Mental health service use (therapy, antidepressants, anxiolytics) shows no significant change for either parent, ruling out a mental health channel and reinforcing that caregiver time demands drive the result. Fathers experience no statistically significant change in earnings, employment, or job transitions across all specifications.&lt;/p&gt;
&lt;p&gt;Subgroup heterogeneity: The earnings penalty is substantially larger for mothers of younger children (under 12): −$9,443 in year 0, equivalent to 25.8% of that subgroup&amp;rsquo;s baseline earnings. For children with above-median healthcare utilization, the year-0 penalty is −$7,826 (21.6%).&lt;/p&gt;
&lt;p&gt;Workplace moderation — three dimensions are examined at the firm, occupation, and industry levels:&lt;/p&gt;
&lt;p&gt;(1) Gender pay gap: Mothers in occupations with below-average gender pay gaps face lower earnings losses ($5,782 vs $8,409; 16.5% vs 18.1%). The effect is significant at the occupation level but not at the firm or industry level.&lt;/p&gt;
&lt;p&gt;(2) Work hour intensity: Mothers in firms with below-median weekly hours face a year-0 earnings loss of $3,240 (9.9%) versus $7,159 (15.6%) in high-hours firms — a difference of $3,919, significant at the firm level. A parallel gap holds at the occupation level. When both firm and occupation are low-hours, the combined loss equals $2,519; when both are high-hours, it reaches $9,357 — a fourfold difference.&lt;/p&gt;
&lt;p&gt;(3) Female representation in the top 20% of earners: Mothers at firms where women are the majority of top-20%-earners suffer a penalty of $3,856 (8.3%) versus $7,799 (23.4%) elsewhere — a $3,943 mitigation at the firm level. At the occupation level the corresponding figures are $4,240 (9.2%) versus $8,356 (25.0%). Female representation in middle or bottom earnings tiers carries no significant moderating effect.&lt;/p&gt;
&lt;p&gt;In the combined specification (all firm- and occupation-level variables simultaneously), female representation in the top 20% and work hour intensity remain jointly significant; the gender pay gap loses significance, consistent with these variables being correlated. In the polar comparison between fully supportive jobs (low hours, high female senior representation, low occupation gender pay gap) and fully unsupportive jobs (opposite), the difference is dramatic: mothers in supportive jobs suffer a −$6,280 year-0 earnings hit that recovers fully by year 1, while mothers in unsupportive jobs face −$10,416 in year 0 widening to −$13,882 in year 3 before partially recovering in year 4.&lt;/p&gt;
&lt;p&gt;Policy implications (with scope conditions): The results support policies that reduce greedy-work norms and increase female representation in senior roles as instruments for attenuating the gendered economic cost of caregiving shocks. The study does not isolate specific workplace policies (e.g., formal paid leave) but identifies observable correlates of supportive environments. Effects are identified among working parents of children requiring chemotherapy; they do not generalize to cancer not requiring chemotherapy or other types of caregiving shocks without further evidence. Notably, fathers&amp;rsquo; outcomes are unresponsive to workplace factors, suggesting that social norms or intra-household bargaining — not workplace barriers per se — are the primary constraints on paternal caregiving adjustment.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The authors use a later-treated dynamic DiD, comparing parents whose children began chemotherapy 2012–2017 (treated) to parents whose children will begin the same treatment 2018–2023 (control), with the control group&amp;rsquo;s placebo treatment assigned six years before their actual treatment. Individual fixed effects absorb time-invariant heterogeneity; year fixed effects absorb macro shocks. The parallel trends assumption is validated by showing: (1) no statistically significant differences in pre-cancer demographic, socioeconomic, or workplace characteristics between treated and control groups (Figure 1); and (2) no pre-trend in earnings or employment in years -4 and -3 relative to baseline (Table A3, estimates small and insignificant). The main threats acknowledged are (a) non-random selection into workplace types — mothers who anticipate greater caregiving loads may sort into more family-friendly jobs — and (b) differences in baseline wage levels across job types. On (a), the authors argue the direction of selection bias goes the wrong way: if selection were driving results, mothers in supportive workplaces (who selected there due to caregiving preferences) would have weaker labor market attachment and larger post-shock earnings declines; instead the opposite is found. On (b), the authors show that absolute dollar declines in less-supportive workplaces also correspond to larger percentage declines relative to baseline, so the pattern is not an artifact of higher baseline wages in high-hour jobs (though Appendix Table A2 confirms mothers in high-hour and high-senior-female firms do have higher baseline earnings of around $46,000–$50,000 vs $32,000–$33,000).&lt;/p&gt;
&lt;h3 id="q2-how-is-the-caregiving-shock-defined-and-what-does-this-imply-for-external-validity"&gt;Q2. How is the caregiving shock defined and what does this imply for external validity?&lt;/h3&gt;
&lt;p&gt;The shock is defined as initiation of chemotherapy by the child, identified from Medicare prescription records using ATC codes beginning with L01 (excluding methotrexate L01BA01) and adding immunomodulators with chemotherapy-like effects. Chemotherapy initiation is treated as a reliable, time-consistent marker because it typically follows immediately from diagnosis of cancers such as acute lymphoid leukemia, astrocytoma, and neuroblastoma. The authors note explicitly that estimates do not represent the effects of childhood cancer not requiring chemotherapy (e.g., early-stage cancers treated with surgery, radiation, or immunotherapy alone). This restriction to chemotherapy-requiring cancers likely selects a sample with above-average caregiving intensity.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-main-mechanism-through-which-the-earnings-decline-operates"&gt;Q3. What is the main mechanism through which the earnings decline operates?&lt;/h3&gt;
&lt;p&gt;The primary mechanism is a reduction in hours worked rather than outright job exit. The employment decline (approximately 4.5–5.0 percentage points in years 0–2 per Table A3) is modest relative to the earnings loss of $5,608. A back-of-envelope calculation in footnote 6 shows that if 5% of mothers left the labor market at average earnings, the implied earnings drop would be only $1,882, far below the observed $5,608. Job-to-job transitions (probability of switching employer) are not significantly elevated. Mental health service use (psychological therapy, antidepressant/anxiolytic/antipsychotic prescriptions) shows no significant change for either parent (Appendix Figure A4), ruling out mental health deterioration as a channel. The persistence of earnings losses beyond the period of peak healthcare service use (which returns to baseline by year 3, per Appendix Figure A2) is consistent with stalled career trajectories — foregone promotions or skill development — or with continued but less-measured caregiving demands.&lt;/p&gt;
&lt;h3 id="q4-at-which-organizational-level-firm-occupation-or-industry-do-workplace-moderators-operate-most-strongly"&gt;Q4. At which organizational level (firm, occupation, or industry) do workplace moderators operate most strongly?&lt;/h3&gt;
&lt;p&gt;Firm and occupation levels are the dominant levels; industry-level measures are consistently insignificant for all three moderating variables. The authors interpret this as follows: industry-level measures are too broad to capture the specific work arrangements and norms that affect caregiving balance. At the occupation level, structural characteristics — profession-wide agreements, flexibility of task-based roles, part-time feasibility — directly govern how feasible it is to reduce hours without exiting employment. At the firm level, immediate workplace culture and specific HR policies apply. The relative contribution of firm vs occupation varies by the moderator: work hour intensity effects are significant at both firm and occupation levels, female senior representation is significant at both, while the gender pay gap effect is significant only at the occupation level.&lt;/p&gt;
&lt;h3 id="q5-why-does-female-representation-in-senior-roles-top-20-of-earners-mitigate-the-earnings-penalty-while-middle-and-bottom-tier-representation-does-not"&gt;Q5. Why does female representation in senior roles (top 20% of earners) mitigate the earnings penalty while middle and bottom tier representation does not?&lt;/h3&gt;
&lt;p&gt;The authors argue that women in the top-20% of earners — effectively leadership positions — are better positioned to advocate for and implement caregiving-supportive policies (paid leave, flexible scheduling). Representation in lower tiers may be indicative of a caregiving-friendly workforce composition but lacks the organizational power to shape policies. This is supported empirically: the moderating interaction is significant and economically large for top-20% female representation at both the firm (mitigating the penalty by $3,943) and occupation levels (mitigating by $4,116), while interactions for the middle 50–80% and bottom 50% earnings tiers are not statistically significant in most specifications.&lt;/p&gt;
&lt;h3 id="q6-why-does-the-occupational-gender-pay-gap-matter-for-the-earnings-penalty-but-not-the-firm-level-or-industry-level-gap"&gt;Q6. Why does the occupational gender pay gap matter for the earnings penalty but not the firm-level or industry-level gap?&lt;/h3&gt;
&lt;p&gt;The authors offer two explanations. First, occupations define the day-to-day nature of work — task structure, required hours, flexibility — in ways that make caregiving more or less compatible. Occupations that accommodate part-time and flexible scheduling tend to attract more women and develop norms that support caregiving, which in turn narrows occupational gender pay gaps. At the firm level, the same firm often contains diverse occupations with heterogeneous norms, so firm-level gender pay gap is a noisier signal. At the industry level, the measure is too aggregated. Second, narrow occupational gender pay gaps may reflect the collective bargaining power of women in female-dominated occupations (e.g., nursing), which translates into formal caregiving protections. A firm or industry may exhibit a wide gender pay gap due to male dominance in senior or high-earning roles even when specific female-dominated occupations within that firm/industry have caregiving-friendly norms. However, in the combined specification including all workplace factors simultaneously, the gender pay gap variable loses statistical significance, suggesting its initial effect was partly mediated by correlated factors (hours intensity and female senior representation).&lt;/p&gt;
&lt;h3 id="q7-how-does-the-combined-supportive-vs-unsupportive-comparison-work-and-what-does-it-show"&gt;Q7. How does the combined &amp;lsquo;supportive vs unsupportive&amp;rsquo; comparison work and what does it show?&lt;/h3&gt;
&lt;p&gt;Supportive jobs are defined as those satisfying all three criteria: low work hour intensity at both firm and occupation levels, high female representation in the top 20% of earners at both firm and occupation levels, and low gender pay gap at the occupation level (N = 2,708 mother-years). Unsupportive jobs are the opposite on all criteria (N = 2,339). Event study estimates (Table A9, Figure 3) show stark divergence. In supportive jobs, the year-0 penalty is −$6,280, and earnings recover quickly to statistically insignificant levels by years 1–4. In unsupportive jobs, the year-0 penalty is −$10,416, it widens to −$10,658 in year 2 and −$13,882 in year 3, before partially recovering in year 4. Pre-treatment estimates are not significantly different from zero in both subsamples, supporting parallel trends within each group.&lt;/p&gt;
&lt;h3 id="q8-what-heterogeneity-is-documented-by-child-and-family-characteristics"&gt;Q8. What heterogeneity is documented by child and family characteristics?&lt;/h3&gt;
&lt;p&gt;Appendix Figure A3 presents two subgroup analyses. Mothers of children under age 12 at diagnosis experience a year-0 earnings loss of −$9,443 (25.8% of baseline earnings of $36,567), substantially larger than the average. Mothers of children with above-median healthcare utilization (measured by number of medical appointments in the year following treatment initiation) experience a year-0 loss of −$7,826 (21.6% of baseline earnings of $36,278). These patterns are consistent with the interpretation that caregiving intensity — driven by child age and treatment severity — scales the maternal earnings penalty.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-conducted"&gt;Q9. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s main robustness arguments are: (1) pre-trend validation (Figures 1 and 2, Table A3) confirming no anticipatory effects and balanced pre-characteristics; (2) the selection-direction argument for workplace heterogeneity — the selection story would predict larger penalties in supportive workplaces but the opposite is found; (3) showing that absolute earnings declines in less-supportive workplaces also represent larger proportional declines relative to baseline, ruling out a level-effect interpretation; (4) the mental health non-result (Appendix Figure A4) confirming earnings effects are not confounded by parental mental health deterioration; (5) separate combined specification (Table A8) testing all workplace moderators simultaneously to address multicollinearity. The paper does not report explicit placebo tests using alternative shocks or falsification samples, nor does it report results restricted to narrow geographic areas or specific cancer types.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-prior-literature-on-caregiving-shocks"&gt;Q10. How does this paper relate to prior literature on caregiving shocks?&lt;/h3&gt;
&lt;p&gt;The paper builds most directly on three prior studies using Nordic or European administrative data: Eriksen et al. (2021, Journal of Health Economics) on childhood health shocks and parental labor supply; Breivik and Costa-Ramon (2024, Review of Economics and Statistics) on children&amp;rsquo;s health shocks and parental earnings and mental health; and Vaalavuo et al. (2023, Demography) on gender inequality from child health shocks on parental trajectories. All three find significant maternal earnings or employment losses and no or small paternal effects. The present paper&amp;rsquo;s contribution relative to these is the explicit examination of how firm-, occupation-, and industry-level workplace characteristics moderate the maternal penalty — a dimension the prior literature has not addressed. It also connects to Fadlon and Nielsen (2019, 2021) on the methodology and to the broader child-penalty literature reviewed by Cortes and Pan (2023, Journal of Economic Literature). On workplace mechanisms it connects to Goldin (2014) on &amp;lsquo;greedy jobs&amp;rsquo; and Goldin and Katz (2016) on pharmacy as a family-friendly profession.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The findings suggest that maternal earnings losses from caregiving shocks can be substantially mitigated by workplace environments characterized by lower work hour intensity and higher female representation in senior earnings tiers. This points to policies promoting: (1) reduced greedy-work norms — discouraging long-hours cultures and enabling part-time flexibility without disproportionate wage penalties; (2) greater female representation in leadership and high-earning positions, which appears to create cultural and policy environments more accommodating of caregiving. Scope conditions: the results apply to working mothers (and fathers) of children requiring chemotherapy in Australia, where Medicare provides universal healthcare coverage and existing social insurance exists. The paper explicitly does not identify specific causal mechanisms (e.g., it cannot isolate the effect of formal paid leave from culture). On fathers, the implication is that workplace factors alone are unlikely to induce fathers to increase caregiving, pointing instead to the need to shift social norms around paternal caregiving and intra-household bargaining.&lt;/p&gt;
&lt;h3 id="q12-how-do-the-australian-institutional-context-and-data-compare-to-european-studies"&gt;Q12. How do the Australian institutional context and data compare to European studies?&lt;/h3&gt;
&lt;p&gt;Australia&amp;rsquo;s PLIDA dataset is exceptional in combining population-level coverage, employer-employee identifiers (enabling firm-level workplace measures), and Medicare healthcare records (enabling both shock identification via chemotherapy and caregiving-intensity proxying via healthcare utilization). The employer identifiers are critical for this paper&amp;rsquo;s contribution — most comparable European studies cannot construct firm-level workplace characteristics. The Australian context differs from Nordic studies in terms of family policy generosity (less universal paid parental leave), but Medicare provides universal healthcare access. Pre-diagnosis earnings ($37,639 for mothers vs $79,702 for fathers) indicate a large pre-existing earnings gap, consistent with a majority-male breadwinner household structure in the sample.&lt;/p&gt;
&lt;h3 id="q13-do-fathers-outcomes-respond-to-any-workplace-factor"&gt;Q13. Do fathers&amp;rsquo; outcomes respond to any workplace factor?&lt;/h3&gt;
&lt;p&gt;In almost all specifications, fathers&amp;rsquo; earnings, employment, and job changes show no statistically significant effects of the caregiving shock and no significant interactions with workplace characteristics (Appendix Tables A4 and A6). One exception: in Table A4, the interaction between the cancer shock and working at a firm with above-median work hours is negative and significant at the 5% level for fathers, suggesting that fathers who work in high-hours firms do experience some earnings reduction — consistent with them reducing hours in an environment that penalizes deviations from long hours. However, the authors note the effect is substantially smaller relative to baseline earnings than the corresponding maternal effect. The broader pattern implies that workplace flexibility does not appear to be the binding constraint preventing fathers from taking on more caregiving; social norms and intra-household bargaining are posited as more important.&lt;/p&gt;
&lt;h3 id="q14-what-are-the-data-limitations-and-caveats"&gt;Q14. What are the data limitations and caveats?&lt;/h3&gt;
&lt;p&gt;First, work hours at the firm and occupation levels are constructed from the 2011 Census, which is a single cross-section; work hour norms may have shifted between 2011 and the 2012–2023 sample period. Occupation and industry codes also come from the 2011 Census, so parents who changed occupation between 2011 and their baseline year may be misclassified. Second, employment status is inferred from positive ATO earnings in a financial year, a coarser measure than actual employment spells. Third, the sample is restricted to firms with at least 10 employees, which excludes small-firm workers. Fourth, the analysis uses dollar earnings levels, not log earnings, which means baseline wage differences across workplace types can affect the interpretation of absolute dollar results (though the authors show percentage effects are also larger in less-supportive workplaces). Fifth, the study identifies workplace correlates of smaller penalties but does not isolate the causal effect of any specific policy. Sixth, the paper covers only cancer requiring chemotherapy — typically more intensive cancers — so results may overstate average caregiving-shock effects.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Caregiving shock&lt;/strong&gt;: In this paper, a sudden, largely unanticipated increase in caregiving demands on parents triggered by a child&amp;rsquo;s initiation of chemotherapy. Distinguished from the chronic caregiving burden of childbirth; specifically refers to health events that arrive later in childhood and impose large, time-intensive care requirements.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Later-treated dynamic DiD&lt;/strong&gt;: The paper&amp;rsquo;s identification design, following Fadlon and Nielsen (2019, 2021), in which the control group consists of parents who will receive the same treatment (child&amp;rsquo;s cancer diagnosis) at a later date. The control group&amp;rsquo;s placebo treatment year is set six years before their actual treatment, enabling estimation of time-path effects relative to diagnosis while accounting for pre-existing differences via individual fixed effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Work hour intensity&lt;/strong&gt;: Median weekly hours worked by employees at a given firm or in a given occupation (from the 2011 Census), used as a proxy for &amp;lsquo;greedy job&amp;rsquo; characteristics — workplaces that reward continuous long-hours presence and penalize deviations. High work hour intensity captures both above-full-time norms and the likely presence of evening and weekend work requirements.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Female representation in the top 20% of earners&lt;/strong&gt;: A binary indicator equal to one when women are the majority (above 50%) of workers in the top quintile of earnings at a given firm or occupation. The paper distinguishes this from female representation in middle and lower earnings tiers to isolate the effect of women&amp;rsquo;s presence in positions with organizational power to influence workplace policies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Supportive job&lt;/strong&gt;: As defined operationally in this paper: a job in which the worker&amp;rsquo;s firm and occupation both have below-median work hour intensity, both have majority female representation in the top 20% of earners, and the occupation has a below-average gender pay gap. Mothers in supportive jobs suffer smaller and shorter-lived earnings penalties following a caregiving shock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Greedy occupation&lt;/strong&gt;: Borrowed from Goldin (2014), and used in this paper to describe occupations that disproportionately reward workers who supply long, often inflexible, hours. In the paper&amp;rsquo;s empirical framework, these are occupations with above-median work hour intensity, which are shown to amplify maternal earnings losses after a caregiving shock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Caregiving intensity&lt;/strong&gt;: The time-varying burden of care associated with a child&amp;rsquo;s illness, proxied in this paper by the volume of child healthcare service utilization (Medicare items: GP visits, specialist consultations, diagnostic imaging, prescriptions). Caregiving intensity peaks at year 0 (treatment initiation), declines significantly by year 2, and returns to baseline by year 3 — yet maternal earnings penalties persist beyond this return to baseline.&lt;/p&gt;
&lt;!-- flags: Employment figures cited in the text (4.9 pp in year 0; peak of 5.6 pp in year 2) differ slightly from Table A3 values (-0.045 = 4.5 pp in year 0; -0.050 = 5.0 pp in year 2). This is a within-paper discrepancy in the IZA working paper version. Layer 1 reports the text-stated figures as authored. --&gt;</description></item><item><title>Business Cycle during Structural Change: Arthur Lewis' Theory from a Neoclassical Perspective</title><link>https://macropaperwarehouse.com/papers/business-cycle-during-structural-change-arthur-lewis-theory-from-a-neoclassical-perspective/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/business-cycle-during-structural-change-arthur-lewis-theory-from-a-neoclassical-perspective/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks why the nature of business cycles changes systematically as economies develop and shed their large agricultural sectors. The motivation is both empirical and theoretical. Empirically, countries with large declining agricultural sectors—most prominently China—exhibit business cycle patterns that depart sharply from the textbook procyclical-employment pattern seen in mature economies: aggregate employment is acyclical with respect to GDP, nonagricultural employment is strongly procyclical, agricultural employment is countercyclical, and the labor productivity gap between nonagriculture and agriculture narrows during booms. These cross-country regularities hold in a sample of 63–66 countries using ILO sectoral employment data over 1970–2015, with the correlation between aggregate employment and GDP declining monotonically as the agricultural employment share rises. The cross-country correlation between the agricultural employment share and log GDP per capita is −0.84. For China specifically over 1978–2012, the correlation between HP-filtered agricultural employment and GDP is −0.69, while the correlation for nonagricultural employment with GDP is 0.73. Agricultural employment fell from about 62.4% of total Chinese employment in 1985 to 33.6% in 2012.&lt;/p&gt;
&lt;p&gt;The authors construct a unified neoclassical model of growth, structural change, and business cycles. The economy produces a CES aggregate of agricultural and nonagricultural output (elasticity of substitution epsilon), with agriculture itself being a CES aggregate of modern and traditional sub-sectors (elasticity omega). Modern agriculture uses capital and labor (Cobb-Douglas), whereas traditional agriculture uses only labor. This nested structure means the effective elasticity of substitution between capital and labor in agriculture is variable and declines as the traditional sector shrinks—formalizing the Lewisian surplus-labor mechanism within a neoclassical framework. A time-invariant tax wedge tau on nonagricultural wages captures rural-urban earnings gaps and keeps agriculture inefficiently large.&lt;/p&gt;
&lt;p&gt;The deterministic model is estimated using Simulated Method of Moments on Chinese data from 1985 to 2012, targeting seven moment sequences: employment share in agriculture, capital share in agriculture, agricultural output-to-GDP ratio, agricultural expenditure share, aggregate GDP growth, the aggregate capital-output ratio path, and the change in the productivity gap. Key findings from estimation: the elasticity of substitution between agricultural and nonagricultural goods epsilon is estimated at 3.6 (significantly greater than 1 at 1% level), and the elasticity between modern and traditional agriculture omega is also very large. The estimated subsistence level in a Stone-Geary extension is small (11% of agricultural production in 1985), so nonhomothetic preferences play only a minor quantitative role. Nonagricultural TFP growth gM is estimated at 6.5% per year; modern-agricultural TFP growth gAM at 6.1% per year; traditional-sector TFP growth gS at 0.9% per year. The estimated labor wedge tau implies persistent misallocation.&lt;/p&gt;
&lt;p&gt;Stochastic TFP shocks (VAR(1) for each of the three sectors) are then estimated from observed data by exploiting the model&amp;rsquo;s equilibrium conditions. The persistence parameters are 0.63 (nonagriculture), 0.90 (modern agriculture), and 0.42 (traditional agriculture). The model, simulated 1,000 times starting in 1980, reproduces the salient Chinese business cycle features: the standard deviation of GDP is 1.7% (matching the data), agricultural employment is countercyclical (model correlation with GDP: −0.25; data: −0.23), nonagricultural employment is strongly procyclical (model: 0.99; data: 0.73), and aggregate employment has a low correlation with GDP (model: 0.42; data: 0.10). A variance decomposition shows nonagricultural TFP shocks account for approximately 95% of GDP fluctuations.&lt;/p&gt;
&lt;p&gt;The key mechanism is that a large traditional sector provides an elastic labor supply to nonagriculture at low marginal cost (a neoclassical Lewisian buffer). Positive TFP shocks to nonagriculture draw labor out of traditional agriculture, raising average capital intensity and labor productivity in agriculture—hence the countercyclical productivity gap. As structural change progresses and the traditional sector shrinks, this labor buffer disappears, the effective labor supply elasticity declines, and business cycle properties converge toward those of a standard neoclassical (Hansen-Prescott) economy. Out-of-sample simulations confirm this convergence: the correlation between total employment and GDP rises from around 40% to near 100% as the agricultural employment share falls below 10%. The paper also shows that positive TFP shocks in agriculture slow structural change, consistent with empirical evidence from the Green Revolution (Foster and Rosenzweig 2004; Bustos et al. 2016; Moscona 2018; Jayachandran 2006).&lt;/p&gt;
&lt;p&gt;Elasticity estimates using CES production functions for the US, Japan, and China from consumption value-added data yield epsilon of 2.49, 1.58, and 1.70 respectively, all significantly above unity at the 1% level—supporting the labor-pull interpretation of structural change. The authors find that imposing the symmetry restriction (epsilon = epsilon_ms) used by Herrendorf et al. (2013) replicates their near-zero estimate for the US, but relaxing that restriction reveals the agriculture-nonagriculture elasticity to be large while the manufacturing-services elasticity is near zero.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-four-key-business-cycle-stylized-facts-documented-for-countries-with-large-agricultural-sectors"&gt;Q1. What are the four key business cycle stylized facts documented for countries with large agricultural sectors?&lt;/h3&gt;
&lt;p&gt;The paper documents four regularities that hold across 63–66 countries (ILO data, 1970–2015): (1) aggregate employment is less correlated with GDP and less volatile; (2) agricultural employment is countercyclical; (3) the labor productivity gap (nonagriculture/agriculture) is negatively correlated with nonagricultural employment; (4) consumption is highly volatile relative to GDP. All four are quantitatively documented for China and compared with the US.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-core-theoretical-mechanism-distinguishing-this-paper-from-earlier-structural-change-models"&gt;Q2. What is the core theoretical mechanism distinguishing this paper from earlier structural-change models?&lt;/h3&gt;
&lt;p&gt;The paper adds an internal split of the agricultural sector into modern (capital-using Cobb-Douglas) and traditional (labor-only) sub-sectors that are imperfect substitutes. This nested structure generates a variable effective elasticity of labor supply to nonagriculture: when the traditional sector is large, labor can be released to industry at near-constant marginal cost (a continuous Lewisian surplus), dampening wage and price fluctuations and decoupling aggregate employment from GDP. As the traditional sector shrinks through capital accumulation and differential TFP growth, the effective labor-supply elasticity falls, progressively transforming the economy into a standard neoclassical one.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-handle-the-lack-of-a-steady-state-for-the-business-cycle-analysis"&gt;Q3. How does the paper handle the lack of a steady state for the business cycle analysis?&lt;/h3&gt;
&lt;p&gt;Because structural change is ongoing in China, approximating the model around a balanced growth path is infeasible. The authors instead solve the model recursively over 250 periods back from an assumed one-sector asymptotic balanced growth path (ABGP), using a 27-state Tauchen Markov chain for the three TFP shocks and piecewise linear decision rules on a 75-point grid for each of the two continuous state variables (kappa and kappa-tilde). They simulate 1,000 economies and compute rolling 28-year window statistics, which are then compared to the data.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-identification-strategy-for-the-elasticity-of-substitution-epsilon-and-what-are-the-main-threats"&gt;Q4. What is the identification strategy for the elasticity of substitution epsilon, and what are the main threats?&lt;/h3&gt;
&lt;p&gt;The primary strategy is Simulated Method of Moments on 143 moment conditions from Chinese data 1985–2012 (28 annual observations each for five moment series plus two level/change moments). A second strategy uses IFGNLS estimation of a Stone-Geary demand system for three countries (US, Japan, China) using both consumption value-added (Herrendorf et al. method) and production value-added (GGDC data). The main threats acknowledged: (a) endogeneity—both sides of the demand equations are driven by unobserved productivity and preference shocks with opposite sign implications (addressed by turning to exogenous Green Revolution shocks); (b) measurement error; (c) the symmetry restriction in prior work; (d) the model is closed-economy and abstracts from demand shocks.&lt;/p&gt;
&lt;h3 id="q5-what-role-do-agricultural-tfp-shocks-versus-nonagricultural-tfp-shocks-play-in-gdp-fluctuations"&gt;Q5. What role do agricultural TFP shocks versus nonagricultural TFP shocks play in GDP fluctuations?&lt;/h3&gt;
&lt;p&gt;A variance decomposition shows nonagricultural TFP shocks (ZM) account for approximately 95% of GDP fluctuations in the benchmark economy over 1985–2012. The logic is that positive TFP shocks to ZM reduce misallocation by drawing labor from the (inefficiently large) agricultural sector to nonagriculture, amplifying the GDP response. In contrast, positive TFP shocks to agriculture partially offset the direct productivity gain by worsening misallocation (labor stays in agriculture), so GDP barely responds. In the low-elasticity (epsilon = 0.5) alternative model, agricultural TFP shocks account for about half of GDP fluctuations—one reason the authors reject this alternative.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-models-prediction-for-business-cycle-evolution-as-structural-change-progresses-compare-to-cross-country-evidence"&gt;Q6. How does the model&amp;rsquo;s prediction for business cycle evolution as structural change progresses compare to cross-country evidence?&lt;/h3&gt;
&lt;p&gt;Using rolling 28-year windows of simulated data from 1985 to 2185, the paper documents four monotone transitions as the agricultural employment share falls: (a) the correlation between agricultural employment and the productivity gap falls toward zero; (b) the correlation between agricultural and nonagricultural employment rises from large and negative (around −0.75 for China&amp;rsquo;s current employment share of 40–50%) toward zero; (c) the correlation between total employment and GDP rises from about 40% to nearly 100%; (d) the volatility of employment relative to GDP rises toward the level of mature economies. All four patterns match the cross-country empirical patterns documented in Figure 5.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-labor-push-versus-labor-pull-debate-imply-for-the-estimated-elasticity-and-how-is-it-resolved"&gt;Q7. What does the labor-push versus labor-pull debate imply for the estimated elasticity, and how is it resolved?&lt;/h3&gt;
&lt;p&gt;With epsilon &amp;gt; 1 (gross substitutes), nonagricultural TFP growth attracts labor from agriculture (labor pull), whereas agricultural TFP growth keeps workers on farms and slows structural change. With epsilon &amp;lt; 1 (complements), agricultural TFP growth would instead push workers into industry. The structural estimate epsilon = 3.6 &amp;gt; 1 strongly favors the labor-pull interpretation. This is confirmed by the Green Revolution evidence: Foster and Rosenzweig (2004), Moscona (2018), Bustos et al. (2016), and Jayachandran (2006) all find that positive agricultural TFP shocks slow industrialization and expand agricultural employment—consistent with epsilon &amp;gt; 1 and inconsistent with epsilon &amp;lt; 1.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-are-run-on-the-business-cycle-model"&gt;Q8. What robustness checks are run on the business cycle model?&lt;/h3&gt;
&lt;p&gt;Four robustness exercises: (1) Low elasticity epsilon = 0.5 with a large food subsistence level—this version fails to generate the observed countercyclicality of the productivity gap and implies an empirically incorrect response to agricultural TFP shocks. (2) Sectoral capital adjustment costs (quadratic, kappa = 2.5)—improves the cyclical behavior of aggregate employment and consumption but makes investment too smooth. (3) Raising the persistence of traditional-sector TFP shocks to match that of modern agriculture (phi_S = phi_AM = 0.90)—reduces aggregate labor volatility and makes the relative volatility of employment monotonically increasing with development. (4) Orthogonal shocks (zero cross-sector correlation)—results are negligibly different from the benchmark. These exercises indicate that the qualitative conclusions are robust across specifications.&lt;/p&gt;
&lt;h3 id="q9-how-is-the-productivity-gap-between-nonagriculture-and-agriculture-generated-by-the-model-and-does-it-match-the-data"&gt;Q9. How is the productivity gap between nonagriculture and agriculture generated by the model, and does it match the data?&lt;/h3&gt;
&lt;p&gt;In the model, the productivity gap (nonagricultural output per worker divided by agricultural output per worker) declines with development because the traditional, labor-intensive sector shrinks, raising average labor productivity in agriculture. This is both a long-run trend prediction and a business-cycle prediction: positive TFP shocks to nonagriculture draw workers from the traditional sector, raising agricultural capital intensity and productivity, thereby reducing the gap. The model successfully captures the falling trend in the productivity gap for China. The correlation between the HP-filtered productivity gap and nonagricultural employment in the model is −0.74, close to the empirical value of −0.54 for China. The model predicts lower volatility of the productivity gap than observed in the data.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-estimated-role-of-nonhomothetic-preferences"&gt;Q10. What is the estimated role of nonhomothetic preferences?&lt;/h3&gt;
&lt;p&gt;The authors extend the baseline homothetic CES model to allow Stone-Geary preferences (agricultural good as a necessity). The estimated subsistence level c-bar corresponds to only 11% of agricultural production in 1985, making the income effect through nonhomotheticity quantitatively small. The estimated epsilon falls only marginally when Stone-Geary preferences are introduced. The remaining structural parameters are virtually unchanged. The authors interpret this as evidence that, at the macroeconomic level, technological factors (TFP growth differences and capital accumulation) rather than nonhomothetic preferences are the primary drivers of structural change in China—a finding consistent with Alvarez-Cuadrado and Poschke (2011).&lt;/p&gt;
&lt;h3 id="q11-how-does-this-paper-relate-to-acemoglu-and-guerrieri-2008-and-herrendorf-et-al-2013"&gt;Q11. How does this paper relate to Acemoglu and Guerrieri (2008) and Herrendorf et al. (2013)?&lt;/h3&gt;
&lt;p&gt;The model builds on Acemoglu and Guerrieri (2008) in having capital deepening and differential TFP growth drive reallocation from agriculture to nonagriculture, but adds the traditional sector (absent in Acemoglu-Guerrieri), which generates the Lewisian surplus-labor mechanism and the declining productivity gap. With respect to Herrendorf et al. (2013): their three-sector CES model imposes a common elasticity across agriculture, manufacturing, and services, yielding a near-Leontief (epsilon near zero) estimate for the US. The authors show this estimate is an artifact of the symmetry restriction: when that restriction is relaxed, the agriculture-nonagriculture elasticity is large (2.32–2.49 for the US) while the manufacturing-services elasticity is near zero. The asymmetric three-sector estimates for the US (2.49), Japan (1.58), and China (1.70) are all above unity at the 1% significance level.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-main-limitations-and-open-questions"&gt;Q12. What are the main limitations and open questions?&lt;/h3&gt;
&lt;p&gt;The paper explicitly identifies several limitations: (1) the business cycle analysis is restricted to productivity (TFP) shocks only and does not include demand shocks; (2) the model is closed-economy and ignores trade; (3) the distinction between traditional and modern agriculture is not directly observed in the data—the traditional sector&amp;rsquo;s TFP process is estimated indirectly, introducing potential measurement error that may exaggerate the volatility and understate the persistence of traditional-sector shocks; (4) the prediction that agricultural value added is positively correlated with nonagricultural labor (and negatively with agricultural labor) is inconsistent with Chinese data, a failure the paper acknowledges. Future work is flagged on demand shocks and open-economy extensions.&lt;/p&gt;
&lt;h3 id="q13-what-cross-country-empirical-evidence-beyond-china-is-presented"&gt;Q13. What cross-country empirical evidence beyond China is presented?&lt;/h3&gt;
&lt;p&gt;Using ILO sectoral employment data for 63–66 countries over 1970–2015 (requiring at least 15 consecutive years of observations), the authors document: the correlation between agricultural and nonagricultural HP-filtered employment shifts from positive for countries with small agricultural sectors to strongly negative for countries with large sectors; the correlation between total employment and GDP declines monotonically with the agricultural employment share; the productivity gap is negatively correlated with nonagricultural employment in countries with large agricultural sectors (correlation of −0.54 for China) but near zero in mature economies; consumption volatility relative to GDP declines with development. The US historical time series (1929–2015) shows that before 1960 NBER recessions were associated with reversals in structural change—mirroring today&amp;rsquo;s China—while this pattern ceased after 1960.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Traditional agriculture (subsistence sector)&lt;/strong&gt;: A sub-sector of the agricultural sector that uses only labor (no capital) and produces an imperfect substitute for modern agricultural output. Its presence generates a reserve pool of labor that can move to nonagriculture at low marginal cost, creating the Lewisian surplus-labor property within a neoclassical framework. As the economy develops, this sector is crowded out by capital-intensive modern agriculture.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Modern agriculture&lt;/strong&gt;: A Cobb-Douglas sub-sector within agriculture that uses both capital and labor. Its expansion—crowding out the traditional sector—constitutes the modernization of agriculture. As workers leave the traditional sector, average capital intensity and labor productivity in agriculture rise, generating the procyclical productivity-gap pattern observed in developing economies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Asymptotic Balanced Growth Path (ABGP)&lt;/strong&gt;: The long-run equilibrium toward which the model economy converges, characterized by a fully modernized (traditional sector vanished), small agricultural sector, constant growth rates of sectoral capitals, and standard neoclassical business cycle properties. The paper establishes conditions under which the ABGP is asymptotically stable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor wedge (tau)&lt;/strong&gt;: An exogenous, time-invariant tax on nonagricultural wages that prevents equalization of marginal products of labor across sectors, standing in for a variety of frictions (migration barriers, rural overpopulation, institutional barriers) that keep agriculture inefficiently large. Its presence means that positive TFP shocks to nonagriculture both raise productivity directly and reduce misallocation by drawing workers out of the oversized agricultural sector.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Elasticity of substitution between agriculture and nonagriculture (epsilon)&lt;/strong&gt;: The elasticity governing substitution between agricultural and nonagricultural goods in aggregate CES production. When epsilon &amp;gt; 1 (gross substitutes, as estimated: epsilon = 3.6 for China), positive TFP shocks to nonagriculture pull labor from agriculture (labor-pull structural change), while positive shocks to agriculture slow structural change—consistent with Green Revolution evidence. When epsilon &amp;lt; 1 (complements), the opposite holds, implying counterfactual predictions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Productivity gap&lt;/strong&gt;: The ratio of average labor productivity in nonagriculture to average labor productivity in agriculture. In the model and the data this gap declines over the course of development (because agriculture modernizes and raises its average productivity) and also narrows during booms in countries undergoing structural change (because booms draw workers from low-productivity traditional agriculture). The model relates the gap formally to the ratio of labor income shares: APLM/APLG = (1−tau) × (LISM/LISA)^(−1).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sullying effect of recessions on agriculture&lt;/strong&gt;: The paper&amp;rsquo;s terminology for the pattern—documented empirically for China by Zhang et al. (2001)—whereby recessions induce workers to return to or remain in the agricultural sector, reversing structural change and lowering average agricultural productivity. This is the cyclical analog of the Lewisian adjustment: in downturns, the labor buffer of traditional agriculture absorbs displaced workers, cushioning aggregate employment but impairing agricultural productivity.&lt;/p&gt;</description></item><item><title>Carbon Pricing and Inequality: A Normative Perspective</title><link>https://macropaperwarehouse.com/papers/carbon-pricing-and-inequality-a-normative-perspective/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/carbon-pricing-and-inequality-a-normative-perspective/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper quantifies the sources and distributional consequences of unexpected carbon price changes for European households using a money-metric welfare framework. The motivation is stark: while carbon taxes enjoy broad support among economists, they face persistent public opposition — exemplified by Australia&amp;rsquo;s 2014 repeal, France&amp;rsquo;s 2018 Yellow Vest protests, and the 2025 rollback of Canada&amp;rsquo;s consumer carbon tax. The authors ask whether average welfare losses are unusually large, and whether the burden falls disproportionately on vulnerable groups, both questions with direct implications for understanding and reducing political resistance.&lt;/p&gt;
&lt;p&gt;The empirical approach rests on the &amp;ldquo;feasible set approach&amp;rdquo; of Del Canto et al. (2025), which applies the Envelope Theorem to show that the first-order welfare impact of a shock on any household is fully summarized by how the shock changes the discounted present value of their future budget sets — through consumption-basket prices, labor income, financial wealth (asset prices and dividends), and government transfers. This money-metric welfare change is preference-free up to first order: behavioral responses drop out, and the measure is independent of specific utility-function assumptions. The framework is appropriate for policy shocks (supply-side) but not for preference shocks.&lt;/p&gt;
&lt;p&gt;The geographic focus is euro-area countries (excluding the Netherlands and Austria due to data gaps) over 1999–2019. The identification strategy follows Känzig (2023): high-frequency shifts in EU ETS carbon futures prices around regulatory events affecting allowance supply are used as instruments in an external-instruments VAR to isolate plausibly exogenous carbon policy shocks. These shocks are then projected onto a wide array of household-level outcomes using local projections (Jordà 2005). The normalization throughout is a 1% increase in the HICP energy component on impact, which corresponds to roughly a 2.5-euro (or about 20%) increase in EU ETS carbon prices. Cross-sectional household budget data come from three Eurostat/ECB surveys: the Household Budget Survey (HBS, 2015 wave) for consumption baskets, EU-SILC (from 2004) for labor and transfer income by demographic group, and the Household Finance and Consumption Survey (HFCS) for household portfolio positions. Demographics are grouped by four age brackets (25–34, 35–49, 50–64, 65+), two education levels (college vs. non-college), three income brackets (bottom quartile = low, middle 50% = mid, top quartile = high), and four geographic regions (Southern, Western, Northern, Eastern Europe).&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. First, aggregate welfare losses are large: a 1% carbon-policy-induced energy price increase causes an average welfare loss of approximately 250 euros, corresponding to about 0.5% of a household&amp;rsquo;s three-year consumption (68% confidence band: 0.06% to 0.94%). Second, decomposing by channel, the direct consumption-price effect accounts for 0.19% of three-year consumption (68% CI: 0.02% to 0.35%); the labor income channel for 0.43% (68% CI: –0.08% to 0.93%); the portfolio channel for –0.04% (a welfare gain; 68% CI: –0.10% to 0.01%); and the transfer income channel for –0.07% (a welfare gain; 68% CI: –0.15% to 0.02%). Labor income is thus the dominant driver — both in aggregate and in the distributional patterns.&lt;/p&gt;
&lt;p&gt;Third, distributional heterogeneity is pervasive and statistically significant (joint F-tests reject uniformity with p-value = 0.00 across all demographic groupings). Non-college-educated households bear welfare losses of roughly 0.6% of three-year consumption, versus roughly 0.3% for college graduates — a gap concentrated in the labor income channel, not the consumption channel (which is broadly similar across groups at around 0.2%). By income, the pattern is U-shaped: young, low-income households suffer the largest losses, exceeding 1% of three-year consumption, while middle-income and older households are the most insulated; high-income households also experience significant losses (around the 0.5% average), driven by their own labor income exposure. Households aged 65 and over suffer welfare losses of only around 0.15%, largely because they are retired from the labor market.&lt;/p&gt;
&lt;p&gt;Fourth, regional heterogeneity is stark. Southern Europe bears the highest burden, with welfare losses of 0.5% to 0.8% for working-age households; Eastern Europe also faces substantial losses; Western Europe stands at around 0.2% to 0.3%; Northern Europe is the most insulated, with losses below 0.2% and not statistically significant. The labor income channel is the primary driver of these regional differences, consistent with more rigid labor markets in Southern and Eastern Europe (stronger employment protection, less flexible wage-setting). Northern Europe is protected partly by its high share of renewable energy, which mutes the carbon-price pass-through. Eastern Europe benefited from disproportionate free ETS allowance allocations over the sample period, dampening direct price impacts.&lt;/p&gt;
&lt;p&gt;These results collectively suggest that public opposition to carbon taxes may stem from legitimate distributional concerns rather than mere ideological resistance or ignorance. The authors conclude with three policy implications: (1) compensation schemes focused only on consumption prices will be insufficient because the dominant channel is labor income; (2) expansionary (green) monetary policy could ease the income burden, though at some inflationary cost; and (3) redistribution should run from older to younger households, since working-age groups bear the disproportionate burden while retirees are largely insulated.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-carbon-policy-shock-and-what-are-the-main-threats-to-identification"&gt;Q1. What is the identification strategy for the carbon policy shock, and what are the main threats to identification?&lt;/h3&gt;
&lt;p&gt;The instrument is the high-frequency shift in EU ETS carbon futures prices around regulatory events affecting allowance supply (following Känzig 2023). The logic is that economic conditions are already priced in prior to the regulatory news, so futures-price movements in a tight window around those events reflect only policy surprises. This instrument is then used in an external-instruments VAR to identify a monthly structural carbon policy shock series (1999–2019). The local projections use 6 lags for monthly outcomes and 2 lags for quarterly outcomes, plus a linear trend and a dummy for the euro sovereign debt crisis (July 2011–March 2012). The main identification threats are: (a) if economic conditions are not fully priced into carbon futures before the regulatory events, the instrument could be correlated with macroeconomic conditions; (b) the framework assumes no preference shocks, which rules out COVID-style demand shifts; (c) the small-noise approximation underlying the feasible-set approach is less suitable for large aggregate shocks.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-feasible-set-approach-not-require-specific-preference-assumptions-and-what-are-its-limitations"&gt;Q2. Why does the feasible-set approach not require specific preference assumptions, and what are its limitations?&lt;/h3&gt;
&lt;p&gt;By the Envelope Theorem applied to household optimization, first-order welfare effects depend only on how the policy changes the prices and quantities in the household&amp;rsquo;s budget constraint — not on how preferences are shaped. Behavioral responses drop out at first order. The welfare metric is money-metric: the willingness-to-pay to avoid the shock, expressed in euros (income units). Limitations: (1) It is a small-noise approximation around a zero-risk limit; large aggregate shocks are not well-handled. (2) It is valid for shocks from the production or policy side but not for preference shocks (e.g., discount rate changes). (3) Accounting properly for idiosyncratic risk requires covariance weights (Theta terms in Proposition 1 of the appendix); Del Canto et al. (2025) estimate these at –0.1 to –0.4, implying somewhat attenuated welfare levels but no meaningful change to the distributional comparisons. (4) Carbon emissions-reduction benefits are excluded from the welfare calculation by design, since the paper focuses on the pecuniary costs side only.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-mechanism-behind-the-labor-income-channel-and-how-does-it-vary-across-demographic-groups-and-regions"&gt;Q3. What is the mechanism behind the labor income channel, and how does it vary across demographic groups and regions?&lt;/h3&gt;
&lt;p&gt;Carbon price increases raise production costs for energy-intensive sectors, reduce output and employment, and depress aggregate wages — a general equilibrium effect that transmits to household labor income over multiple quarters. The average labor income response peaks at around 1% below trend. For non-college-educated households the peak fall exceeds 1%, while for college graduates the response is more muted. By income group, low-income households face the sharpest falls — around 2–4% over the three-year horizon — whereas middle-income households fall by approximately 0.5–1% and high-income households by about 1%. These effects are larger than those estimated by Del Canto et al. (2025) for oil price shocks on US households (approximately 0.3% welfare loss from labor income after a 10% oil price increase), which the authors attribute to more rigid European labor markets: strong employment protection limits wage cuts but discourages hiring and prolongs unemployment spells, amplifying extensive-margin adjustments. In Southern and Eastern Europe, rigidities are most pronounced, generating the largest regional labor-income responses. Northern and Western Europe show more muted responses.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-role-of-the-portfolio-channel-and-who-gains-or-loses-through-it"&gt;Q4. What is the role of the portfolio channel, and who gains or loses through it?&lt;/h3&gt;
&lt;p&gt;Stock prices fall by a peak of about 5% and dividends decline by about 3% after a carbon policy shock. Bond prices initially decline then partially recover. House prices decline substantially but with a lag. The welfare effect of asset price changes depends on whether a household is a net buyer or net seller of the asset. Younger households in the accumulation phase gain from falling asset prices (they can buy cheaply); older households planning to dis-save lose. The portfolio channel is quantitatively modest: average welfare gain of about 0.04%, most pronounced for younger college-educated households. The channel is not large enough to offset labor income or consumption-price losses for any group.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-role-of-the-transfer-income-channel-and-which-groups-benefit-most"&gt;Q5. What is the role of the transfer income channel, and which groups benefit most?&lt;/h3&gt;
&lt;p&gt;Transfer income — which the paper splits into inflation-indexed pension income and other government transfers (unemployment, sickness, disability, education benefits) — generates a welfare gain of about 0.07% on average. Pensions are indexed to inflation and rise as carbon pricing lifts headline prices; this benefit accrues primarily to older households (aged 65+), who have large pension income. Other transfers show an increase post-shock but the responses are generally not statistically significant at conventional levels. High-income households show a negative transfer response. Northern and Southern Europe benefit more from the transfer channel, consistent with more generous welfare programs; Eastern Europe shows little or negative transfer response, consistent with weaker automatic stabilizers.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-u-shaped-pattern-of-welfare-losses-by-income-and-what-explains-it"&gt;Q6. What is the U-shaped pattern of welfare losses by income, and what explains it?&lt;/h3&gt;
&lt;p&gt;The paper finds that low-income and young households suffer the largest losses (exceeding 1% of three-year consumption), middle-income and older households are most insulated, and high-income households also face significant losses (broadly around the 0.5% average). The U-shape arises from the labor income channel: low-income households are concentrated in sectors and employment types most exposed to carbon pricing contractions; high-income households also have substantial labor income (in absolute terms) that contracts; middle-income households appear more buffered, possibly due to sector composition or greater employment stability. The consumption channel contributes approximately uniformly across income groups (around 0.2%), so does not generate the U-shape.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-differ-methodologically-from-prior-distributional-studies-of-carbon-taxes"&gt;Q7. How does this paper differ methodologically from prior distributional studies of carbon taxes?&lt;/h3&gt;
&lt;p&gt;Prior work such as Andersson and Atkinson (2020) and Beznoska et al. (2012) focused on direct consumption-price incidence, following Poterba (1989) and using static input-output methods or cross-sectional spending data to estimate first-round price effects. The present paper differs in three ways: (1) it instruments for unexpected carbon price shocks, isolating exogenous variation; (2) it incorporates indirect channels — labor income, asset prices, and transfers — in addition to direct consumption prices; (3) it estimates dynamic IRFs directly, capturing the persistence of effects over a three-year horizon. The key novel finding is that indirect labor income effects are the dominant driver of both the level and the distribution of welfare losses, and that neglecting these indirect channels substantially understates both the size and the regressiveness of carbon pricing.&lt;/p&gt;
&lt;h3 id="q8-why-are-regional-differences-in-welfare-loss-so-large-and-what-drives-northern-europes-relative-insulation"&gt;Q8. Why are regional differences in welfare loss so large, and what drives Northern Europe&amp;rsquo;s relative insulation?&lt;/h3&gt;
&lt;p&gt;Regional differences are driven primarily by differential pass-through from carbon prices to consumer prices and by differential labor market rigidity. Northern Europe sources a large share of energy from renewables, so a carbon price increase has a smaller pass-through to domestic energy costs. Eastern Europe was allocated disproportionate free ETS allowances over the 1999–2019 sample period, also dampening direct price impacts — consistent with Känzig and Konradt (2024). Southern and Eastern Europe have more rigid labor markets (stronger employment protection, less flexible wage-setting), amplifying the labor-income contraction. Northern and Western Europe have more flexible labor markets. Additionally, Northern and Southern Europe have more generous welfare programs that partially cushion losses via the transfer channel; Eastern Europe lacks this buffer.&lt;/p&gt;
&lt;h3 id="q9-what-data-sources-does-the-paper-combine-and-what-are-the-key-sample-restrictions"&gt;Q9. What data sources does the paper combine, and what are the key sample restrictions?&lt;/h3&gt;
&lt;p&gt;The paper combines three Eurostat/ECB household surveys: (1) the Household Budget Survey (HBS), 2015 wave, for consumption basket shares by COICOP categories for demographic groups; (2) EU-SILC (2004 onward for some countries, 2005 for most) for annual labor income and transfer income time series by group, converted to quarterly frequency via Chow-Lin interpolation; (3) HFCS (conducted every 4 years by the ECB) for household portfolio positions. Time-series macro data on HICP components, house prices, bond prices, stock prices, and dividends come from Eurostat and ECB/Bloomberg. The sample covers euro-area countries (excluding Netherlands and Austria for data reasons) over 1999–2019. Households are restricted to ages 25–75; top and bottom 1% by net worth are excluded from portfolio statistics. The base year for all life-cycle variables is 2015.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Three main implications are drawn: (1) Public resistance to carbon taxes is not merely ideological — the estimated welfare losses are sizable (about 0.5% of three-year consumption for a 1% energy-price increase), so opposition reflects genuine economic concerns. (2) Standard compensation via energy-bill rebates or consumption-basket adjustments is insufficient because the dominant channel is labor income (0.43% vs. 0.19% for consumption). Compensation schemes should include labor-market policies; the authors also suggest expansionary (green) monetary policy as a tool to ease the income burden, though at some inflationary cost. (3) The intergenerational dimension is important: working-age households (especially young, less-educated, lower-income ones) bear the brunt while retirees are largely shielded. Redistribution should run from old to young, not just from rich to poor. Scope conditions: the estimates are derived from the EU ETS context (European carbon market, euro area, 1999–2019), rely on a small-shock linear approximation, and focus on short-to-medium-run impacts (three-year horizon). The benefits of reduced carbon emissions are excluded from the welfare calculation.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-handle-inference-given-the-short-time-series-and-estimation-uncertainty"&gt;Q11. How does the paper handle inference given the short time series and estimation uncertainty?&lt;/h3&gt;
&lt;p&gt;The sample runs from 1999 to 2019, which is relatively short for the IRF exercises. The paper reports 68% and 90% confidence bands throughout (rather than the conventional 95%), using the lag-augmentation approach of Montiel Olea and Plagborg-Møller (2021) to account for serial correlation. For the money-metric welfare calculations, inference uses a parametric bootstrap that draws from the estimated distribution of IRFs (assuming block-wise uncorrelatedness across variables, justified by low cross-residual correlations averaging 0.16). Cross-sectional group shares are treated as given. The authors explicitly acknowledge considerable uncertainty: the 68% confidence band on the aggregate welfare loss spans 0.06% to 0.94%. They conduct joint F-tests for homogeneity of welfare effects across demographic groups; in all cases the null is rejected with p-value = 0.00. Only 68% bands are reported for welfare calculations given short sample and estimation uncertainty.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-heterogeneous-labor-income-irf-magnitudes-for-different-groups-and-are-they-statistically-significant"&gt;Q12. What are the heterogeneous labor income IRF magnitudes for different groups, and are they statistically significant?&lt;/h3&gt;
&lt;p&gt;Average labor income falls by about 1% at the peak (imprecisely estimated). Non-college-educated peak fall exceeds 1%; college-educated peak fall is more muted. By income group: low-income households see falls of roughly 2–4% over three years; high-income households see a fall of about 1%; middle-income households fall by approximately 0.5–1%. These effects are noted to be larger than analogous results for oil shocks in the US (Del Canto et al. 2025), attributed to European labor market rigidity. The responses are described as featuring &amp;lsquo;a considerable degree of persistence but only imprecisely estimated&amp;rsquo; at the average level. The welfare calculations based on these IRFs have wide confidence bands, reflecting this imprecision.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-consumer-price-dynamics-following-a-carbon-policy-shock"&gt;Q13. What are the consumer price dynamics following a carbon policy shock?&lt;/h3&gt;
&lt;p&gt;Energy prices (HICP energy component) rise by 1% on impact and remain elevated for approximately one year before returning toward baseline. Housing and utilities experience a significant, persistent increase, remaining approximately 0.5% above baseline three years after the shock. Transport prices increase by 0.5% on impact but revert within a year. Food prices rise to a lesser extent. Restaurants and hotels, recreation and culture, and clothing also show significant impact-period increases, though most effects become insignificant after 12 months. Two exceptions at 12 months: housing and utilities remain significantly elevated; education and communication prices actually fall, possibly reflecting adverse general-equilibrium wage and employment effects.&lt;/p&gt;
&lt;h3 id="q14-how-is-the-welfare-analysis-limited-to-short-to-medium-run-effects-and-what-longer-run-effects-are-left-unaddressed"&gt;Q14. How is the welfare analysis limited to short-to-medium-run effects, and what longer-run effects are left unaddressed?&lt;/h3&gt;
&lt;p&gt;The welfare calculations are restricted to a three-year horizon because statistical power in the local projections declines beyond that point given the available sample (1999–2019). The paper explicitly notes that the estimates may miss unemployment hazard effects (i.e., transitions into and out of employment), borrowing cost effects induced by carbon taxes, and any long-run structural adjustments (sectoral reallocation, green investment, capital formation). The benefits of reduced carbon emissions — which may be very large in welfare terms but are realized over much longer horizons — are also excluded by design.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Feasible Set Approach&lt;/strong&gt;: A welfare-measurement methodology (from Del Canto et al. 2025) that applies the Envelope Theorem to show that the first-order welfare impact of any shock on a household equals the change in the discounted present value of that household&amp;rsquo;s budget set — encompassing consumption prices, labor income, asset income, and transfers. The measure is preference-free at first order and is expressed in money-metric (income-equivalent) units.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Money-Metric Welfare Loss&lt;/strong&gt;: In this paper, the number of euros a household would be willing to pay to avoid exposure to the carbon policy shock, computed as a share of total three-year consumption. It is derived from the feasible-set formula and expressed in income units, making it directly interpretable and comparable across demographic groups without requiring preference parameters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Carbon Policy Shock&lt;/strong&gt;: An exogenous, unexpected change in carbon prices driven by regulatory events affecting the supply of EU ETS emission allowances, identified via high-frequency shifts in carbon futures prices around those events used as instruments in an external-instruments VAR. Distinguished from demand-driven carbon price fluctuations correlated with the business cycle.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor Income Channel&lt;/strong&gt;: The indirect welfare effect of a carbon price shock that operates through general-equilibrium changes in aggregate wages and employment. It is the dominant welfare channel in the paper (0.43% of three-year consumption on average, versus 0.19% for direct consumption-price effects), and the primary driver of both the aggregate welfare loss and the distributional heterogeneity across education, income, and regional groups.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption Channel (Direct Effect)&lt;/strong&gt;: The welfare impact arising from higher prices for goods in the household&amp;rsquo;s consumption basket following a carbon price increase. Weighted by the household&amp;rsquo;s nominal expenditure on each good. Broadly similar across demographic groups (clustering around 0.2% of three-year consumption), so it does not generate the observed distributional heterogeneity — in contrast to the labor income channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Portfolio Channel&lt;/strong&gt;: The welfare effect transmitted through changes in asset prices (equities, bonds, housing) after a carbon shock. The sign depends on whether a household is a net buyer or net seller of the asset: younger households in the accumulation phase gain from falling asset prices; older households in the dis-saving phase lose. Quantitatively small on average (net welfare gain of about 0.04%), most pronounced for younger, college-educated households.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Transfer Channel&lt;/strong&gt;: The welfare effect operating through government transfer income (unemployment and other social benefits) and inflation-indexed pension payments. Because pensions are indexed to the price level, carbon-induced inflation raises pension income and benefits older households. Other transfer income tends to rise post-shock but the responses are generally imprecisely estimated. On average the channel generates a modest welfare gain (about 0.07% of three-year consumption), primarily for the elderly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Greenflation&lt;/strong&gt;: The phenomenon, documented empirically by Bettarelli et al. (2025) and referenced in this paper, whereby carbon-tax shocks contribute to broader consumer price inflation beyond the direct energy-price impact — through pass-through to housing, transport, food, and other categories, and by raising inflation expectations and triggering tighter monetary policy, which in turn depresses bond and house prices.&lt;/p&gt;</description></item><item><title>Diet, Economic Development and Climate Change</title><link>https://macropaperwarehouse.com/papers/diet-economic-development-and-climate-change/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/diet-economic-development-and-climate-change/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Food production accounts for roughly one-third of global greenhouse gas (GHG) emissions, and richer nations contribute disproportionately through meat-intensive diets and input-intensive farming. This paper asks how much of that disparity will be exported to the developing world as it grows, and which policies can most cost-effectively reduce agricultural emissions during that transition. The answer requires separately identifying two distinct channels—demand-side dietary change and supply-side technological change—and tracing their general equilibrium consequences through global food markets.&lt;/p&gt;
&lt;p&gt;The authors build a quantitative multi-country general equilibrium model calibrated to 90 countries (plus a rest-of-world aggregate) and 47 food products for 2010. The demand side features nested non-homothetic CES preferences, which allow income elasticities to differ across food products—the core mechanism of the nutrition transition. The supply side, built on Farrokhi and Pellegrina (2023), operates at a granular grid-cell level covering the Earth&amp;rsquo;s surface, with producers on each plot choosing both which crop to grow and whether to use a modern, input-intensive (higher-GHG) technology or a traditional, labor-intensive one—the core mechanism of agricultural modernization. GHG emissions are tracked from both production and transportation. Data on calorie intake come from FAO Food Balance Sheets; emissions from Poore and Nemecek (2018) and EDGAR-FOOD; yields from FAO-GAEZ (approximately 1.1 million fields).&lt;/p&gt;
&lt;p&gt;A key methodological contribution is an identification result for income elasticities that requires no price data. In open-economy models, trade shares provide a sufficient statistic for consumer prices, so the model&amp;rsquo;s implicit Marshallian demand equations can be estimated using only expenditure shares and bilateral trade flows—a cleaner identification than prior closed-economy approaches. Structural elasticity estimates are validated against reduced-form regressions that regress product-level log absorption on log GDP per capita interacted with the product&amp;rsquo;s GHG intensity; the cross-method correlation has a slope of 0.64–0.77 and R² of 0.93–0.95.&lt;/p&gt;
&lt;p&gt;Four empirical patterns motivate the model. First, diet composition alone drives large variation in emissions: if the whole world adopted the US diet (holding total calories fixed), the food share of global GHG emissions would rise from 30% to 42%; adopting the Argentinian diet would raise it to 74%; adopting the Ethiopian diet would lower it to 12%. Second, GHG emissions per capita from food rise strongly with GDP per capita (elasticity 0.39 in the cross-section); about one-third of this is a pure scale effect (more calories) and two-thirds is a compositional shift toward higher-emission foods (elasticity of emissions per calorie with respect to GDP per capita is 0.23–0.28). Third, products with higher GHG emissions per calorie have higher income elasticities; a 1% rise in a product&amp;rsquo;s GHG intensity is associated with a 0.17–0.21% higher income elasticity, robust to excluding all meat products. Fourth, emissions from fertilizers and energy use as a share of total agricultural emissions rise with GDP per capita (slope 0.82), indicating that agricultural modernization independently amplifies GHG emissions within each crop.&lt;/p&gt;
&lt;p&gt;Model decompositions reveal that about two-thirds of the cross-sectional correlation between food emissions per capita and GDP per capita is attributable to intrinsic dietary preferences (culture, religion, demographics) rather than to income itself, and about one-half of the correlation for emissions per calorie. This implies that the causal effect of economic growth on emissions is substantially smaller than raw correlations suggest.&lt;/p&gt;
&lt;p&gt;Policy counterfactuals (Table 4) are the paper&amp;rsquo;s centerpiece. A uniform 10% TFP shock across all modern agricultural, non-agricultural, and input producers raises global welfare by 14.9% and increases global agricultural GHG emissions by 5.0% (approximately 0.6 Gt CO₂ from production, 0.004 Gt from transport). Shutting down the nutrition transition channel reduces this emission increase by 28%; shutting down agricultural modernization reduces it by a further 16%; shutting both down reduces it by 42%—so the two mechanisms together account for more than one-third of the growth-induced emission increase. Crucially, ignoring general equilibrium supply responses would overstate the emission impact of economic growth by 100%: higher food demand raises production prices, which dampens both consumption growth and further technology adoption.&lt;/p&gt;
&lt;p&gt;For dietary restrictions: a global no-beef mandate would reduce agricultural GHG emissions by 20%, at a global welfare cost of 0.6%, with large concentrated losses in major beef-producing and consuming countries (Argentina −3–5%; Uruguay −4%). A global vegetarian mandate would reduce emissions by 30% (approximately the same 20% figure is given in the abstract with apparent inconsistency but Table 4 column 3 shows −20% for no-beef and −30% for vegetarian), at a welfare cost of 2.8% globally and with greater inequality impacts for developing countries. Back-of-the-envelope calculations that ignore general equilibrium overstate the emission reductions from dietary restrictions by roughly one-third.&lt;/p&gt;
&lt;p&gt;For food trade policy: raising trade costs enough to cut transportation emissions by 75% reduces total agricultural GHG emissions by 11.9%, but at a global welfare cost of 17.8%—a ratio far worse than dietary policies. The welfare loss is highly unequal: countries in the bottom quartile of the GDP per capita distribution face welfare losses of up to 41% (the abstract states this figure; Table 4 col. 2 shows the Q4/Q1 inequality worsening by 4.9 percentage points in the eat-local scenario). The conclusion is that dietary policies dominate food trade policies on both effectiveness and equity grounds.&lt;/p&gt;
&lt;p&gt;Transportation emissions account for only about 5% of agricultural GHG (0.7 Gt CO₂ vs. 16.5 Gt from production), so policies targeting transport emissions alone have limited aggregate impact.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-for-income-elasticities-and-why-is-it-novel"&gt;Q1. What is the core identification strategy for income elasticities, and why is it novel?&lt;/h3&gt;
&lt;p&gt;Standard non-homothetic CES estimation requires price data because the demand equation depends on price indices. In a closed economy this problem is severe. The authors show that in an open economy, bilateral trade shares provide a sufficient statistic for variety price indices: averaging trade shares across a country&amp;rsquo;s import partners yields a geometric mean of production prices that can be differenced out using fixed effects. The key estimating equation (40) regresses an adjusted expenditure share on log income per capita, with fixed effects absorbing production-price variation through the set of import partners. No price data is needed. This is exact—not an approximation—unlike the approximate methods in Comin et al. (2021) or Caron and Fally (2022), which either impose additional assumptions about price variation across consumer groups or require proxies for crop-specific trade costs such as gravity variables.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-threats-to-identification-and-how-are-they-addressed"&gt;Q2. What are the main threats to identification and how are they addressed?&lt;/h3&gt;
&lt;p&gt;The key concern is that income is correlated with prices and preference shifters that also affect food expenditure shares. In the reduced-form regressions (equation 1), country-year and product-year fixed effects control for country-specific factors (including regional technology change) and global product-specific factors (including product-specific technological progress). In the structural estimation (equation 40), the model&amp;rsquo;s functional form is used to control fully for endogeneity arising through prices, since trade shares substitute out unobservable price indices exactly. The close agreement between reduced-form and structural income elasticity estimates (slope 0.64–0.77, R² 0.93–0.95 in cross-validation) is reassuring that the two quite different identifying assumptions yield similar results. One remaining concern is unobservable preference shifters (ai,k and ã_i,s), which appear as residuals; identification requires income variation orthogonal to these shifters, and the authors follow the precedent of assuming fixed effects are sufficient. Household-level data from Brazil&amp;rsquo;s Consumer Expenditure Survey (POF) bolster the reduced-form patterns using within-country income variation.&lt;/p&gt;
&lt;h3 id="q3-how-are-the-nutrition-transition-and-agricultural-modernization-distinguished-empirically-and-in-the-model"&gt;Q3. How are the nutrition transition and agricultural modernization distinguished empirically and in the model?&lt;/h3&gt;
&lt;p&gt;These are fundamentally different economic mechanisms. The nutrition transition operates through demand: as incomes rise, consumers shift toward food products that, for reasons of taste or nutrition, happen to have higher GHG emissions per calorie. It is a between-product phenomenon captured by non-homothetic income elasticities. Agricultural modernization operates through supply: as wages rise, producers substitute away from labor-intensive traditional technologies toward input-intensive modern technologies (fertilizers, machinery) that emit more GHG per calorie of output, for any given crop. It is a within-product phenomenon captured by the endogenous technology-choice margin in the agricultural production model. In the counterfactual decompositions, the authors shut down each channel independently: the nutrition transition is shut down by setting all within-sector income elasticity parameters (ε_k) equal; agricultural modernization is shut down by fixing the land share in each technology exogenously. Doing so reveals that the nutrition transition accounts for 28% and modernization for 16% of the emission increase from a 10% TFP shock (jointly 42%), with the remainder attributable to scale effects and general equilibrium price responses.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-role-of-general-equilibrium-supply-responses-and-why-do-they-matter-so-much"&gt;Q4. What is the role of general equilibrium supply responses and why do they matter so much?&lt;/h3&gt;
&lt;p&gt;A central finding is that ignoring supply-side equilibrium price responses would overstate the emission impact of economic growth by 100%. The mechanism is straightforward: economic growth raises income and thus food demand, which pushes up production prices (because agricultural supply is upward-sloping due to limited land and heterogeneous productivity across grid cells). Higher prices dampen consumption, which partially offsets the demand-driven emission increase. For dietary restriction policies, back-of-the-envelope calculations that simply remove the GHG attributable to banned food products overstate the emission reduction by roughly one-third, because consumers substitute toward other food products and global agricultural production reorganizes. The model&amp;rsquo;s general equilibrium structure is therefore essential for obtaining credible policy counterfactuals, and a main conclusion of the paper is that the literature&amp;rsquo;s existing back-of-the-envelope calculations in environmental science substantially overstate both the emission risks from growth and the emission benefits from dietary policies.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented-across-countries-and-products"&gt;Q5. What heterogeneity is documented across countries and products?&lt;/h3&gt;
&lt;p&gt;Across countries: diet composition varies enormously. Counterfactual calculations show that if all countries adopted the Argentinian diet (holding total calories fixed), the global food share of total emissions would rise to 74%; adopting the Ethiopian diet would lower it to 12%, compared to the factual 30%. The income elasticity of the agricultural sector as a whole is 0.39, close to Comin et al. (2021)&amp;rsquo;s 0.37. Rich countries have a higher share of modern technology in production, higher fertilizer and energy use per unit of land, higher food GHG per capita, and higher food GHG per calorie. About two-thirds of the cross-sectional gradient in food GHG per capita is attributable to intrinsic preferences rather than income per se. Religion is documented as one driver: Islamic-majority countries show lower preference for pork; Hindu-majority countries show higher preference for lamb, mutton, and poultry relative to other meats. Across products: GHG emissions per 1,000 kcal range from above 35 kg CO₂ for beef and coffee to below 5 kg CO₂ for wheat and rye. Income elasticity parameters (ε_k) range from lowest for staples (yams, sweet potatoes, millet, sorghum, rice) to highest for luxury fruits and vegetables (berries, asparagus, cucumbers, watermelon). Notably, the income-GHG gradient persists after excluding all meat products: vegetables and fruits have higher GHG per calorie than staples, so the nutrition transition is broader than a simple meat-consumption story.&lt;/p&gt;
&lt;h3 id="q6-how-do-the-diet-restriction-and-food-trade-policy-counterfactuals-compare-on-welfare-and-effectiveness"&gt;Q6. How do the diet restriction and food trade policy counterfactuals compare on welfare and effectiveness?&lt;/h3&gt;
&lt;p&gt;Diet restriction (no-beef): global GHG emissions fall 20%, global welfare falls 0.6%. The welfare effect is highly concentrated—Argentina experiences −3–5% welfare loss, Uruguay approximately −4% in the no-beef scenario, because they are large meat producers and exporters. Inequality between rich (Q4) and poor (Q1) countries worsens by 1.0 percentage point. Diet restriction (vegetarian): global GHG emissions fall 30%, global welfare falls 2.8%. Inequality worsens by 6.0 percentage points, indicating developing countries bear more of the cost because a larger share of their income goes to food, and their income sources (agriculture) are more directly affected. Food trade policy (&amp;rsquo;eat local&amp;rsquo;, raising trade costs to cut transportation emissions by 75%): global GHG emissions fall 11.9%, but global welfare falls 17.8%—roughly 25–30 times the welfare cost per percentage point of emission reduction compared to dietary policies. Inequality worsens substantially more: Q4/Q1 ratio worsens by 4.9 percentage points. Countries in the bottom GDP quartile face welfare losses up to 41%. The paper concludes that dietary restrictions are both substantially more effective in reducing GHG emissions and far more equitable in their welfare consequences than food trade policies.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-share-of-agricultural-ghg-from-transportation-versus-production-and-what-are-the-implications"&gt;Q7. What is the share of agricultural GHG from transportation versus production, and what are the implications?&lt;/h3&gt;
&lt;p&gt;In the 2010 data, GHG emissions from food transportation account for approximately 5% of total agricultural GHG (0.7 Gt CO₂ out of approximately 17.2 Gt total). Production accounts for 95% (16.5 Gt CO₂). This has two implications. First, in the economic growth counterfactual, transportation emissions increase by 2.2%, but because transportation is only 5% of total, its contribution to total emission growth (0.004 Gt) is negligible. Second, it implies that policies targeting food &amp;lsquo;food miles&amp;rsquo; or local eating are poorly targeted: even a dramatic 75% reduction in transportation emissions only mechanically eliminates 4.6% of total agricultural GHG, and the actual general equilibrium reduction (11.9%) comes mostly from production effects (agricultural trade restrictions reduce global production and consumption), accompanied by very large welfare costs.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-and-validation-exercises-are-conducted"&gt;Q8. What robustness checks and validation exercises are conducted?&lt;/h3&gt;
&lt;p&gt;The paper provides several validation exercises. (1) The reduced-form income elasticity regressions are run both with all crops and excluding all meat products (beef, lamb and mutton, pig meat, poultry), yielding nearly identical coefficients of 0.176 and 0.175 (columns 1 and 2 of Table 1), and with country-year and product-year fixed effects (columns 3–4), showing similar results across specifications. (2) The structural income elasticities are compared to the reduced-form estimates, with a cross-method slope of 0.64–0.77 and R² of 0.93–0.95, reassuring given the two methods make different identifying assumptions. (3) Model fit is checked against six untargeted empirical regularities (Figure 6): declining agricultural employment share, rising input cost share, rising modern technology land share, rising food GHG per capita, rising calories per capita, and rising food GHG per calorie—all with GDP per capita. The model matches the sign and approximate magnitude of each relationship. (4) Household-level estimates using Brazil&amp;rsquo;s POF survey replicate the cross-country finding that higher-GHG products have higher income elasticities, controlling for fixed effects, food price proxies, and excluding meat. (5) The decomposition of the cross-sectional income-emissions gradient shows that equalizing comparative advantage (column 3) or trade costs (column 4) across countries leaves the gradient approximately unchanged, supporting the focus on preferences and technology.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-prior-work-and-where-does-it-depart-from-it"&gt;Q9. How does this paper relate to prior work and where does it depart from it?&lt;/h3&gt;
&lt;p&gt;The paper sits at the intersection of several literatures. It builds on Farrokhi and Pellegrina (2023) for the granular grid-cell production model with technology choice; on Costinot, Donaldson, and Smith (2016) for the agricultural field structure; and on Comin, Lashkari, and Mestieri (2021) for non-homothetic CES preferences and the identification of income elasticities. Key departures: (a) Relative to Comin et al. (2021), the authors extend identification to nested CES preferences and to an open-economy without requiring price data—their method is exact rather than approximate. (b) Relative to the environmental science literature (e.g., Hoolohan et al., 2013; Perignon et al., 2017; Tilman et al., 2011), the paper endogenizes general equilibrium supply responses, which the authors show dramatically attenuate the effect of both income growth and dietary policies on emissions. (c) Relative to prior quantitative spatial models of climate change (e.g., Shapiro 2016 on trade costs and CO₂), this paper focuses on agricultural emissions specifically and introduces nutrition transition and technology choice. (d) The authors claim to be the first to analyze both dietary restrictions and food trade policies on agricultural emissions within quantitative trade models. (e) Relative to Chen et al. (2022), who use a computable general equilibrium model with general equilibrium supply adjustments, this paper includes far more food products (47 vs. their smaller set) and endogenizes technology choice, both of which are quantitatively important for capturing the nutrition transition.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-papers-mechanism-for-why-vegetable-and-fruit-consumption-also-raises-ghg-emissions-as-income-rises-even-without-meat"&gt;Q10. What is the paper&amp;rsquo;s mechanism for why vegetable and fruit consumption also raises GHG emissions as income rises, even without meat?&lt;/h3&gt;
&lt;p&gt;The paper notes in footnote 1 that the positive correlation between income elasticities and GHG emissions per calorie persists even when meat products are excluded from the sample (Table 1, columns 3–4). The reason is that vegetables and fruits—which become more preferred as countries grow richer—emit more GHG per calorie than staple foods such as yams and potatoes. Staples require little processing or refrigeration and are typically produced with traditional, low-input technologies. By contrast, fresh fruits and vegetables (especially high-value items such as berries, asparagus, grapes, and coffee) require more energy-intensive transportation, storage, and sometimes greenhouse production. This means that the nutrition transition generates rising emissions not merely through the beef channel emphasized in much of the public debate, but through a broader shift away from calorie-dense staples toward diverse, lower-calorie-density products that happen to have higher GHG footprints per calorie.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-model-imply-about-the-environmental-kuznets-curve-for-food-emissions"&gt;Q11. What does the model imply about the Environmental Kuznets Curve for food emissions?&lt;/h3&gt;
&lt;p&gt;The paper explicitly tests for and finds no evidence of an Environmental Kuznets Curve (EKC) in food emissions—that is, no inverse-U shape in which emissions per capita eventually decline as countries become very rich, as might be expected if wealthy nations adopt more sustainable diets or stricter environmental regulations. The income-emission relationship is found to be approximately log-linear across all levels of development (footnote 8). This is consistent with the broader empirical literature on the EKC (cited survey by Dinda, 2004). The implication is that there is no automatic &amp;lsquo;greening&amp;rsquo; of diets as countries develop; active policy intervention would be needed.&lt;/p&gt;
&lt;h3 id="q12-how-is-economic-development-modeled-in-the-policy-counterfactuals-and-what-are-the-scope-conditions"&gt;Q12. How is economic development modeled in the policy counterfactuals, and what are the scope conditions?&lt;/h3&gt;
&lt;p&gt;Economic development is modeled as a uniform 10% increase in TFP for three types of agents: (i) modern agricultural producers, (ii) non-agricultural producers, and (iii) agricultural input producers (fertilizers, machinery, pesticides). Traditional agricultural technology is not subject to productivity growth, following Gollin, Parente, and Rogerson (2007). This creates both income effects (via higher wages) and substitution effects (via changes in relative input prices that favor modern, input-intensive technology). The scope conditions are important: the results apply specifically to a uniform global TFP shock, not to individual-country development. For individual-country TFP shocks, the analytical decomposition (equation 34) shows that general equilibrium income spillovers to foreign countries can attenuate the nutrition transition if foreign incomes fall (e.g., due to terms-of-trade effects). The model does not incorporate dynamics (it is a static model calibrated to 2010), so it cannot directly speak to transition paths or time horizons for emission convergence.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-welfare-implications-for-developing-countries-under-different-policies-and-why-do-dietary-policies-dominate"&gt;Q13. What are the welfare implications for developing countries under different policies, and why do dietary policies dominate?&lt;/h3&gt;
&lt;p&gt;Under economic growth (10% TFP shock), global welfare rises 14.9% with a modest increase in Q4/Q1 inequality of 0.4 percentage points, indicating relatively even welfare gains. Under no-beef, global welfare falls 0.6% but inequality worsens by 1.0 pp; under vegetarian, welfare falls 2.8% and inequality worsens by 6.0 pp—developing countries lose more because more of their income is spent on food and the agricultural sector is a larger share of their economy. Under eat-local (food trade restrictions), welfare falls 17.8% and the Q4/Q1 ratio worsens by 4.9 pp, with countries in the bottom GDP quartile facing losses up to 41%. The stark dominance of dietary policies over trade policies reflects two structural features: (a) food trade restrictions reduce the gains from comparative advantage in food production, which are particularly large for food-exporting developing countries; and (b) the welfare cost per unit of GHG reduction is far higher for trade policies because they distort production allocation without addressing the underlying demand-side emissions driver.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Nutrition Transition&lt;/strong&gt;: As defined and used in this paper: the demand-side process by which rising income causes consumers to shift their caloric intake away from staple foods (yams, potatoes, rice, millet) toward food products with higher GHG emissions per calorie (meats, fruits, vegetables, coffee). The transition is captured in the model by non-homothetic income elasticity parameters ε_k that are higher for more emissions-intensive products and is operative even after excluding all meat products.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agricultural Modernization&lt;/strong&gt;: As defined and used in this paper: the supply-side process by which rising wages induce producers to substitute from traditional, labor-intensive agricultural technology (τ=0, no purchased intermediate inputs) toward modern, input-intensive technology (τ=1, fertilizers, machinery, pesticides), which emits more GHG per calorie of output. This operates within each crop and is captured in the model by endogenous technology choice at the plot level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-Homothetic CES Preferences (Nested)&lt;/strong&gt;: A three-tier preference structure in which the expenditure share of a food product k depends on income through a product-specific parameter ε_k that governs how fast the product&amp;rsquo;s preference weight grows with utility. Products with higher ε_k have higher income elasticities; the overall income elasticity of the agricultural sector (0.39 in this paper&amp;rsquo;s calibration) is an expenditure-weighted average of the ε_k values. The nested structure allows the agricultural sector&amp;rsquo;s income elasticity relative to non-agriculture to be determined separately from the income elasticities of individual food products within agriculture.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implicit Marshallian Demand&lt;/strong&gt;: The demand equation derived from non-homothetic CES preferences by substituting out unobservable price indices using a base good, yielding a demand specification that depends on observable expenditure shares and income rather than on prices directly. In this paper&amp;rsquo;s open-economy extension, trade shares further substitute out unobservable variety price indices, making the estimation equation fully price-data-free.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GHG Emission Intensity (per calorie)&lt;/strong&gt;: In this paper: the parameter φ_k (crop-specific) and φ_τ (technology-specific), where φ_kτ = φ_k × φ_τ is the kg CO₂-equivalent emitted per 1,000 kcal of crop k produced under technology τ. This is the key cross-product heterogeneity that, combined with income elasticity heterogeneity, drives the environmental consequences of the nutrition transition. In the data: ranges from below 5 kg CO₂ per 1,000 kcal for wheat and rye to above 35 kg for beef and coffee.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Grid-Cell Production Model&lt;/strong&gt;: A representation of the agricultural supply side in which the Earth&amp;rsquo;s land surface is divided into approximately 1.1 million fields (FAO-GAEZ), each with agro-climatically determined potential yields by crop and technology that are independent of market conditions. Within each field, a continuum of plots is allocated to crops and technologies via Fréchet productivity draws, yielding smooth aggregate supply functions and allowing for realistic specialization patterns and technology gradients across geography.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Back-of-the-Envelope (Demand Mechanism) Benchmark&lt;/strong&gt;: In this paper: a partial-equilibrium counterfactual calculation that takes observed or baseline food demand quantities and simply attributes changes to them from a policy without allowing supply prices, production, or trade flows to adjust. The paper systematically compares model general equilibrium results against this benchmark (column 9 of Table 4) to quantify how much supply-side adjustments matter, finding that the back-of-the-envelope approach overstates the emission impact of economic growth by approximately three times, and overstates the emission reduction from dietary policies by roughly one-third.&lt;/p&gt;</description></item><item><title>Dispersion Over the Business Cycle: Passthrough, Productivity, and Demand</title><link>https://macropaperwarehouse.com/papers/dispersion-over-the-business-cycle-passthrough-productivity-and-demand/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/dispersion-over-the-business-cycle-passthrough-productivity-and-demand/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Carlsson, Clymo, and Joslin use Swedish manufacturing firm-level microdata for 1998–2013 to separately identify and characterize the cyclical behavior of physical productivity (TFPQ) shocks and demand shocks at the firm level, two forces that are observationally equivalent under the standard CES-demand benchmark. The paper&amp;rsquo;s central contribution is threefold: it documents new empirical facts about dispersion cyclicality, estimates a non-constant-elasticity (non-CES) demand curve directly from firm-level price and quantity data, and embeds those estimates into a quantitative heterogeneous-firm model to study the aggregate consequences of each type of dispersion shock.&lt;/p&gt;
&lt;p&gt;The data combine four Swedish register sources: the Företagens Ekonomi (FEK) survey for bookkeeping variables; the Industrins Varuproduktion (IVP) survey for 8-digit product-level price and quantity data used to construct firm-level price indices; the Konjunkturstatistik för Industrin (KFI) survey for quarterly capacity-utilization data; and additional investment deflators. The unbalanced panel contains 3,181 unique manufacturing firms and 15,044 firm-year observations. TFPQ is measured using a Cobb-Douglas value-added production function with factor utilization adjustment; factor elasticities are estimated via cost shares at the 2-digit sector level, yielding an average labor share of 0.735.&lt;/p&gt;
&lt;p&gt;Demand is estimated using the Gopinath-Itskhoki-Rigobon (GIR) flexible demand curve, which nests CES as the limiting case. TFPQ innovations instrument for price in a second-order approximation, following Foster, Haltiwanger, and Syverson (2008). The main-sample estimates yield theta = 2.94 (average elasticity) and eta = 4.27 (super-elasticity), both significant at the 1% level. The second-order price term is statistically significant at the 5% level in all three samples, decisively rejecting CES. These estimates imply that a 5% price increase raises the demand elasticity from 2.94 to 3.74, while a 5% price reduction reduces it to 2.42, creating a &amp;ldquo;real rigidity&amp;rdquo; in the sense of Ball and Romer (1990): raising price loses many customers while lowering it gains few.&lt;/p&gt;
&lt;p&gt;Incomplete passthrough of TFPQ shocks is a central empirical finding. OLS estimates yield beta_z = -0.124; first-difference estimates yield -0.097. Even in the subsample of firms that adjusted all product-level prices in a given year, TFPQ passthrough remains near -0.10, ruling out Calvo or menu-cost price stickiness as the sole driver. Longer-horizon (two- and three-year) first-difference regressions produce similar estimates, ruling out Rotemberg gradual adjustment as well. The non-CES demand curve alone implies a static-optimal passthrough of theta/(theta + eta) = 3/(3 + 4.3) = 41%, so real rigidity explains most of the incompleteness even before accounting for adjustment costs. Demand shocks pass through to prices at a rate of 0.209-0.235, a non-zero result rationalized in the quantitative model by input adjustment costs.&lt;/p&gt;
&lt;p&gt;On cyclicality of dispersion, both TFPQ and demand shock dispersion are countercyclical, but demand dispersion rises by more and is more robust across recession episodes. In 2009 (the Great Recession), the IQR of demand shock growth was 56% above its non-recession average, while the IQR of TFPQ shock growth rose 36%. Sales dispersion rose 58% (IQR) in 2009. A semi-structural variance decomposition shows that demand shocks account for 63% of average sales growth dispersion and approximately 80% of its increase in 2009; TFPQ dispersion contributes only marginally to sales dispersion because the TFPQ variance is shrunk by a factor of roughly 25 on its way to sales growth through the chain of low passthrough and demand elasticity. Demand accounts for about 50% of average price growth dispersion and 40% of its cyclical increase in 2009; TFPQ accounts for about 10% of price dispersion on average.&lt;/p&gt;
&lt;p&gt;The quantitative heterogeneous-firm model extends Bloom (2009) and Bloom et al. (2018) to continuous time with both TFPQ and demand shocks, non-CES demand (theta = 3, eta = 4.3 from the estimates), and non-convex input adjustment costs on a composite scale factor covering both capital and labor. The resale loss kappa = 0.3565 is taken from Bloom et al. (2018). The model is calibrated to match IQRs of 0.2 for TFPQ and demand shock log-changes in the low-uncertainty state, consistent with pre-crisis Swedish data. For the high-uncertainty state, the calibration targets the Great Recession peaks: a 30% rise in TFPQ dispersion (sigma_z(2) = 1.38 sigma_z(1)) and a 60% rise in demand dispersion (sigma_epsilon(2) = 1.90 sigma_epsilon(1)), reflecting the empirical finding that demand dispersion increases more.&lt;/p&gt;
&lt;p&gt;A simulated transition to the high-uncertainty state causes aggregate output to fall by 3.5%. Decomposing into the Bloom (2009) &amp;ldquo;volatility effect&amp;rdquo; (realized shocks drawn from the high-dispersion distribution, firms believe low) and &amp;ldquo;uncertainty effect&amp;rdquo; (firms believe high, shocks drawn from low distribution), the paper finds both effects are negative in the non-CES model, in sharp contrast to Bloom (2009) where the volatility effect is positive (the Oi-Hartman-Abel effect). Non-CES demand amplifies the total output decline by approximately 40% relative to the CES model (peak fall 2.5% vs. 1.75%), primarily by reversing the sign of the volatility effect. Increased demand dispersion drives almost all of the first-year output decline and the majority of the uncertainty effect; TFPQ dispersion is the main driver of the negative volatility effect via markup dispersion. The inaction rate among firms jumps from 50% to 95% on impact of the uncertainty shock, then recovers within one year. TFPQ uncertainty induces little wait-and-see behavior because firms optimally adjust inputs by only 23% of the TFPQ shock size (versus 200% under CES), so uncertainty about TFPQ translates mainly into markup uncertainty. Demand uncertainty triggers strong wait-and-see behavior because demand directly maps one-for-one into desired input use.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-core-identification-strategy-for-separating-tfpq-and-demand-shocks-and-what-are-the-main-threats"&gt;Q1. What is the paper&amp;rsquo;s core identification strategy for separating TFPQ and demand shocks, and what are the main threats?&lt;/h3&gt;
&lt;p&gt;The authors identify TFPQ from a utilization-adjusted Cobb-Douglas value-added production function, then estimate demand using TFPQ innovations as instruments for price. TFPQ innovations are valid instruments because they shift marginal cost without directly shifting demand, tracing out the demand curve. The utilization adjustment (from the KFI managerial survey) is critical: without it, demand shocks that reduce utilization would appear as negative TFPQ shocks, biasing demand elasticity estimates upward and breaking instrument validity. The paper validates the adjustment by showing that firms reporting &amp;lsquo;insufficient demand&amp;rsquo; exhibit 15% lower utilization on average, and 23% lower during the Great Recession. A second threat is quality change in firm-level prices; the authors address this with (a) robustness using the Eslava et al. (2023) CUPI quality-adjusted price index and (b) a single-product-firm subsample. Demand and passthrough results are similar across all three price index approaches. The within-firm focus (demeaning by firm and sector-year fixed effects throughout) mitigates cross-sectional comparability issues but limits misallocation-level analyses analogous to Hsieh and Klenow (2009).&lt;/p&gt;
&lt;h3 id="q2-how-is-the-non-ces-demand-curve-identified-and-what-exactly-does-the-super-elasticity-parameter-eta-measure"&gt;Q2. How is the non-CES demand curve identified, and what exactly does the super-elasticity parameter eta measure?&lt;/h3&gt;
&lt;p&gt;The GIR demand curve is q = (1 - eta * log p)^(theta/eta). A second-order approximation around the firm&amp;rsquo;s average price yields log q = -theta * p_hat - (eta&lt;em&gt;theta/2) * p_hat^2 + fixed effects + epsilon, where p_hat is the firm&amp;rsquo;s demeaned log relative price. Regressing real sales on p_hat and p_hat^2, instrumented by demeaned TFPQ and its square, recovers theta = -b1 and eta = 2&lt;/em&gt;b2/b1. Because p_hat is demeaned at the firm level, the estimates capture within-firm nonlinearity in the price-sales relationship, not cross-sectional heterogeneity in elasticity levels. The parameter eta is the &amp;lsquo;super-elasticity&amp;rsquo;: it measures how much the demand elasticity itself changes with the price. When eta &amp;gt; 0, a firm that raises its price faces an increasingly elastic demand curve (loses customers rapidly), and one that lowers its price faces a less elastic curve (gains customers slowly). The estimated eta = 4.27 in the main sample is roughly half the value of 10 studied (but not estimated) in Klenow and Willis (2016) and larger than the approximately 2 used in Berger and Vavra (2019).&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-distinguish-the-volatility-effect-from-the-uncertainty-effect-in-the-quantitative-model"&gt;Q3. How does the paper distinguish the &amp;lsquo;volatility effect&amp;rsquo; from the &amp;lsquo;uncertainty effect&amp;rsquo; in the quantitative model?&lt;/h3&gt;
&lt;p&gt;Following Bloom (2009), the paper simulates two counterfactuals. The uncertainty effect holds shocks drawn from the low-dispersion distribution (s=1) but lets firms believe that the high-uncertainty state (s=2) has arrived; this isolates the precautionary wait-and-see channel. The volatility effect draws shocks from the high-dispersion distribution (s=2) but lets firms believe they are in the low-uncertainty state; this isolates the direct effect of realizing more extreme shocks on aggregate output. In the non-CES model, both effects are negative. The uncertainty effect is dominated by demand uncertainty because demand shocks directly affect desired input use one-for-one, so uncertainty about future demand creates strong incentives to pause investment. TFPQ uncertainty induces little wait-and-see behavior because the optimal scale adjustment to a TFPQ shock is only 23% of the shock magnitude (vs. 200% under CES). The volatility effect is dominated by TFPQ dispersion because realized TFPQ shocks generate markup dispersion via incomplete passthrough, creating misallocation. Under CES, the volatility effect from TFPQ is positive (OHA effect: convex output-productivity relationship); non-CES demand makes the output-productivity relationship concave for eta large enough, flipping the sign.&lt;/p&gt;
&lt;h3 id="q4-what-mechanism-makes-tfpq-passthrough-so-low-in-both-the-data-and-the-model"&gt;Q4. What mechanism makes TFPQ passthrough so low in both the data and the model?&lt;/h3&gt;
&lt;p&gt;Two mechanisms operate. First, non-CES demand itself: when eta &amp;gt; 0, raising price increases the demand elasticity, and lowering price decreases it. This means the benefit to revenue from a price cut (following a productivity gain that reduces costs) is muted because the firm gains fewer customers than under CES. The static optimal passthrough is theta/(theta + eta) = 3/(7.3) = 41%. Second, non-convex input adjustment costs further reduce passthrough by making firms reluctant to change their scale in response to TFPQ shocks. In the model, the investment threshold is nearly flat across a wide range of TFPQ values (shown in Figure 6, left panel), reflecting that optimal scale barely responds to productivity. Together these mechanisms reproduce TFPQ passthrough of 20-30% in model-simulated data vs. 10-24% in the actual data, both far below the CES benchmark of 100%. The paper also verifies that low passthrough persists in the subsample of flexible-price firm-years, ruling out sticky prices as the primary driver.&lt;/p&gt;
&lt;h3 id="q5-why-does-demand-shock-dispersion-rather-than-tfpq-dispersion-dominate-the-variance-decompositions-of-sales-and-price-growth"&gt;Q5. Why does demand shock dispersion, rather than TFPQ dispersion, dominate the variance decompositions of sales and price growth?&lt;/h3&gt;
&lt;p&gt;The contribution of TFPQ dispersion to sales dispersion is (1-theta)^2 * beta_z^2 * Var(z). With beta_z = -0.097 and theta = 2.99, the TFPQ variance is shrunk by approximately (1-2.99)^2 * (0.097)^2 = 4 * 0.0094 ≈ 0.04, so only about 4% of TFPQ variance propagates to sales variance. This extremely small multiplier reflects two successive attenuation steps: low TFPQ passthrough to prices (beta_z^2 ≈ 0.01) and a small price-to-sales elasticity. Demand shocks, by contrast, affect sales directly through the demand curve without a price intermediary: the contribution is ((1-theta)*beta_epsilon + 1)^2 * Var(epsilon). With beta_epsilon = 0.209 and theta = 2.99, the multiplier is ((1-2.99)*0.209 + 1)^2 = (1 - 0.416)^2 = 0.34, about eight times larger than for TFPQ even though both shocks have similar variance. The cyclical increase is even more skewed toward demand because demand dispersion rises by 56% vs. 36% for TFPQ in 2009.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-relate-to-tfpr-dispersion-and-what-does-it-say-about-using-tfpr-as-a-sufficient-statistic"&gt;Q6. How does the paper relate to TFPR dispersion, and what does it say about using TFPR as a sufficient statistic?&lt;/h3&gt;
&lt;p&gt;TFPR = p * z. For arbitrary passthrough, TFPR growth = beta_epsilon * delta_epsilon + (beta_z + 1) * delta_z. Because passthrough from both shocks is incomplete, TFPR growth reflects a mixture of both underlying shocks. The paper shows via a variance decomposition of TFPR that TFPQ is the main driver of TFPR growth dispersion—accounting for roughly 60% on average—because low passthrough means prices move little, leaving TFPQ changes to dominate TFPR. However, this finding obscures the importance of demand shocks for aggregate outcomes: demand dispersion is the dominant driver of sales growth dispersion and wait-and-see behavior, yet TFPR growth dispersion mostly reflects TFPQ. A researcher relying on TFPR dispersion to infer uncertainty would correctly detect productivity uncertainty but would miss the more cyclically important demand uncertainty channel.&lt;/p&gt;
&lt;h3 id="q7-how-do-the-oi-hartman-abel-oha-and-wait-and-see-mechanisms-work-differently-under-non-ces-vs-ces-demand"&gt;Q7. How do the Oi-Hartman-Abel (OHA) and wait-and-see mechanisms work differently under non-CES vs. CES demand?&lt;/h3&gt;
&lt;p&gt;Under CES demand, sales of each firm are s = z^(theta-1) * exp(epsilon), and aggregate output is E[z^(theta-1)] which is convex in z, so a mean-preserving spread in TFPQ raises aggregate output (OHA effect). Under the estimated non-CES parameters (theta=3, eta=4.3), the approximate relationship yields output proportional to z^0.82, which is concave, so a mean-preserving spread in TFPQ reduces aggregate output. The mechanism is that under non-CES demand, TFPQ shocks pass through incompletely to prices and thus create markup dispersion: high-productivity firms have high markups, low-productivity firms have low markups, and the resulting misallocation reduces total output even relative to a social planner who would set p=mc. For wait-and-see: under CES, optimal input adjustment to a TFPQ shock equals (theta-1) times the shock, which is 200% for theta=3; under non-CES with eta=4.3, it is only (theta^2/(theta+eta) - 1) * shock = 0.233 * shock = 23%. This means firms adjust scale very little in response to TFPQ uncertainty, dampening the wait-and-see channel for TFPQ. TFPQ uncertainty then causes uncertainty about markups, which is costly but does not trigger large investment adjustments.&lt;/p&gt;
&lt;h3 id="q8-what-role-do-adjustment-costs-play-and-how-robust-are-the-results-to-the-structure-of-those-costs"&gt;Q8. What role do adjustment costs play, and how robust are the results to the structure of those costs?&lt;/h3&gt;
&lt;p&gt;Non-convex adjustment costs on a composite firm-scale factor x = k^alpha * l^(1-alpha) create an inaction region: firms neither invest nor disinvest until shocks are sufficiently large. In the low-uncertainty state, the model generates a yearly inaction rate of 25.4% (consistent with pre-crisis Swedish data showing roughly 15%). When uncertainty rises, the inaction region widens, the inaction rate jumps to 95% on impact, and firms let their scale shrink via depreciation. The baseline calibration uses the resale loss kappa = 0.3565 from Bloom et al. (2018). The paper also calibrates kappa to the Swedish inaction rate (kappa = 0.1165), which delivers qualitatively identical dynamics but a smaller amplitude recession (1.7pp vs. 3.5pp output fall). The paper also solves a version with adjustment costs only on capital (as in Bachmann and Bayer, 2013): the wait-and-see effect is dampened but the qualitative results hold—demand uncertainty still dominates TFPQ uncertainty in driving wait-and-see, and non-CES demand still reverses the sign of the OHA effect.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-role-of-the-price-wedge-and-time-varying-passthrough"&gt;Q9. What is the role of the price wedge and time-varying passthrough?&lt;/h3&gt;
&lt;p&gt;The passthrough equation residual (price wedge, tau) captures price changes unexplained by TFPQ and demand shocks. It could reflect un-modeled shocks (e.g., financial constraints, as Gilchrist et al. (2017) document for Sweden), markup decisions, or measurement error. The price wedge makes a meaningful contribution to both average sales/price dispersion and to the rise in 2009. Time-varying passthrough is also documented: TFPQ passthrough is countercyclical (more negative in recessions), while demand passthrough is procyclical (falls in recessions when firms receive more extreme idiosyncratic demand shocks). Redoing the variance decomposition with year-by-year passthrough estimates makes demand&amp;rsquo;s contribution to sales dispersion in 2009 even larger, because firms adjust prices less to demand shocks during the recession, leaving more of the demand shock impact in sales.&lt;/p&gt;
&lt;h3 id="q10-what-heterogeneity-is-documented-across-industries-and-firm-types"&gt;Q10. What heterogeneity is documented across industries and firm types?&lt;/h3&gt;
&lt;p&gt;Sectoral demand elasticity estimates from the pooled 22-sector sample yield an average theta of 3.89 and median of 2.73 for the linear CES model; for the non-linear model, average theta is 3.26 and average eta is 7.42, with substantial positive skew. The median non-linear eta of 5.37 is larger than the pooled estimate of 4.27, indicating the pooled estimate is pulled down by some sectors with smaller deviations from CES. Key empirical results (greater cyclicality of demand dispersion, incomplete TFPQ passthrough) hold within each major sector and across balanced panels, the single-product subsample, and the CUPI price-index sample. Time-varying passthrough is also found to be systematically higher by about 25% in the post-2008 period compared to the pre-2008 period, suggesting a structural shift in how demand shocks transmit to prices, though the paper does not investigate the source of this change.&lt;/p&gt;
&lt;h3 id="q11-what-robustness-checks-are-run-on-the-demand-and-passthrough-estimates"&gt;Q11. What robustness checks are run on the demand and passthrough estimates?&lt;/h3&gt;
&lt;p&gt;Demand estimation robustness: (1) piece-wise linear specification (elasticity of 2 below average price, 4 above average price, significant at 0.1% level); (2) balanced panel; (3) excluding the Great Recession; (4) using Statistics Sweden firm identifiers instead of authors&amp;rsquo; own; (5) CUPI price index; (6) single-product firms; (7) sector-by-sector estimation; (8) including firm and sector-year fixed effects directly in the nonlinear regression (rather than pre-demeaning). All exercises confirm statistically significant eta and broadly similar theta. Passthrough robustness: (1) OLS vs. IV (lagged shocks) vs. first-differences; (2) balanced panel; (3) single-product subsample; (4) two-period lagged instruments (beta_z = -0.294, beta_epsilon = 0.249); (5) flexible-price subsample; (6) longer-horizon (two- and three-year) first differences for TFPQ. Corroboration: TFPQ innovations are positively associated with reported process innovations in Eurostat CIS data (7% greater TFPQ growth for process innovators); negative demand shocks are correlated with managers reporting &amp;lsquo;insufficient demand&amp;rsquo; in KFI data (8% lower demand growth).&lt;/p&gt;
&lt;h3 id="q12-how-does-this-paper-differ-from-and-relate-to-bloom-2009-and-bloom-et-al-2018"&gt;Q12. How does this paper differ from and relate to Bloom (2009) and Bloom et al. (2018)?&lt;/h3&gt;
&lt;p&gt;Bloom (2009) and Bloom et al. (2018) model a single composite firm-level shock (implicitly TFPR) in a CES-demand economy, finding that uncertainty shocks reduce output through wait-and-see behavior but generate a positive volatility effect (OHA) that partly offsets the uncertainty effect. The present paper adds two departures: (1) it separates TFPQ and demand shocks and shows they have distinct empirical and aggregate implications; (2) it replaces CES demand with an estimated non-CES demand curve. Departure (2) reverses the OHA effect, amplifying the total output decline by around 40% relative to the CES model. Departure (1) shows that the uncertainty channel operates primarily through demand, while TFPQ operates primarily through the volatility channel. The quantitative model uses the same non-convex adjustment cost structure and calibration approach as Bloom et al. (2018) to ensure comparability. The paper also relates to Bachmann and Bayer (2013) and Mongey and Williams (2017), who find smaller aggregate effects with adjustment costs only on capital; the present paper notes that adjustment costs on both capital and labor are needed for large wait-and-see effects, but qualitative conclusions are unchanged with capital-only costs.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-policy-and-theoretical-implications-of-the-findings"&gt;Q13. What are the policy and theoretical implications of the findings?&lt;/h3&gt;
&lt;p&gt;First, policies aimed at reducing firm-level demand uncertainty (e.g., demand stabilization, aggregate demand management) have larger aggregate output effects than policies addressing productivity uncertainty, because demand uncertainty triggers wait-and-see investment behavior while TFPQ uncertainty is largely absorbed in markups without changing investment much. Second, TFPQ dispersion is still harmful but through misallocation: policies that reduce markup dispersion induced by productivity differentials can raise aggregate output without requiring reduced dispersion per se. Third, the finding that TFPR dispersion is a poor proxy for demand shock dispersion has implications for how researchers use TFPR as a measure of misallocation or uncertainty: it conflates two distinct forces with different aggregate implications. Fourth, the estimated super-elasticity provides a data-disciplined input for calibrating models with real rigidities, directly relevant for the Ball-Romer nominal non-neutrality question—higher real rigidities amplify the output effects of monetary policy shocks. The authors flag this as a natural extension. The scope conditions are: Swedish manufacturing, annual data 1998-2013, partial equilibrium model (aggregate price level exogenous), firms with matching price and utilization data (large-firm bias).&lt;/p&gt;
&lt;h3 id="q14-what-additional-findings-are-documented-regarding-the-cyclicality-of-other-firm-level-variables"&gt;Q14. What additional findings are documented regarding the cyclicality of other firm-level variables?&lt;/h3&gt;
&lt;p&gt;Beyond TFPQ and demand dispersion, the paper documents that dispersion of sales growth, price growth, labor, intermediate goods, and capacity utilization are all countercyclical. The IQR of sales growth was 58% above the non-recession average in 2009 and 9% above in 2001; the IQR of price growth was 83% above in 2009 and 5% above in 2001. The one notable exception is investment, which displays procyclical dispersion (less dispersed during the Great Recession). The paper also documents that roughly 30% of firms report insufficient demand at all their plants in the survey data; average capacity utilization is 88% with median 91% and standard deviation of 14.1%; and about 25% of firm-year observations involve utilization at or above 100%.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Physical total factor productivity (TFPQ)&lt;/strong&gt;: Firm-level quantity productivity: output per unit of inputs, measured from a utilization-adjusted Cobb-Douglas value-added production function. Distinct from revenue TFP (TFPR = p*z) because it abstracts from demand conditions and price-setting. In this paper, TFPQ is estimated within firm over time using the cost-share approach and a capacity-utilization correction from managerial survey data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Demand shock (epsilon)&lt;/strong&gt;: The idiosyncratic component of a firm&amp;rsquo;s demand curve that captures its ability to sell more (or fewer) units at a given price in a given year, reflecting changes in customer base size or customers&amp;rsquo; willingness to pay. Estimated as the residual from the GIR demand curve after controlling for firm fixed effects, sector-time fixed effects, and the firm&amp;rsquo;s own price.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-CES demand curve / super-elasticity (eta)&lt;/strong&gt;: A demand specification adapted from Gopinath, Itskhoki, and Rigobon (2010) in which the demand elasticity is not constant but rises with the firm&amp;rsquo;s price. The parameter eta (estimated at 4.27 in the main sample) governs how fast the elasticity rises with the price: when eta &amp;gt; 0, firms gain few customers by cutting price (elasticity falls as price falls) and lose many customers by raising price (elasticity rises as price rises). This is the source of &amp;lsquo;real rigidity&amp;rsquo; that makes incomplete TFPQ passthrough optimal.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Incomplete TFPQ passthrough&lt;/strong&gt;: The empirical finding that firms reduce their prices by far less than one-for-one in response to a productivity gain (estimated beta_z = -0.097 to -0.124, far from the CES benchmark of -1). The paper attributes this primarily to non-CES demand real rigidity (which implies an optimal static passthrough of only 41% given the estimated parameters) and secondarily to adjustment costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Oi-Hartman-Abel (OHA) effect&lt;/strong&gt;: The positive &amp;lsquo;volatility effect&amp;rsquo; in standard CES-demand uncertainty models: because output is a convex function of TFPQ under CES, a mean-preserving spread in productivity raises aggregate output (lucky firms expand more than unlucky firms contract). The paper overturns this result by showing that with non-CES demand (eta sufficiently large), the output-productivity relationship becomes concave, so TFPQ dispersion reduces aggregate output via markup misallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wait-and-see channel&lt;/strong&gt;: The mechanism by which uncertainty about future shocks causes firms with non-convex input adjustment costs to pause investment: firms prefer to remain inactive and let inputs depreciate rather than invest or disinvest, at the risk of having to pay an irreversibility cost if the shock turns out to have been in the opposite direction. In this paper, this channel is driven primarily by demand uncertainty because demand shocks determine how many units a firm can sell and hence its desired input level; TFPQ uncertainty does not trigger strong wait-and-see behavior because the optimal scale response to TFPQ shocks is small under non-CES demand.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Markup dispersion / misallocation&lt;/strong&gt;: Dispersion across firms in the ratio of price to marginal cost, arising in this paper from incomplete TFPQ passthrough: firms with high productivity set high markups rather than passing through productivity gains as price cuts. The resulting wedge between prices and marginal costs means that resources are misallocated (too little output at high-productivity firms relative to the social optimum), reducing aggregate output. This is the channel through which TFPQ dispersion harms the aggregate economy in the model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Price wedge (tau)&lt;/strong&gt;: The residual from the passthrough regression: the component of firm price changes unexplained by the estimated TFPQ and demand shocks. Interpreted as capturing un-modeled shocks (financial constraints, markup adjustments) and potentially measurement error. The price wedge makes a meaningful contribution to both average sales/price dispersion and to the Great Recession increase in dispersion.&lt;/p&gt;</description></item><item><title>Distributional Consequences of Becoming Climate-Neutral</title><link>https://macropaperwarehouse.com/papers/distributional-consequences-of-becoming-climate-neutral/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/distributional-consequences-of-becoming-climate-neutral/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates how the EU&amp;rsquo;s Fit-for-55 climate package will affect aggregate output and distribute its costs across the income distribution. The question matters because energy is a necessity good — poorer households devote a larger share of spending to energy — so policies that raise energy prices are regressive in their first-order incidence. Despite a large literature on the aggregate macroeconomics of the green transition, distributional consequences have received limited attention.&lt;/p&gt;
&lt;p&gt;The authors build a parsimonious dynamic general-equilibrium model with two infinitely-lived households (rich and poor), a standard output-producing firm that treats energy as a complementary CES input alongside the capital-labor aggregate, and an energy-producing sector that combines a carbon-intensive brown technology with a carbon-free green technology as imperfect substitutes (CES with elasticity of substitution calibrated to 3 following Papageorgiou et al. 2017). The novel feature is Price Independent Generalized Linearity (PIGL) non-homothetic preferences following Boppart (2014), which generate nonlinear Engel curves: the poor agent&amp;rsquo;s energy expenditure share exceeds the rich agent&amp;rsquo;s, matching Eurostat Household Finance and Consumption Survey data (2015) showing the bottom income quintile has more than twice the energy expenditure share of the top quintile. The model targets an 18% energy expenditure share for the poor agent and 7.5% for the rich agent. The rich agent holds all financial wealth; the poor agent lives on labor income alone. The government taxes the brown technology and recycles revenue as a green-technology subsidy under a balanced budget, representing the ETS. Agents have perfect foresight. The paper simulates perfect-foresight transitions from an initial steady state to a new climate-neutral steady state, with the transition path endogenously determining the new steady state — a nonstandard feature arising from non-homothetic preferences.&lt;/p&gt;
&lt;p&gt;In the baseline scenario (linear tax ramp over 25 years), achieving an 85% reduction in brown energy use requires a 168% tax on the brown technology. This drives the price of energy services up by 49%, GDP down by 9.3% in the new steady state, energy as a production input down by 10.9%, and capital input down by 9.3%, while the real wage falls by roughly 7% and the real interest rate is nearly unchanged (dropping by only 0.02 percentage points transiently). The welfare cost measured in expenditure-equivalent terms is a 10.8% loss for the rich agent and a 16.2% loss for the poor agent — the poor agent suffers approximately 50% more. To finance consumption during the transition the poor agent accumulates debt equal to 38.8% of annual income.&lt;/p&gt;
&lt;p&gt;Results are highly sensitive to the brown-green substitution elasticity: raising it from 3 to 5 roughly halves the required tax (to 78.6%) and halves GDP losses (to 4.7%); lowering it to 2 roughly doubles the tax (to 354%) and GDP losses (to 17.7%). Non-homothetic preferences matter quantitatively: switching to homothetic preferences (while preserving different expenditure shares) shrinks aggregate GDP losses by 26% and eliminates nearly all distributional disparity, confirming that the non-homotheticity — not merely different expenditure levels — is the operative distributional mechanism. If the Fit-for-55 energy efficiency improvement target of 1.49% per year is simultaneously achieved, the required tax falls to 136%, the price of energy actually declines by 5.5%, and GDP rises by 1.1% in the new steady state, with the poor agent benefiting slightly more and accumulating assets (4% of annual income) rather than debt.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-modeling-and-calibration-strategy-and-what-are-the-main-threats"&gt;Q1. What is the core modeling and calibration strategy, and what are the main threats?&lt;/h3&gt;
&lt;p&gt;The paper is a quantitative theory exercise with no econometric identification. Calibration targets HFCS Eurostat data (2015) for energy expenditure shares by income quintile, the Papageorgiou et al. (2017) estimate of the brown-green substitution elasticity (ρE = 3), and stylized facts on wealth and income distribution from Krueger, Mitman, and Perri (2016). The main threat is parameter uncertainty around ρE, which the paper acknowledges is poorly identified empirically and which drives the results almost one-for-one. The sensitivity analysis explores ρE ∈ {2, 3, 5}, a range the paper concedes is narrow relative to the literature&amp;rsquo;s full dispersion.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-generating-the-distributional-gap-between-rich-and-poor"&gt;Q2. What are the main mechanisms generating the distributional gap between rich and poor?&lt;/h3&gt;
&lt;p&gt;Three reinforcing channels: (1) Non-homothetic preferences give the poor agent a higher energy expenditure share (18% vs. 7.5%), so the 49% energy price increase hits the poor&amp;rsquo;s budget much harder as a share of income. (2) The poor agent cannot buffer the shock through wealth drawdowns (holding zero net assets initially), forcing it to accumulate debt of 38.8% of annual income. (3) Non-homothetic preferences alter the labor supply response: as expenditures fall, the poor agent&amp;rsquo;s labor supply declines less than the rich agent&amp;rsquo;s (the rich agent decreases labor supply by 0.2 percentage points more), reflecting that leisure is a luxury good in this preference system. In the new steady state the rich agent&amp;rsquo;s consumption of the consumption good drops sharply while the rich agent front-loads consumption at the announcement, immediately jumping 2% higher.&lt;/p&gt;
&lt;h3 id="q3-how-are-non-homothetic-preferences-distinguished-empirically-and-in-the-model-from-simply-having-different-expenditure-shares"&gt;Q3. How are non-homothetic preferences distinguished empirically and in the model from simply having different expenditure shares?&lt;/h3&gt;
&lt;p&gt;Section 4.4 runs a counterfactual with homothetic preferences (ε = 0) but preserves identical initial expenditure shares for each agent (7.5% and 18%) by making ν agent-specific. Under homotheticity the expenditure shares do not vary with income as the transition unfolds. The comparison shows that GDP losses shrink by 26% (from 9.3% to 6.9%) and the distributional gap nearly vanishes — both agents experience almost identical welfare losses. This decomposition isolates the effect of non-homotheticity itself: it is the income-dependent adjustment of expenditure shares during the transition, not merely the different initial levels, that drives both larger aggregate losses and the distributional disparity.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented-and-along-what-dimensions"&gt;Q4. What heterogeneity is documented and along what dimensions?&lt;/h3&gt;
&lt;p&gt;Heterogeneity is modeled along two dimensions: initial wealth (rich holds all assets; poor holds zero) and energy expenditure shares (18% for poor, 7.5% for rich) arising from non-homothetic preferences. The model produces no within-group heterogeneity by construction (two-agent framework). The paper documents the time paths of consumption, expenditures, expenditure equivalents, energy expenditure shares, and wealth shares for each agent separately along the transition, showing that both agents cut energy consumption by roughly 15% while the poor agent cuts consumption-good spending by substantially more than the rich agent.&lt;/p&gt;
&lt;h3 id="q5-what-alternative-transition-timing-paths-are-explored-and-what-do-they-imply"&gt;Q5. What alternative transition timing paths are explored and what do they imply?&lt;/h3&gt;
&lt;p&gt;Three alternatives supplement the linear baseline: tax introduction after 1 year, after 12.5 years, and after 25 years of the announcement. Key findings: (a) the required final tax rate is nearly insensitive to timing — the 25-year-delayed scenario requires 172% vs. 168% in the baseline; (b) conditional on excluding climate damages, it is always welfare-superior to delay implementation, with the poor agent gaining close to 3.5 percentage points in expenditure equivalent welfare by delaying to 25 years vs. implementing after 1 year; (c) gradual vs. immediate introduction yields similar welfare outcomes in the benchmark without adjustment costs, but with investment adjustment costs (χ = 10) a sudden implementation causes a brief sharp drop in the real interest rate without large quantity effects.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-gdp-measure-differ-from-aggregate-output-in-the-model"&gt;Q6. How does the GDP measure differ from aggregate output in the model?&lt;/h3&gt;
&lt;p&gt;GDP is defined to exclude the share of final output used as input into energy production. Aggregate output Y falls 7.3% in the new steady state, but GDP falls 9.3%. The gap (approximately 2 percentage points) reflects the increased resource cost of energy production under the green transition: because the brown and green technologies are imperfect substitutes, satisfying the emission reduction target requires devoting a larger share of final output to producing energy services, a real resource drain captured in the GDP definition but excluded from raw output Y.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-energy-efficiency-scenario-imply-and-what-is-its-key-caveat"&gt;Q7. What does the energy efficiency scenario imply, and what is its key caveat?&lt;/h3&gt;
&lt;p&gt;If energy efficiency improves at 1.49% per year over 25 years (a 45% cumulative gain in energy-producing-firm total factor productivity), the required tax falls to 136.3%, the price of energy declines by 5.5% (rather than rising 49%), and GDP rises 1.1% rather than falling 9.3%. The poor agent benefits more from the efficiency gains and accumulates assets worth 4% of annual income rather than debt. The critical caveat is that the efficiency improvement is modeled as purely exogenous and costless. The paper explicitly acknowledges that achieving these efficiency gains may require investment that is not modeled, so the results should be interpreted as an upper bound on the offsetting potential.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-relate-to-and-differ-from-the-most-closely-related-prior-work"&gt;Q8. How does the paper relate to and differ from the most closely related prior work?&lt;/h3&gt;
&lt;p&gt;Ascari et al. (2025) is the closest related paper (developed independently). Differences: (i) Ascari et al. use a Bewley-type incomplete-markets model generating heterogeneity through random discount factors, whereas this paper uses a two-agent complete-markets construct with exogenously fixed initial wealth; (ii) this paper allows endogenous labor supply, which increases short-run flexibility; (iii) this paper does not consider transfer schemes to redistribute away from distributional consequences. Results are described as broadly consistent. Fried, Novan, and Peterman (2018) and Boehl and Budianto (2024) use OLG models and find inequality implications but focus on inter-generational rather than intra-generational distributional effects.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The core implications are: (1) the Fit-for-55 emission tax alone is regressive — the poor bear a welfare loss 50% larger than the rich and end up with 38.8% of annual income in additional debt; (2) delaying tax implementation (with early announcement) is welfare-improving in the absence of climate damage modeling — the welfare difference is nearly 3.5 percentage points for the poor between fastest and latest implementation; (3) if energy efficiency targets are met exogenously, the transition is nearly costless and distributional concerns vanish; (4) the regressive result is conditional on the government recycling tax revenues to green-technology subsidies rather than to household transfers. All these implications are conditional on European economies where climate damages are plausibly small and the model abstracts from open-economy dynamics, endogenous technology, and within-income-group heterogeneity.&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-are-reported"&gt;Q10. What robustness checks are reported?&lt;/h3&gt;
&lt;p&gt;Five robustness exercises are reported: (1) investment adjustment costs raised from χ = 0 to χ = 10 — minimal effect on welfare or quantities in the smooth baseline, though sudden tax introduction produces a brief interest-rate plunge; (2) homothetic preferences counterfactual while maintaining initial expenditure shares (Section 4.4); (3) elasticity of substitution between brown and green technology at ρE = 2 and ρE = 5 (Section 4.3, Table 2); (4) alternative transition timing (1 year, 12.5 years, 25 years post-announcement; Section 4.2); (5) simultaneous energy efficiency improvement of 1.49% per year (Section 4.5). A New Keynesian extension with Rotemberg price adjustment costs and a Taylor rule (Appendix B) is also provided for robustness on inflation dynamics.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-main-caveats-or-limitations-acknowledged-by-the-authors"&gt;Q11. What are the main caveats or limitations acknowledged by the authors?&lt;/h3&gt;
&lt;p&gt;Climate damages are excluded, so the paper understates the case for early action and cannot provide a full welfare comparison between acting early and acting late. Energy efficiency improvement is modeled as exogenous and costless, overstating the net gain from that channel. The two-agent framework abstracts from within-group heterogeneity and overlapping generations. Open-economy dynamics are not modeled; the brown-technology structure serves as a reduced-form for energy imports but does not capture international price feedback. The elasticity of substitution between brown and green technology is uncertain, and results are nearly proportional to this parameter. The model has no endogenous innovation or directed technical change, limiting applicability to long-run transition analysis.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Non-homothetic PIGL preferences&lt;/strong&gt;: Preferences of the Price Independent Generalized Linearity class (Boppart 2014) where energy expenditure shares depend on income level, making energy a necessity good (share declining in income) and consumption goods a luxury. Parameter ε ∈ (0,1) controls non-homotheticity; ε = 0 recovers homothetic preferences. The paper calibrates γ = 0.639 from CEX data, implying an elasticity of substitution between consumption and energy goods of approximately 0.4.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Brown vs. green technology&lt;/strong&gt;: Two imperfectly substitutable technologies for producing energy services within the model&amp;rsquo;s energy sector. The brown technology converts units of final output into energy services using a carbon-intensive (emission-producing) process; the green technology is emission-free. They enter a CES aggregator for energy production with elasticity ρE calibrated to 3. Imperfect substitutability means the green transition raises the cost of energy services even with subsidies to green technology.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Expenditure equivalent loss&lt;/strong&gt;: The welfare metric used in the paper: the percentage change in expenditures in the initial steady state (without any tax) that would make an agent indifferent between remaining in the initial steady state and living through the actual transition path. Defined implicitly by equating flow utility at scaled initial expenditures to flow utility along the transition. Baseline results: -10.8% for the rich agent and -16.2% for the poor agent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tax on the brown technology&lt;/strong&gt;: The policy instrument modeled as capturing the essence of EU ETS and national carbon schemes. It raises the unit cost of the emission-intensive energy input; revenue is recycled as a subsidy to the green technology within a balanced government budget rather than distributed to households. A 168% tax achieves the 85% emission reduction target in the baseline, implying fossil fuel prices nearly triple.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous final steady state&lt;/strong&gt;: The model&amp;rsquo;s new steady state after the green transition is not predetermined; it depends on the wealth distribution that emerges endogenously during the transition. Because markets are complete and preferences are non-homothetic, different transition paths generate different terminal wealth distributions and therefore different aggregate outcomes in the new steady state. This prevents backward solution and requires a fully nonlinear transition path solver.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Energy expenditure share by income quintile&lt;/strong&gt;: The empirical regularity, documented from Eurostat HFCS data (2015), that the bottom income quintile devotes more than twice the fraction of disposable income to energy (electricity, gas, fuels for personal transport) as the top quintile. This fact calibrates the non-homotheticity of preferences (targeting 18% for the poor agent and 7.5% for the rich agent) and motivates the paper&amp;rsquo;s focus on distributional consequences.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Elasticity of substitution between brown and green technology (ρE)&lt;/strong&gt;: The key production-side parameter governing how easily the energy sector can switch from fossil-fuel to clean inputs. Calibrated to ρE = 3 from Papageorgiou et al. (2017). Results are nearly proportional to this parameter: ρE = 5 halves and ρE = 2 roughly doubles the required tax, GDP losses, and welfare costs. The paper identifies this as the dominant source of quantitative uncertainty.&lt;/p&gt;</description></item><item><title>Forecasting with Feedback</title><link>https://macropaperwarehouse.com/papers/forecasting-with-feedback/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/forecasting-with-feedback/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper develops a strategic model of point forecast production in environments where the forecast itself influences the outcome being predicted — what the authors call &amp;ldquo;forecasting with feedback.&amp;rdquo; The canonical example is Federal Reserve staff (Greenbook) inflation forecasts: these forecasts guide FOMC interest rate decisions, and those rate decisions in turn affect realized inflation. The central theoretical claim, proved formally, is that even a forecaster with purely quadratic (mean-squared-error) loss will optimally produce biased forecasts in such environments, provided there is some uncertainty about how strongly the decision maker (DM) will react to the forecast. This finding offers a third interpretation of observed forecast biases — beyond the two dominant explanations in the prior literature, namely forecaster irrationality and asymmetric loss functions.&lt;/p&gt;
&lt;p&gt;The model has three components. First, an outcome equation: y_{t+1} = theta_t + a_t + epsilon_{t+1}, where theta_t is a private signal (the state of the economy) observed only by the forecaster, a_t is the DM&amp;rsquo;s action, and epsilon_{t+1} is unforecastable noise. Second, a DM reaction function: a_t = x_t * [y_T - E(theta_t | f_t)], analogous to a Taylor rule, where y_T is a known target, and x_t is a strength-of-reaction multiplier drawn from a distribution with mean mu and variance tau^2; x_t is the DM&amp;rsquo;s private information. Third, the forecaster minimizes expected squared error, anticipating the DM&amp;rsquo;s endogenous response. The model is linear and closed-form solutions are derived.&lt;/p&gt;
&lt;p&gt;The key mechanism is a bias-variance tradeoff. Because the DM&amp;rsquo;s action responds to the forecast, the variance of the realized outcome itself becomes a function of the forecast. When the DM&amp;rsquo;s reaction strength x_t is uncertain (tau^2 &amp;gt; 0), this variance-of-outcome term is not trivially minimized by an unbiased forecast. The forecaster reduces outcome volatility by attenuating the sensitivity of the forecast to the state — shrinking the forecast slope toward zero relative to what an unbiased forecast would require — at the cost of introducing systematic bias. When tau^2 = 0 (no uncertainty about the DM&amp;rsquo;s reaction), the forecaster can perfectly anticipate and correct for the DM&amp;rsquo;s response, and the optimal forecast is unbiased. Feedback alone, without uncertainty, does not produce bias.&lt;/p&gt;
&lt;p&gt;The paper derives equilibrium forecasts in a Perfect Bayesian Equilibrium where the DM holds correct (rational) beliefs about the forecasting rule. Key analytical results include: (i) the equilibrium exists when tau^2 &amp;lt;= 1/4; (ii) the equilibrium conditional bias equals [(1 - sqrt(1 - 4*tau^2))/2] * (theta_t - y_T), which changes sign depending on whether the state is above or below the target — the forecaster gravitates toward the target; (iii) the Mincer-Zarnowitz (MZ) regression slope (the slope from regressing realized outcomes on forecasts) can be large and positive, close to zero, or even negative, depending on mu and tau^2; (iv) when mu = 1 (the DM on average fully closes the gap to the target), the equilibrium MZ slope is exactly zero for any tau^2 value.&lt;/p&gt;
&lt;p&gt;The paper motivates these results with two documented empirical patterns in Greenbook 4-quarter-ahead inflation forecasts from 1980q1 to 2019q4. First, using 40-quarter rolling windows, bias in Greenbook forecasts is persistent but sign-changing over time — a pattern consistent with the model&amp;rsquo;s prediction that the sign of bias tracks whether the state theta_t is above or below the inflation target y_T. Second, the MZ slope (from 40-quarter rolling-window regressions) hovers near unity in the mid-1980s through early 1990s, returns to unity by the late 1990s, then drops sharply to significantly negative territory by the mid-2000s, before becoming indistinguishable from zero in the final portion of the sample — a pattern consistent with the model&amp;rsquo;s prediction that the MZ slope shifts radically with changes in mu and tau^2. Both facts are computed using the last revision of the GDP deflator.&lt;/p&gt;
&lt;p&gt;The policy and methodological implications are significant. Standard forecast rationality tests (Mincer-Zarnowitz regressions, bias tests) are designed to detect irrationality or asymmetric loss, but in feedback environments these same test statistics can indicate &amp;ldquo;failure&amp;rdquo; even when the forecaster is fully rational under quadratic loss. Studies conducting rationality tests or estimating loss functions must either explicitly assume away feedback (and justify that assumption) or account for the feedback mechanism.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-identification"&gt;Q1. What is the identification strategy, and what are the main threats to identification?&lt;/h3&gt;
&lt;p&gt;The paper is primarily theoretical: it derives closed-form equilibrium forecasting rules and forecast statistics from first principles within a stylized game-theoretic model. There is no econometric identification exercise. The Greenbook evidence is descriptive and motivational — rolling-window bias estimates and MZ slope estimates are presented as stylized facts consistent with the theory, not as causal identification. The main caveat the authors themselves make is that the model is not claimed to be an exclusive or exhaustive explanation of the documented GB forecast patterns. Inflation forecasting is complex, and many other factors (learning, structural breaks, regime changes in monetary policy, data revisions) could contribute to the observed patterns. The authors explicitly disclaim any claim to exclusivity.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-core-mathematical-mechanism-and-how-does-uncertainty-play-a-necessary-role"&gt;Q2. What is the core mathematical mechanism, and how does uncertainty play a necessary role?&lt;/h3&gt;
&lt;p&gt;The forecaster&amp;rsquo;s MSE decomposes into a conditional variance term and a squared-bias term: MSE = Var[a*(f_t) | theta_t] + bias^2(f_t | theta_t) + sigma^2. The critical insight is that when x_t (the reaction-strength multiplier) is uncertain, the variance of the DM&amp;rsquo;s action — and hence of the outcome — depends on the level of the forecast itself. Specifically, Var[a*(f_t) | theta_t] = tau^2 * (y_T - f_t/c + b/c)^2. So choosing a larger or smaller forecast changes not just the bias term but also the variance term. The optimal resolution of this tradeoff requires an attenuated (biased) forecast slope. When tau^2 = 0 (no uncertainty), the variance term vanishes entirely and the forecaster can correct for feedback in full by solving a fixed-point problem, producing an unbiased forecast. The paper explicitly proves (taking limits as tau^2 to 0 in the bias and MZ slope formulas) that both return to zero and one respectively, confirming that uncertainty is a necessary condition for bias.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-equilibrium-concept-and-what-are-its-properties"&gt;Q3. What is the equilibrium concept and what are its properties?&lt;/h3&gt;
&lt;p&gt;The equilibrium is a linear Perfect Bayesian Equilibrium (PBE). The DM conjectures that the forecast is a linear function f_t = b + c*theta_t, uses that conjecture to form expectations E(theta_t | f_t) = (f_t - b)/c, and chooses her action optimally. Equilibrium requires that the DM&amp;rsquo;s conjectured intercept and slope (b, c) coincide with those actually used by the forecaster. The paper shows (Corollary 1) that such a linear PBE exists when tau^2 &amp;lt;= 1/4, and that the equilibrium is fully revealing — the DM can learn the true state theta_t from the forecast because the forecast is a one-to-one function of the state. Two linear equilibria exist: the paper focuses on the Pareto-preferred one (lower forecaster loss, lower absolute bias), which is also the one whose limit as tau^2 approaches 0 corresponds to the natural optimal forecast.&lt;/p&gt;
&lt;h3 id="q4-what-sign-and-magnitude-patterns-does-the-equilibrium-bias-exhibit"&gt;Q4. What sign and magnitude patterns does the equilibrium bias exhibit?&lt;/h3&gt;
&lt;p&gt;From Corollary 2(a), the conditional equilibrium bias is: E(y_{t+1} - f_t^dagger | theta_t) = [(1 - sqrt(1 - 4&lt;em&gt;tau^2)) / 2] * (theta_t - y_T). The multiplier (1 - sqrt(1 - 4&lt;/em&gt;tau^2))/2 is always positive (for tau^2 in (0, 1/4]), so the sign of the bias is determined entirely by the sign of (theta_t - y_T). When theta_t &amp;gt; y_T (state above target), bias is positive — the forecaster underpredicts, shrinking the forecast toward the target. When theta_t &amp;lt; y_T, bias is negative — the forecaster overpredicts, again gravitating toward the target. This sign-change mechanism, driven by changing economic conditions relative to a fixed target, is cited as consistent with the persistent but sign-changing bias observed in Greenbook inflation forecasts from 1980 to 2019.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-model-predict-about-the-mincer-zarnowitz-slope-and-how-variable-can-it-be"&gt;Q5. What does the model predict about the Mincer-Zarnowitz slope, and how variable can it be?&lt;/h3&gt;
&lt;p&gt;From Corollary 2(b), the MZ slope in equilibrium is a highly nonlinear function of mu and tau^2. Figure 3 in the paper (discussed in the text) shows that the slope can be large and positive, positive but close to zero, negative, or even very steeply negative, for different combinations of mu and tau^2. A key special case: when mu = 1 (DM fully closes the gap to target on average), E(y_{t+1} | f_t^dagger) = y_T for all values of the forecast, giving an MZ slope of exactly zero and intercept equal to y_T. The authors note that when mu is close to 1 and tau^2 is small, even small deviations of mu from unity can produce large positive or negative MZ slopes. The model can thus account for the dramatic shift in the GB MZ slope documented in the paper — from around unity in the 1980s-1990s, to significantly negative territory in the mid-2000s, to approximately zero thereafter.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-relationship-between-the-dms-reaction-function-and-the-taylor-rule-and-how-is-it-microfounded"&gt;Q6. What is the relationship between the DM&amp;rsquo;s reaction function and the Taylor rule, and how is it microfounded?&lt;/h3&gt;
&lt;p&gt;The DM&amp;rsquo;s reaction function is a_t* = x_t * [y_T - E(theta_t | f_t)], directly analogous in spirit to a Taylor rule (Taylor, 1993). Online Appendix A provides a formal microfoundation: if the DM minimizes a quadratic loss in (y_{t+1} - y_T)^2 plus a quadratic adjustment cost w_t * a_t^2 — where w_t is a private, randomly drawn adjustment cost parameter — then the optimal action is precisely a_t* = x_t * [y_T - E(theta_t | f_t)] with x_t = 1/(1 + w_t). This microfoundation connects the model to the literature on central bank optimal control and provides a rational justification for the reaction function structure used throughout the paper.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-the-crawford-sobel-1982-cheap-talk-model"&gt;Q7. How does this paper relate to and differ from the Crawford-Sobel (1982) cheap talk model?&lt;/h3&gt;
&lt;p&gt;The paper borrows the sender-receiver communication game structure from Crawford and Sobel (1982), with the forecaster as sender and the DM as receiver. However, it departs in two important ways. First, in Crawford-Sobel, the sender&amp;rsquo;s payoff depends only on the state and the action, not directly on the message (the forecast). In this paper, the forecast enters the forecaster&amp;rsquo;s loss function directly through the outcome equation (y = theta + a + epsilon, and the forecast determines a which determines y which enters the loss), making it a model of &amp;lsquo;costly talk&amp;rsquo; in the sense of Kartik, Ottaviani, and Squintani (2007). Second, in standard communication games the realized outcome is exogenous — the DM&amp;rsquo;s action affects only her own payoff but not the variable being forecast. Here, the DM&amp;rsquo;s action causally determines the realized outcome that the forecaster was trying to predict. This feedback causality is absent in the standard setup and is the source of the paper&amp;rsquo;s novel results.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-bernanke-and-woodford-1997"&gt;Q8. How does this paper relate to Bernanke and Woodford (1997)?&lt;/h3&gt;
&lt;p&gt;Bernanke and Woodford (1997) also study professional inflation forecasts and monetary policy in a rational expectations equilibrium framework, and raise the question of whether an informative equilibrium exists — concluding it may not. This paper differs in three respects: it assumes the forecaster has private information (state theta_t) that the DM cannot directly observe; it works in an environment with uncertainty about the DM&amp;rsquo;s reaction (x_t is random); and rather than focusing on equilibrium existence, it derives the statistical properties of equilibrium forecasts — the bias formula, MZ regression coefficients — which Bernanke and Woodford do not. The authors describe their work as providing &amp;rsquo;the first formal treatment of the statistical properties of forecasts&amp;rsquo; in feedback environments.&lt;/p&gt;
&lt;h3 id="q9-what-heterogeneity-and-parameter-sensitivity-is-documented"&gt;Q9. What heterogeneity and parameter sensitivity is documented?&lt;/h3&gt;
&lt;p&gt;The paper documents sensitivity of forecast properties to mu (mean policy reaction strength) and tau^2 (variance of policy reaction strength). The DM&amp;rsquo;s average aggressiveness mu affects both the sign and magnitude of the MZ slope: for cautious DMs (mu near 0.1), the equilibrium MZ slope is relatively close to unity; for aggressive DMs (mu near 1), the slope can flatten toward zero; for moderate but increasing mu (with tau^2 above a threshold of approximately 0.05), the slope flattens monotonically. A higher tau^2 at given mu generally attenuates the slope toward zero, but the relationship is nonlinear. When mu is precisely one, the MZ slope is exactly zero regardless of tau^2. The equilibrium bias magnitude scales with [(1 - sqrt(1 - 4*tau^2))/2], which increases in tau^2. The sign of bias is determined by the direction of (theta_t - y_T). The paper does not present cross-sectional or time-series panel heterogeneity — the parametric sensitivity analysis in Figure 3 constitutes the heterogeneity exercise.&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-are-run-for-the-greenbook-empirical-patterns"&gt;Q10. What robustness checks are run for the Greenbook empirical patterns?&lt;/h3&gt;
&lt;p&gt;The authors state (in a footnote) that the documented patterns — persistent but sign-changing bias in 4-quarter-ahead GB inflation forecasts from 1980q1 to 2019q4 — are robust to using the second release of the GDP deflator rather than the last release. The main results use the last release. The choice of 40-quarter (10-year) rolling window is applied uniformly for both the bias plot and the MZ slope plot. No additional robustness checks (alternative window lengths, alternative forecast horizons, formal structural break tests) are explicitly documented in the paper, though the authors cite Rossi and Sekhposyan (2016), who use formal rationality tests and confirm that GB forecast rationality breaks down around 2005 — consistent with the pattern the authors document via the rolling MZ slope.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-model-say-about-the-forecasters-inability-to-commit-and-could-commitment-help"&gt;Q11. What does the model say about the forecaster&amp;rsquo;s inability to commit, and could commitment help?&lt;/h3&gt;
&lt;p&gt;In the baseline model, the forecaster cannot commit to a fixed forecasting rule ex ante because the state theta_t is not directly observable by the DM. The authors note in Section 3.3 that modeling forecasters with commitment is a straightforward extension, and that commitment can actually increase forecaster welfare in equilibrium. However, this extension is not formally developed in the paper. The intuition is that if the forecaster could credibly commit to a more informative forecast rule, the DM could react more precisely, reducing the variance of outcomes; but without commitment, the strategic equilibrium involves an attenuated (biased) forecast.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-implications-for-forecast-rationality-tests-and-loss-function-estimation"&gt;Q12. What are the implications for forecast rationality tests and loss function estimation?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s central methodological warning is that standard forecast rationality tests (MZ regression tests for zero intercept and unit slope; bias tests) and loss function estimation exercises are contaminated in environments with policy feedback. If feedback is present and x_t is uncertain, a fully rational forecaster with quadratic loss will produce forecasts that fail standard rationality tests — showing nonzero bias, non-unit MZ slopes (potentially even negative), and forecast errors correlated with the forecaster&amp;rsquo;s own information. Researchers conducting such tests must either: (a) explicitly assume no feedback applies (and justify this assumption in their specific application), or (b) carefully model the feedback mechanism and account for it. Studies that interpret GB forecast irrationality (e.g., Rossi and Sekhposyan 2016) or asymmetric loss (e.g., Capistran 2008) as the explanation for observed GB forecast properties may be confounded by the feedback mechanism identified in this paper.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-conditions-under-which-a-linear-equilibrium-does-or-does-not-exist"&gt;Q13. What are the conditions under which a linear equilibrium does or does not exist?&lt;/h3&gt;
&lt;p&gt;From Corollary 1 and Remark 3 following it: a linear PBE exists if and only if tau^2 &amp;lt;= 1/4. When tau^2 &amp;gt; 1/4, the forecaster always wants to attenuate the slope more than the DM expects, so no fixed-point equilibrium in linear strategies exists. The paper also notes a sufficient condition for equilibrium existence: if the support of x_t is contained in [0, 1] (the DM never overreacts and never underreacts by more than half), then tau^2 &amp;lt;= 1/4 is automatically satisfied and an equilibrium always exists. Two linear equilibria exist when tau^2 &amp;lt;= 1/4, but the paper focuses on the Pareto-preferred one, which has lower forecaster loss, lower absolute bias, and a natural limiting behavior as tau^2 approaches 0.&lt;/p&gt;
&lt;h3 id="q14-what-scope-conditions-limit-the-applicability-of-the-results"&gt;Q14. What scope conditions limit the applicability of the results?&lt;/h3&gt;
&lt;p&gt;Several scope conditions are made explicit: (1) The outcome equation is linear; nonlinear outcome determination would change quantitative results but the feedback mechanism would persist qualitatively. (2) The model is a single-period (point-in-time) game, not a multi-period learning model — it does not analyze how beliefs about mu and tau^2 evolve over time. (3) The independence assumption between x_t and theta_t is a benchmark; if policy aggressiveness varies with economic conditions, additional effects arise. (4) The focus on linear equilibria rules out non-linear forecasting strategies. (5) The results apply to unconditional forecasts (where the forecaster anticipates the DM&amp;rsquo;s response); conditional forecasts (conditioned on a pre-specified action) behave differently. (6) The empirical Greenbook evidence is illustrative, not a formal test of the model — the authors explicitly state they do not claim their model provides an exclusive explanation of GB forecast properties.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Forecasting with feedback&lt;/strong&gt;: A forecasting environment in which the DM&amp;rsquo;s action — taken in response to the forecast — causally affects the realized value of the variable being forecast, so that the forecast influences its own target outcome. Distinguished from no-feedback environments (e.g., weather forecasting) where decisions made on the basis of the forecast do not affect the outcome.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Unconditional forecast&lt;/strong&gt;: A forecast that anticipates and factors in the expected response of the decision maker to the forecast itself, rather than being conditioned on a pre-specified (potentially counterfactual) action. The paper&amp;rsquo;s model produces unconditional forecasts; conditional forecasts (conditioned on a given policy path) are a distinct and narrower concept.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bias-variance tradeoff (in feedback forecasting)&lt;/strong&gt;: The tradeoff that arises when the DM&amp;rsquo;s reaction to the forecast is uncertain: a less informative (attenuated) forecast reduces the variance of the outcome (by inducing a less volatile policy action) but introduces systematic bias. The optimal forecast under quadratic loss resolves this tradeoff by attenuating the forecast slope below what an unbiased forecast would require, producing an optimally biased forecast.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reaction function (DM&amp;rsquo;s)&lt;/strong&gt;: The rule by which the decision maker translates a forecast into a policy action: a_t* = x_t * [y_T - E(theta_t | f_t)], analogous to a Taylor rule. The multiplier x_t captures the strength of the policy response and is drawn from a distribution with mean mu and variance tau^2; it is the DM&amp;rsquo;s private information and a key source of the forecaster&amp;rsquo;s uncertainty.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mincer-Zarnowitz (MZ) regression&lt;/strong&gt;: The linear regression of the realized outcome on the forecast: y_{t+1} = alpha + beta * f_t + error. Under the canonical null of rational forecasting with quadratic loss and no feedback, the intercept alpha should be zero and the slope beta should be one. The paper shows that under optimal forecasting with feedback, alpha and beta can take a wide range of values, including negative beta, even when the forecaster is rational.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Equilibrium forecast slope (c-dagger)&lt;/strong&gt;: The slope of the linear forecasting rule in Perfect Bayesian Equilibrium, given by c^dagger = (1/2) - mu + sqrt(1 - 4*tau^2)/2. This slope is less than one and can be negative depending on mu and tau^2, reflecting the attenuation of the forecast toward the policy target that arises from the bias-variance tradeoff under uncertain DM reactions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Greenbook (GB) inflation forecasts&lt;/strong&gt;: Inflation forecasts produced by Federal Reserve staff (now called Tealbook forecasts), used as empirical motivation in the paper. The paper documents two stylized facts for 4-quarter-ahead GB forecasts from 1980q1 to 2019q4: (i) persistent but sign-changing bias in rolling 40-quarter windows, and (ii) a dramatic shift in the rolling MZ slope from approximately unity in the 1980s-1990s to significantly negative in the mid-2000s and approximately zero in the final part of the sample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy feedback (as a confound for rationality tests)&lt;/strong&gt;: The paper&amp;rsquo;s use of this term to describe the mechanism by which the presence of feedback invalidates the standard interpretation of forecast rationality test outcomes: a forecaster who is fully rational (quadratic loss, no private agenda) and operating in a feedback environment will systematically produce forecasts that fail standard MZ-based rationality tests, not because of irrationality or asymmetric loss, but because of the optimal bias-variance tradeoff induced by uncertain policy reactions.&lt;/p&gt;</description></item><item><title>Global Value Chains and Labor Standards: The Race-to-the-Bottom Problem</title><link>https://macropaperwarehouse.com/papers/global-value-chains-and-labor-standards-the-race-to-the-bottom-problem/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/global-value-chains-and-labor-standards-the-race-to-the-bottom-problem/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Im and McLaren (2025) ask whether globalization induces governments to weaken labor standards for workers — the so-called &amp;ldquo;race to the bottom&amp;rdquo; (RTB) hypothesis. The question has high stakes: advocates point to events such as the 1,136-worker Rana Plaza factory collapse in Bangladesh (2013) and to India&amp;rsquo;s deregulation campaign after 2014 (associated with approximately 6,500 workplace deaths in 2015–2020) as evidence that competition for global capital systematically erodes safety and working conditions. The paper builds a stylized many-country equilibrium model of labor-market integration adapted from the Grossman and Rossi-Hansberg (2008) tasks framework. Output requires a continuum of tasks z in [0,1], performable in any of N countries; labor requirements per task follow a Weibull distribution (shape parameter nu &amp;gt; 0), independently across tasks and countries. Working conditions (kappa_i) enter the cost function multiplicatively — better conditions reduce worker productivity at the relevant margin. Utility is separable in wages and conditions with both components strictly concave, and Assumption 1 (x&lt;em&gt;xi&amp;rsquo;(x) and x&lt;/em&gt;mu&amp;rsquo;(x) strictly decreasing) ensures conditions are normal goods and second-order conditions hold. The unregulated equilibrium task allocation is equivalent to CES cost minimization with elasticity of substitution 1/(1-rho) &amp;gt; 1, rho = nu/(1+nu). Governments set minimum standards non-cooperatively in Nash equilibrium.\n\nThe paper&amp;rsquo;s results fall into two conceptually distinct categories. &amp;ldquo;Globalization in the large&amp;rdquo; (autarky vs. open economy): whether standards are market-determined or government-set, integrating two previously autarkic countries raises labor standards in both (Proposition 1). Under autarky, market and government-optimal conditions coincide — all costs of better standards are borne domestically. Under trade, wages rise (income channel: conditions are a normal good), and governments gain a terms-of-trade incentive: tightening kappa_i makes domestic effective labor scarcer and shifts part of the cost onto foreign consumers, inducing government standards to strictly exceed market standards. Formally, for each country i: autarky level = market level under autarky &amp;lt; market level under integration &amp;lt; government level under integration.\n\n&amp;quot;Globalization at the margin&amp;quot; with symmetric countries (Proposition 2): as more identical countries join (N increasing), both market-set and government-set standards rise monotonically. The terms-of-trade motive does not vanish because each country specializes in an increasingly narrow value-chain slice, retaining market power regardless of N. Government standards exceed market standards for every N &amp;gt;= 2 and grow strictly with N — a race to the top — and are shown to be above the social optimum because each country externally imposes part of its improvement costs on others.\n\n&amp;quot;Globalization at the margin&amp;quot; with a North-South structure (Proposition 3): when Southern host countries (i = 2,&amp;hellip;,N) have perfectly correlated productivity draws (close substitutes for one another), the result reverses for N &amp;gt; 2. Integration of two countries initially raises Southern standards via both channels. But as additional similar Southern competitors join, competition depresses Southern wages and erodes both the income-based demand for better conditions and the terms-of-trade motive (unilateral tightening redirects demand to competitors without cost-shifting benefit). Both market and government standards fall monotonically as N rises beyond 2. As N approaches infinity, both converge to autarky levels. Critically, however, for any finite N, Southern standards remain strictly above their autarky levels — the race to the bottom, even when operative, never fully materializes while integration is incomplete. The efficiency implication is counter-intuitive: government-set standards are inefficiently strict under GVCs because each country over-provides standards by externalizing costs onto trading partners.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-models-formal-structure-and-how-does-it-generate-tractable-results"&gt;Q1. What is the model&amp;rsquo;s formal structure and how does it generate tractable results?&lt;/h3&gt;
&lt;p&gt;The model adapts Grossman and Rossi-Hansberg (2008). Output requires a unit measure of tasks; labor requirement for task z in country i is A_i * a^i_z, where A_i = bar_A_i * kappa_i, so working conditions raise unit labor costs. Each a^i_z is drawn Weibull(nu, 1) independently. A result (adapted from Anderson et al. 1987, applied by Artuç and McLaren 2015) is that the cost-minimizing task allocation is equivalent to minimizing cost with a CES aggregate of national effective labor supplies, with elasticity of substitution 1/(1-rho) and rho = nu/(1+nu). This reduces the multi-dimensional problem to a standard CES factor-demand problem, yielding closed-form wage equations and tractable Nash equilibrium characterizations.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-channels-driving-globalization-in-the-large-raising-standards-above-autarky"&gt;Q2. What are the two channels driving &amp;lsquo;globalization in the large&amp;rsquo; raising standards above autarky?&lt;/h3&gt;
&lt;p&gt;Two reinforcing channels. First, the income channel: integration raises real wages (gains from specialization), and since working conditions are a normal good under Assumption 1 (utility sufficiently concave), demand for better conditions rises. Second, the terms-of-trade channel: tightening kappa_i makes domestic effective labor more expensive and scarcer; part of the resulting cost increase is borne by foreign consumers and workers via the unit cost identity rather than solely by domestic workers. This cost-shifting gives governments an incentive to tighten standards beyond what the unregulated market sets. The mechanism is formally analogous to the policy externalities in Bagwell and Staiger (2001) and the terms-of-trade motive in Chau and Kanbur (2006), though the latter has no value chains.&lt;/p&gt;
&lt;h3 id="q3-why-does-the-terms-of-trade-motive-for-over-regulation-persist-even-as-the-number-of-symmetric-countries-approaches-infinity"&gt;Q3. Why does the terms-of-trade motive for over-regulation persist even as the number of symmetric countries approaches infinity?&lt;/h3&gt;
&lt;p&gt;As more countries join, each specializes in an increasingly narrow slice of the value chain in which it has comparative advantage. This deepening specialization preserves market power: the wage derivative dw_1/d_kappa_1 converges to a limit proportional to rho*w/kappa (strictly greater than the pure autarky productivity effect -w/kappa) rather than to zero. So even in the limit with infinitely many symmetric countries, each country retains some terms-of-trade gain from tightening its standard, and government standards keep rising above market standards.&lt;/p&gt;
&lt;h3 id="q4-under-what-precise-conditions-does-the-race-to-the-bottom-result-hold"&gt;Q4. Under what precise conditions does the race-to-the-bottom result hold?&lt;/h3&gt;
&lt;p&gt;The RTB result (Proposition 3) requires that competing host countries be close substitutes for one another. The paper operationalizes this with the extreme case of perfectly correlated productivity draws across Southern countries (a^i_z = a^2_z for all i &amp;gt;= 2 and all tasks z). Under this structure, as N increases from 2 onward, Southern market and government standards fall monotonically toward autarky levels. The mechanism: competition among near-identical countries means unilateral tightening of kappa_2 redirects Northern demand to competitors without generating a terms-of-trade gain for Country 2, so the wage falls and conditions deteriorate. The RTB thus requires high substitutability among competitors, not just trade openness.&lt;/p&gt;
&lt;h3 id="q5-does-the-race-to-the-bottom-ever-drive-standards-below-autarky-levels"&gt;Q5. Does the race to the bottom ever drive standards below autarky levels?&lt;/h3&gt;
&lt;p&gt;No. Proposition 3 parts (i) and (ii) establish that for any finite N &amp;gt;= 2, both market-set and government-set standards in Southern countries remain strictly above their autarky levels. The race is toward (but never below) the autarky benchmark. Only in the limit as N approaches infinity do standards converge to the autarky level (Proposition 3, part iii). For any realistic finite degree of globalization, even the worst-case RTB scenario leaves standards strictly above autarky.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-efficiency-implication-of-nash-equilibrium-government-set-standards"&gt;Q6. What is the efficiency implication of Nash equilibrium government-set standards?&lt;/h3&gt;
&lt;p&gt;Government-set standards under GVCs are inefficiently strict. Each government maximizes domestic welfare ignoring the cost its tightening imposes on foreign consumers and workers. Because tightening kappa_i raises costs partly borne abroad, each government over-provides standards relative to the global social optimum. This is a race to the top that generates a negative international externality — the mirror image of the usual RTB externality. The implication is that international coordination, if it occurred, would likely reduce Nash equilibrium standards toward the optimum, not raise them further.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-papers-setting-differ-from-prior-theoretical-work-on-the-race-to-the-bottom"&gt;Q7. How does the paper&amp;rsquo;s setting differ from prior theoretical work on the race to the bottom?&lt;/h3&gt;
&lt;p&gt;Prior RTB models (Chau and Kanbur 2006; Felbermayr et al. 2012; Chen and Dar-Brodeur 2020) model countries competing for export markets — competing to sell goods to a common importer — rather than competing to host tasks in global value chains. The current paper frames globalization as an increase in the number of countries that can supply tasks to a common production process, a qualitatively different competitive margin. Prior work also largely takes the degree of globalization as fixed, while this paper explicitly traces out effects as N changes. The distinction between similar versus different competitors as a determinant of the direction of the RTB is also new. The companion paper Im and McLaren (NBER WP 31363) extends the framework to collective-bargaining rights with an empirical component.&lt;/p&gt;
&lt;h3 id="q8-what-heterogeneity-is-documented-and-what-does-it-imply"&gt;Q8. What heterogeneity is documented and what does it imply?&lt;/h3&gt;
&lt;p&gt;The paper develops two polar cases of country heterogeneity: (1) symmetric countries with independent productivity draws — produces a race to the top as N rises; (2) North-South structure with correlated (identical) Southern productivity draws — produces a race to the bottom as N rises beyond 2. The contrast is the central result: the direction of the marginal effect of globalization on standards depends on the degree of substitutability among competing host countries. The authors connect this to observed patterns — Korean firms relocating only to East Asian affiliates (similar countries) when domestic minimum wages rose, and Chan and Ross (2003) noting that competition is &amp;lsquo;most vicious not between North and South, but among nations of the South.&amp;rsquo;&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The core implication is that trade restrictions justified by RTB concerns lack general theoretical support — globalization relative to autarky always raises standards. However, the model validates a targeted RTB concern: when a country faces competition from many similar low-wage countries (e.g., Mexico competing with China in labor-intensive sectors), standards can erode relative to the peak reached under limited integration. The appropriate response in that case is to integrate with structurally different partners (as Mexico did via NAFTA with the US) rather than restrict trade. Since Nash equilibrium standards already exceed the global optimum, international agreements that ratchet standards up further could be welfare-reducing. The paper explicitly cautions that causation is hard to establish in the Mexico-China-NAFTA example, treating it as suggestive illustration rather than proof.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-main-limitations-and-threats-to-the-conclusions"&gt;Q10. What are the main limitations and threats to the conclusions?&lt;/h3&gt;
&lt;p&gt;The paper is entirely theoretical; no empirical test is conducted for working conditions (the authors cite data scarcity as the reason, having a companion empirical paper on collective-bargaining rights instead). Key assumptions include: (a) Weibull, independent task-productivity draws (ensure tractability but are untested); (b) working conditions always reduce productivity at the margin (rules out the many cases where safety improvements also raise output — e.g., Alfaro-Ureña et al. 2021 find no productivity effect of responsible sourcing in Costa Rica, suggesting the trade-off assumption is plausible but not universal); (c) citizen activism, which empirically affects labor standards (Harrison and Scorse 2010; Koenig and Poncet 2019, 2022), is abstracted away; (d) the model has a single final good and no intermediate goods trade beyond the task-allocation interpretation, limiting applicability to multi-sector settings.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Labor standards (kappa_i)&lt;/strong&gt;: In the paper&amp;rsquo;s specific sense, the quality of working conditions that (i) raise worker utility holding wages fixed and (ii) increase unit labor costs for employers. Explicitly restricted to improvements that involve a trade-off — e.g., safety provisions, clean bathrooms, break times — excluding complementary improvements that raise both utility and productivity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Globalization in the large&lt;/strong&gt;: The paper&amp;rsquo;s term for the comparison of any open-economy equilibrium (N &amp;gt;= 2 countries integrated) against autarky. Result: labor standards are always strictly higher in the open economy whether market-set or government-set, because income rises and the terms-of-trade motive activates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Globalization at the margin&lt;/strong&gt;: The paper&amp;rsquo;s term for the effect on labor standards of adding one more country to an already-integrated economy (increasing N by 1). This effect is ambiguous: it raises standards when new entrants are dissimilar (symmetric model) and lowers them when new entrants are similar (North-South model).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Terms-of-trade effect (labor-standards channel)&lt;/strong&gt;: The mechanism by which tightening a country&amp;rsquo;s labor standard (raising kappa_i) reduces domestic effective labor supply, raises the relative price of domestic tasks, and shifts part of the cost improvement onto foreign consumers and workers. This creates an incentive for governments to set standards above the market level and above the global social optimum — producing standards that are too strict from an efficiency standpoint.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Normal good (working conditions)&lt;/strong&gt;: The property implied by Assumption 1 (both x&lt;em&gt;xi&amp;rsquo;(x) and x&lt;/em&gt;mu&amp;rsquo;(x) strictly decreasing in x) that workers&amp;rsquo; marginal valuation of working conditions relative to wages is higher at higher income levels. This ensures that any source of income gains — including gains from trade — mechanically raises equilibrium demand for better working conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Race to the top&lt;/strong&gt;: The paper&amp;rsquo;s characterization of the symmetric-countries equilibrium: as N increases, both market-set and government-set labor standards rise monotonically, because market power persists through value-chain specialization and the terms-of-trade motive remains strong. Government standards also exceed the social optimum, making this over-regulation an externality imposed on trading partners.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Race to the bottom (conditional)&lt;/strong&gt;: The result in the North-South model where additional similar Southern host countries erode Southern labor standards as N rises beyond 2. The race is toward autarky levels but never below them for finite N. The RTB requires high substitutability among competing host countries and does not hold as a general consequence of globalization.&lt;/p&gt;</description></item><item><title>Illuminating the Global South</title><link>https://macropaperwarehouse.com/papers/illuminating-the-global-south/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/illuminating-the-global-south/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Satellite nighttime lights (luminosity) are the dominant remote-sensing proxy for local economic conditions in low-income countries, yet their accuracy at fine spatial scales and over time has remained contested. This paper by Chiovelli, Michalopoulos, Papaioannou, and Regan makes two linked contributions. First, it constructs a standardized, annual, global panel of nighttime lights from 1992 to 2023, integrating the legacy DMSP-OLS satellite series (1992–2013) with the higher-quality VIIRS series (2013–onward) after applying three adjustments to the noisier DMSP data: cross-sensor inter-calibration (following Li et al. 2020), top-coding correction (following Bluhm and Krause 2022, using a truncated Pareto distribution to replace pixels with Digital Number ≥ 55), and blooming correction (following Cao et al. 2019, modeling light spillover as spatial decay and subtracting predicted pseudo-light). VIIRS is then downgraded to DMSP-comparable units using an ensemble machine-learning method — extremely randomized trees trained on the single year of full overlap (2013) — yielding an out-of-sample RMSE of 1.50 versus 3.27 for the Li et al. sigmoid approach and 1.57 for the Nechaev et al. convolutional neural network; the F1 score for the binary lit/unlit classification is 0.72 versus 0.51 and 0.71 for those alternatives, with recall = 0.95 and precision = 0.58 against an actual lit-pixel share of only 8.6 percent globally. At the cross-country level — a sample of 173 countries — the adjusted series retains an elasticity of luminosity to GDP of approximately 0.85 and an R² around 0.9 in cross-section; for Africa specifically the elasticity is 0.7 and R² remains around 0.9. In long-difference panel regressions over 1992–2019, the luminosity-GDP elasticity is approximately 0.25–0.24, broadly consistent with Henderson et al. (2012)&amp;rsquo;s estimate of 0.30–0.33, while at the five-year panel frequency the elasticity is around 0.15–0.17. The second contribution is a systematic validation of the new series against multiple local development proxies across four low-income settings. Using 139 georeferenced DHS surveys from 34 African countries (gridcells of ~28km × 28km), the adjusted series yields cross-sectional coefficients of approximately 0.6 standard deviations for schooling, electricity access, and improved sanitation, and approximately 1 standard deviation for the composite wealth index, between lit and unlit gridcells; in within-gridcell panel regressions, the adjusted log-lights coefficient on schooling is approximately double that of the unadjusted series (~0.02 versus ~0.01), and lit/unlit panel coefficients are statistically significant only with the adjusted series — gridcells turning lit see schooling rise by ~0.05 standard deviations (~0.125 schooling years), wealth index rise by ~0.05 SD, and electricity access rise by ~0.05 SD. In Mozambique, using all post-civil-war censuses (1997, 2007, 2017) across 1,126 admin-4 localities, schooling and non-agricultural employment are at least 0.5 standard deviations higher in lit than unlit localities, equivalent to approximately 0.5 years of schooling and 10 percentage points of non-agricultural employment; within-locality changes in lights co-move significantly with schooling changes, with the difference in schooling gain between localities that turn lit versus stay unlit being about half a year even controlling for admin-3 fixed effects. In Indonesia, panel estimates for public goods across more than 60,000 PODES villages show the adjusted series yields a positive and significant coefficient on the composite wealth index while the unadjusted series yields a counterintuitively negative coefficient. In India, across more than 550,000 SHRUG villages and towns, the adjusted series consistently produces stronger cross-sectional and panel associations with non-farm, manufacturing, and services employment. A key empirical regularity across all settings is that the adjusted series outperforms the unadjusted one most sharply at finer spatial resolutions and in over-time (panel) comparisons, while at coarse aggregation levels (large administrative units or large grid squares) differences between the two series are minor, as spatial averaging attenuates measurement error in the unadjusted data too. Blooming correction delivers most of the improvement in the African context, where top-coding is rare (fewer than 2% of lit DMSP pixels in Africa approach the 63 DN ceiling). The paper also replicates three canonical studies — Michalopoulos and Papaioannou (2013) on precolonial ethnic institutions, Michalopoulos and Papaioannou (2014) on national institutions and split ethnic homelands, and Hodler and Raschky (2014) on regional favoritism — confirming that qualitative conclusions are robust to the data revision while documenting that the adjusted series sharpens several estimates, particularly those exploiting within-region over-time variation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper is a measurement and validation study rather than a causal identification exercise. Its core design is correlational: it regresses local development proxies on nighttime luminosity across gridcells and administrative units, conditioning on country-year fixed effects in cross-section and on unit fixed effects in panel regressions. The main threats are (a) reverse causation (luminosity and development are jointly determined), which the authors acknowledge but do not attempt to address — they are explicit that the goal is proxy validation, not causal estimation; (b) measurement error in both the luminosity variable and the development outcomes (DHS wealth index, census schooling, PODES public goods), which the paper addresses by comparing adjusted versus unadjusted luminosity series and interpreting attenuation bias reduction as evidence of improved measurement; (c) the binary transformation of luminosity (lit/unlit) produces non-classical measurement error — an explicit point drawn from econometric theory (Aigner 1973; Meyer and Mittag 2017) — which partly motivates the adjusted continuous series; and (d) spatial autocorrelation and systematic geographic patterns in prediction error, which the authors check by regressing prediction errors on latitude and longitude and find that the ERT-downgraded series reduces the latitude coefficient to 10% of its magnitude in the unadjusted VIIRS specification for log lights and to 35% for the lit indicator.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-dmsp-deficiencies-corrected-and-what-are-the-specific-methods-used"&gt;Q2. What are the three DMSP deficiencies corrected and what are the specific methods used?&lt;/h3&gt;
&lt;p&gt;Cross-sensor inter-calibration: DMSP data come from six satellites; Li et al. (2020) supply a cross-calibrated series using a second-order polynomial fitted on overlapping satellite years, which the paper adopts as its &amp;lsquo;unadjusted&amp;rsquo; baseline. Top-coding: DMSP records 8-bit Digital Numbers (DN) 0–63, so radiance above a ceiling is truncated. Pixels with DN ≥ 55 are subject to &amp;lsquo;implicit&amp;rsquo; top-coding (averages of potentially top-coded sub-readings). The correction uses the radiance-calibrated (RC) vintage available for seven years, ranks the top-coded pixels by the RC series from the nearest year, then replaces them with &amp;lsquo;structural values&amp;rsquo; drawn from a truncated Pareto distribution with parameters α = 1.5, L = 55, H = 2000. Blooming: the DMSP sensor stretches edge pixels and can be spatially displaced up to 3 km, causing light spillover. Following Cao et al. (2019), pseudo-light pixels (PLPs) — lit pixels neighboring at least one dark pixel — are identified. An OLS regression of PLP light on the inverse-squared-distance weighted sum of neighbors&amp;rsquo; light within a 7 × 7 window is estimated separately for broad global regions. The predicted blooming contribution is subtracted from each lit pixel, negative residuals are set to zero, and a local 3 × 3 mean smoothing is applied. Globally, the blooming correction raises the share of unlit pixels from 92% to 95% in 1992 and from 88% to 91% in 2012.&lt;/p&gt;
&lt;h3 id="q3-how-is-viirs-downgraded-and-harmonized-with-dmsp-and-what-does-extremely-randomized-trees-mean"&gt;Q3. How is VIIRS downgraded and harmonized with DMSP, and what does &amp;rsquo;extremely randomized trees&amp;rsquo; mean?&lt;/h3&gt;
&lt;p&gt;Because VIIRS records 14-bit DN at 15-arc-second resolution with far superior sensor quality, it is not directly comparable to the 8-bit, 30-arc-second DMSP. The authors&amp;rsquo; preferred approach downgrades VIIRS to match the DMSP scale. They use an ensemble machine-learning method called &amp;rsquo;extremely randomized trees&amp;rsquo; (Geurts et al. 2006), a variant of random forests that, instead of choosing the best splits from the training sample, picks split thresholds randomly, which further reduces variance and improves computational efficiency. Features used to predict DMSP-like values from VIIRS include: pixel statistics (mean, median, min, max of the four VIIRS sub-pixels within each DMSP 30-arc-second cell), statistics of neighboring pixels within windows of 3, 4, 7, 9, 11, 13, 17, and 21 pixel widths, and regional dummies for broad world regions. The model is trained on 2013 (the one full year of DMSP-VIIRS overlap) and its out-of-sample performance is assessed by retraining on 2012 and predicting 2013. Four merged series are produced corresponding to the four versions of DMSP (unadjusted; blooming only; top-coding only; both). The authors&amp;rsquo; approach outperforms both the Li et al. (2020) sigmoid-function method (RMSE 3.27 globally vs. 1.50) and the Nechaev et al. (2021) CNN approach (RMSE 1.57), especially in the low-to-middle luminosity range most relevant for low-income countries.&lt;/p&gt;
&lt;h3 id="q4-what-development-proxies-are-used-in-validation-and-across-what-samples"&gt;Q4. What development proxies are used in validation and across what samples?&lt;/h3&gt;
&lt;p&gt;Africa (DHS, 34 countries, 139 surveys, ~28km × 28km gridcells): mean years of schooling (respondents aged 15–39), DHS composite household wealth index, share of households with improved sanitation, share with electricity connection. All outcomes are standardized to mean zero, SD one. Mozambique (Census 1997, 2007, 2017, 1,126 admin-4 localities): mean years of schooling (aged 15–39) and non-agricultural employment (aged 15–24 or 19–24). Indonesia (PODES village census waves 1996–2018, 60,000+ villages): binary measures for garbage disposal, toilet use, drinking water access, gas/electricity for cooking, paved roads, and counts of kindergartens, primary, middle, and secondary schools — aggregated into a first principal component (eigenvalue ~3.5, capturing ~1/3 of variance). India (SHRUG dataset, 550,000+ towns and villages, Population Censuses 1991/2001/2011, Economic Censuses 1990/1998/2005/2013): population count, total non-farm employment, manufacturing employment, services employment.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented"&gt;Q5. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Spatial resolution: adjusted series outperforms unadjusted most at fine resolutions (2×2 gridcell blocks, ~56km × 56km at the equator); at coarse levels (12×12 blocks, ~336km × 336km), both series yield similar coefficients, as spatial aggregation attenuates noise in the unadjusted series. Urban vs. rural: cross-sectional estimates are similarly significant in urban and rural DHS samples. Panel estimates are statistically significant only with the adjusted series; urban panel coefficients are consistently larger than rural ones, echoing Asher et al. (2021)&amp;rsquo;s India finding. The adjustment matters more in rural areas than in urban areas in cross-section. Local variation (spatial RDD / fine fixed effects): with unadjusted series, panel wealth-index coefficients are statistically indistinguishable from zero until spatial fixed effects cover areas at least 7×7 gridcells (~200km × 200km at equator); with the adjusted series, coefficients remain significantly positive at all fixed-effect sizes including the finest 2×2 blocks. Top-coding vs. blooming: most of the improvement in Africa derives from blooming correction; top-coding correction has minor impact because fewer than 2% of lit African DMSP pixels approach the DN ceiling. Country-ethnic homelands (large areas, avg. 25,547 km²): adjustments matter little because spatial averaging already reduces noise. Applications replication: the precolonial institutions result (Michalopoulos and Papaioannou 2013) is robust and essentially unchanged because the units are very large. The national-institutions-at-border result (Michalopoulos and Papaioannou 2014) is strengthened in within-ethnicity specifications (coefficient marginally significant at 90% with adjusted series vs. p ≈ 0.15 with unadjusted); capital-proximity heterogeneity is sharpened. The regional-favoritism result (Hodler and Raschky 2014) strengthens: the log-lights lagged-leader coefficient rises from 0.038 to 0.058, and the lit-probability coefficient rises from ~3 to ~7 percentage points.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-and-specification-variations-are-run"&gt;Q6. What robustness checks and specification variations are run?&lt;/h3&gt;
&lt;p&gt;The paper compares four luminosity series (unadjusted Li et al.; blooming only; top-coding only; both combined + VIIRS fusion) to isolate each correction&amp;rsquo;s contribution. It checks the luminosity-GDP nexus at annual, five-year, and long-difference frequencies. It examines seven African countries&amp;rsquo; co-evolution of the harmonized series with electrification share (Kenya, DRC, Ghana, Tanzania, Nigeria, Mozambique, and one other) and finds no discontinuity at the 2012/2013 DMSP-VIIRS transition year. Spatial aggregation robustness: coefficients are computed across aggregation blocks ranging from 2×2 to 12×12 gridcells, showing stability in cross-section (~0.18) and mild size dependence in panel (~0.075, slightly rising with coarser units). Local variation robustness: fixed effects of increasing spatial coverage (2×2 to 12×12 cells) are added while the outcome remains at the gridcell level. Results replicated for schooling and electricity access (Appendix Section B.2) beyond the primary wealth-index outcome. Confounding by latitude in the ML model is assessed via regressions of prediction errors on latitude and longitude with and without country fixed effects. Median regressions confirm the OLS elasticity estimates at the cross-country level. The India analysis is replicated for both towns (urban) and villages (rural) separately.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Henderson et al. (2012): pioneer the use of luminosity as a cross-country GDP proxy and estimate a long-difference elasticity of 0.30–0.33 across 188 countries; this paper estimates 0.25–0.24 over a comparable specification, consistent but slightly lower. Gibson et al. (2021): show that VIIRS is superior to DMSP but find weak GDP-lights correlations outside cities for the early DMSP period in China, Indonesia, and South Africa; this paper addresses the concern by adjusting DMSP and merging it with VIIRS. Asher et al. (2021): validate luminosity as a strong proxy in India and find stronger urban-luminosity links; this paper replicates and extends those findings to Africa, Mozambique, and Indonesia and shows the adjusted series strengthens the Asher et al. patterns. Chen et al. (2024): find strong cross-sectional but weak panel associations; this paper&amp;rsquo;s adjusted series substantially strengthens panel associations. Bluhm and Krause (2022): provide the top-coding correction method adopted here. Cao et al. (2019): provide the blooming correction method. Nechaev et al. (2021): propose a CNN-based DMSP-VIIRS fusion but apply it to the unadjusted DMSP; this paper outperforms their RMSE slightly (1.50 vs. 1.57) and improves on their F1 score (0.72 vs. 0.71), with greater advantage in low-light regions. Li et al. (2020): propose a sigmoid-based fusion calibrated for high-light pixels; this paper substantially outperforms it (RMSE 1.50 vs. 3.27) particularly in low-luminosity areas. The paper thus synthesizes and extends multiple strands: it unifies the corrections of Bluhm-Krause and Cao et al., pairs them with state-of-the-art ensemble ML fusion, and provides by far the most comprehensive multi-country, multi-context validation of the resulting series.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary policy implication is methodological: researchers studying development in low-income countries should use the adjusted and harmonized nighttime lights series rather than raw DMSP data, and should be especially careful at fine spatial scales (e.g., spatial regression discontinuity designs, granular village-level analyses) and in panel specifications. The gains from adjustment are largest precisely where applied development research is moving — toward local identification strategies and over-time variation. For practitioners and statistical agencies, the series provides a low-cost annual proxy for local economic conditions in environments with weak administrative data, particularly across sub-Saharan Africa, South Asia, and Southeast Asia. Scope conditions: (a) Correlations are far from perfect — binary lit/unlit classification misses much variation in the many-zeros low-income context. (b) At large aggregate units (admin-1, country-ethnic homelands), the adjustments yield minimal additional improvement since noise averages out. (c) The series does not resolve the fundamental limitation that most of sub-Saharan Africa remains unlit (98.4% of DMSP pixels in Africa in 1992), so it captures variation among already-lit areas better than the development gradient at the zero-light frontier. (d) Future research blending nighttime lights with daytime imagery (traffic, built structures) is flagged as a promising extension, though daytime data are often proprietary.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-main-findings-from-the-three-replication-exercises"&gt;Q9. What are the main findings from the three replication exercises?&lt;/h3&gt;
&lt;p&gt;Michalopoulos and Papaioannou (2013) — precolonial ethnic institutions and contemporary development: Replication across 682 country-ethnic homelands confirms that areas with higher precolonial political centralization (as measured by a 0–4 jurisdictional hierarchy index) have significantly higher contemporary luminosity, conditional on country constants and geographic controls. With the adjusted series, the unlit share among homelands rises from 24% to 29% (because blooming correction removes spurious light), but the coefficients on political centralization are still highly significant, somewhat smaller in magnitude, and similar qualitatively. The main conclusion is robust because the units are large and spatial averaging already reduces noise in the raw series. Michalopoulos and Papaioannou (2014) — national institutions and split-border ethnic development: Replication across 38,427 gridcells of 220 systematically partitioned ethnic homelands. Cross-sectional results show a one-point increase in the rule-of-law index (range −2.5 to 2.5) is associated with a ~10 pp higher probability of a gridcell being lit. The within-ethnicity coefficient drops by more than half (~0.025). With the adjusted series, this within-ethnicity coefficient is marginally significant at 90% versus a p-value of ~0.15 with unadjusted. Spatial RDD coefficients remain small and insignificant regardless of adjustment. Capital-proximity heterogeneity: the positive association between rule of law and luminosity is significant only for ethnically split groups where both portions are close to their respective capitals, and this finding is more precisely estimated with the adjusted series; the effect is nil far from capitals in both series. Hodler and Raschky (2014) — regional favoritism: Panel replication across 38,427 subnational regions in 126 countries, 1992–2009. The lagged-leader dummy coefficient (log lights specification) rises from 0.038 to 0.058 with the adjusted series. The linear-probability-model lit indicator rises from ~3 to ~7 percentage points. All specifications with the adjusted series are at least two standard errors above zero, matching or exceeding the precision of the original.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-limitations-and-caveats-acknowledged-by-the-authors"&gt;Q10. What are the limitations and caveats acknowledged by the authors?&lt;/h3&gt;
&lt;p&gt;First, the correlations between luminosity and development are &amp;lsquo;far from perfect&amp;rsquo; — the binary lit/unlit transformation in particular fails to capture the significant continuous variation in assets, education, and public goods across regions that are all formally &amp;rsquo;lit.&amp;rsquo; Second, bottom-coding (under-recording of low-light areas) is acknowledged but not corrected; no existing method addresses it, though the authors note that their corrections nonetheless improve elasticities even in rural African regions with very low light. Third, downgrading VIIRS to DMSP by construction sacrifices some of the VIIRS data quality; the long-difference VIIRS elasticity for Africa (0.4) shrinks to 0.35 in the downgraded series. Fourth, daytime satellite imagery and combinations with nighttime lights (Jean et al. 2016; Yeh et al. 2020; Rossi-Hansberg and Zhang 2025) can better capture local wealth but are often proprietary and not replicable in standard economic research. Fifth, the top-coding correction in Africa is minor because very few pixels approach the DN=63 ceiling (0.98–1.7% of lit pixels in 1992–2012), so the main African improvement comes from blooming; other regions with denser urban cores may benefit more from top-coding correction. Sixth, the cross-sensor inter-calibration step is taken &amp;lsquo;off-the-shelf&amp;rsquo; from Li et al. (2020) and further investigation of sensor calibration is left to future work.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Top coding (DMSP)&lt;/strong&gt;: The truncation of Digital Number values at the 8-bit ceiling of 63 in DMSP-OLS data, caused by sensor calibration for cloud detection. Pixels with DN ≥ 55 also suffer &amp;lsquo;implicit&amp;rsquo; top coding because they represent averages of multiple potentially top-coded sub-readings. The paper corrects this by replacing top-coded pixels with structural values drawn from a truncated Pareto distribution, using the radiance-calibrated DMSP vintage to rank pixels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Blooming (spatial spillover of light)&lt;/strong&gt;: A measurement artifact in DMSP data whereby light from bright pixels spills into neighboring dark areas due to the sensor&amp;rsquo;s imprecise spatial accuracy and possible displacement of up to 3 km. The paper identifies pseudo-light pixels (lit pixels adjacent to at least one dark pixel), models the spillover as an inverse-squared-distance weighted function of neighboring lights, and subtracts the predicted blooming from each lit pixel. This correction raises the global unlit pixel share from 92% to 95% in 1992.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extremely randomized trees (ERT)&lt;/strong&gt;: An ensemble machine-learning method used to downgrade VIIRS luminosity data to the DMSP scale. Unlike standard random forests that find the best split thresholds within a random feature subset, ERT selects split thresholds randomly, reducing variance and improving computational efficiency. The authors train it on pixel statistics (mean, median, min, max) and neighborhood statistics within windows of varying sizes to predict DMSP-like values for 2014 onward from VIIRS readings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Harmonized (adjusted + fused) luminosity series&lt;/strong&gt;: The authors&amp;rsquo; main output: an annual global panel of nighttime lights from 1992 to 2023 that applies inter-sensor calibration, top-coding correction, and blooming correction to DMSP data (1992–2013), then uses the ERT ensemble model to convert post-2013 VIIRS data into DMSP-comparable units, yielding four variants (unadjusted, blooming only, top-coding only, both corrections) merged into a continuous time series at 30-arc-second (~1 km²) resolution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pseudo-light pixels (PLPs)&lt;/strong&gt;: In the blooming correction procedure, PLPs are defined as lit pixels (DN &amp;gt; 0) that have at least one dark neighbor (DN = 0). They are the pixels most likely to contain spurious light from neighboring bright areas. PLP light values are regressed on the inverse-squared-distance weighted sum of surrounding pixels to estimate the blooming decay function.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DHS composite wealth index&lt;/strong&gt;: Used in the validation analysis as a local development proxy: a principal-component aggregation of household characteristics including roof quality and ownership of consumer assets, constructed by the Demographic and Health Surveys program across African countries. The paper standardizes this and other outcomes to mean zero and standard deviation one for cross-outcome coefficient comparisons.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Spatial RDD (regression discontinuity design) using nighttime lights&lt;/strong&gt;: As applied in Michalopoulos and Papaioannou (2014) and referenced throughout, a design that restricts estimation to gridcells within a narrow band (e.g., 50 km) of a political or administrative border to compare otherwise similar areas on opposite sides, using luminosity as the outcome. The paper notes that such fine-resolution, localized comparisons are exactly the setting where measurement error in the unadjusted DMSP series is most consequential and where the adjusted series yields the largest improvement.&lt;/p&gt;</description></item><item><title>Labour Market Power and the Effects of Fiscal Policy</title><link>https://macropaperwarehouse.com/papers/labour-market-power-and-the-effects-of-fiscal-policy/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/labour-market-power-and-the-effects-of-fiscal-policy/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper proposes a novel fiscal transmission channel through which government spending expansions reduce employer monopsony power in the labor market, generating larger fiscal multipliers and stronger distributional consequences than standard models predict.&lt;/p&gt;
&lt;p&gt;Standard New Keynesian models rely on two transmission channels with contested empirical support: a negative wealth effect on labor supply (which moves workers to supply more hours when taxes rise) and countercyclical price markups (which fall in booms, raising labor demand). The evidence on both is ambiguous. This paper introduces a third channel — countercyclical monopsony power — that operates independently of, and interacts with, the other two.&lt;/p&gt;
&lt;p&gt;The theoretical framework is a Two-Agent New Keynesian (TANK) model, extending Cantore and Freund (2021). There are two household types: workers (fraction λ = 0.8), who supply labor and have limited financial market access, and capitalists (fraction 1 − λ = 0.2), who earn profit income. Intermediate-good firms compete monopsonistically in local labor markets, paying wages below the marginal revenue product. The wage markdown μ = η/(η+1), where η is the wage elasticity of labor supply to the individual firm. Workers value both pay and non-pay job characteristics (firm location, culture, flexibility), with heterogeneous idiosyncratic preferences drawn from a type-1 extreme value distribution. This differentiation, following Card et al. (2018), gives firms wage-setting power because they cannot observe individual preferences.&lt;/p&gt;
&lt;p&gt;The key mechanism is that η depends endogenously on workers&amp;rsquo; labor earnings (wt·nt) and their marginal utility of income (uW_c,t): η = θ·uW_c,t·wt·nW_t + 1/φ. When government spending rises, it increases both labor income and — because higher current or future taxes reduce lifetime net income — workers&amp;rsquo; marginal valuation of income. Both forces unambiguously raise η, flattening the firm-level labor supply curve, reducing the marginal cost of labor for firms seeking to attract workers, and driving wages up toward the marginal revenue product. Employment and output rise; profits fall and are redistributed toward workers.&lt;/p&gt;
&lt;p&gt;In the calibrated baseline (steady-state markdown μ = 2/3, i.e., wages at two-thirds of marginal revenue products, calibrated to Yeh et al. 2022), the impact fiscal multiplier is approximately 0.6 under monopsonistic competition compared to slightly less than 0.4 under perfect competition — a difference attributable entirely to the countercyclical-monopsony channel. The wage markdown rises by approximately 0.3 percentage points on impact following a 1% of GDP government spending shock, roughly twice the response observed when the steady-state markdown is 0.9 rather than 0.67.&lt;/p&gt;
&lt;p&gt;The amplification from countercyclical monopsony is strongest when the wealth effect on hours worked is near zero — the baseline calibration consistent with Schmitt-Grohé and Uribe (2012) and Galí et al. (2012). As the wealth elasticity of hours increases, the markdown and output response to spending shocks weaken, because a larger hours response implies a smaller consumption response, which reduces the marginal utility channel. The degree of price stickiness has little effect on the markdown response.&lt;/p&gt;
&lt;p&gt;The channel is amplified when workers bear more of the fiscal burden — either through profit redistribution to workers (amplification rises from approximately 0.25 in the no-redistribution baseline to approximately 0.4 when half of profit income is redistributed to workers) or through regressive taxation. Progressively redistributing the tax burden toward capitalists weakens the countercyclical-monopsony channel, which runs counter to the standard cyclical-inequality channel (Bilbiie 2020) that predicts larger multipliers with progressive taxation.&lt;/p&gt;
&lt;p&gt;The empirical validation uses an expectations-augmented VAR estimated on quarterly U.S. data from 1981Q3 to 2019Q4 (macroeconomic variables) and 2000Q4 to 2019Q4 (monopsony measure). Government spending shocks are identified via recursive ordering (government spending ordered first), controlling for professional forecasters&amp;rsquo; spending growth expectations (following Auerbach-Gorodnichenko 2012), the real interest rate using the Wu-Xia shadow policy rate, and the average tax rate. The inverse monopsony measure — the wage elasticity of worker-firm separations — is estimated by extending Langella and Manning (2021) to quarterly frequency using SIPP microdata, controlling for demographics, industry, occupation, human capital, and time effects via complementary log-log regressions month by month. The VAR impulse responses confirm the model&amp;rsquo;s central prediction: government spending expansions raise the wage elasticity of separations (reducing employer market power), raise labor income, reduce profits, and generate substantial output increases.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-in-the-empirical-var-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy in the empirical VAR and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses a recursive (Cholesky) identification scheme with government spending ordered first, following Blanchard and Perotti (2002). The identifying assumption is that government spending does not respond to economic conditions within the same quarter due to decision and implementation lags. Anticipation effects are addressed by including a fiscal news variable — professional forecasters&amp;rsquo; one-period-ahead spending growth forecast from the Survey of Professional Forecasters — following Auerbach and Gorodnichenko (2012). The innovation in government spending orthogonal to this forecast is taken as the exogenous surprise shock. The real interest rate (Wu-Xia shadow federal funds rate, which captures unconventional monetary policy at the zero lower bound) and the average tax rate are included to control for monetary policy stance and financing mix. A key threat the paper acknowledges concerns the separation elasticity estimates: the monopsony literature recognizes biases from insufficient controls for alternative wage offers, unobserved heterogeneity, and lack of firm-level exogenous wage variation. The authors follow Langella and Manning (2021) in arguing that these biases are roughly constant over time, so changes in the estimated separation elasticity still reflect changes in true monopsony power.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-key-mechanism-through-which-government-spending-reduces-monopsony-power"&gt;Q2. What is the key mechanism through which government spending reduces monopsony power?&lt;/h3&gt;
&lt;p&gt;Two reinforcing forces simultaneously raise the wage elasticity of labor supply to individual firms (η). First, higher government spending raises labor income, which increases the dollar magnitude of pay differences between firms, making workers more responsive to relative pay. Second, higher current or future taxes reduce workers&amp;rsquo; lifetime net income, raising their marginal valuation of income (marginal utility of consumption, uW_c,t). Workers facing a tighter budget place greater relative weight on pay versus non-pay job characteristics, further increasing their responsiveness to firm-level wages. Both effects increase η unambiguously for government spending shocks (unlike productivity shocks, where the two forces can offset each other). Higher η flattens the firm-level labor supply curve, compresses the gap between the marginal cost of labor and the wage, and induces firms to raise wages toward the marginal revenue product. Employment and output rise while profits decline, redistributing income from capitalists to workers.&lt;/p&gt;
&lt;h3 id="q3-how-is-monopsony-modeled-and-why-does-the-paper-use-a-discrete-choice-rather-than-ces-approach"&gt;Q3. How is monopsony modeled, and why does the paper use a discrete choice rather than CES approach?&lt;/h3&gt;
&lt;p&gt;The paper adopts a discrete workplace choice model following Card et al. (2018), where workers draw idiosyncratic preferences over non-pay job characteristics from a type-1 extreme value distribution each period. Firms cannot observe individual preferences and set a posted wage. Standard logit calculations yield the wage elasticity of firm-level labor supply as η = θ·uW_c,t·wt·nW_t + 1/φ, where θ is the inverse importance of non-pay characteristics and 1/φ is the intensive-margin (hours) elasticity. Under CES preferences (used by Berger et al. 2022, Alpanda and Zubairy 2021), the wage markdown is constant in equilibrium — analogous to constant price markups under CES monopolistic competition — which eliminates the time variation in monopsony power that is the paper&amp;rsquo;s central object of study. The discrete choice framework generates endogenous variation in η through the endogenous terms wt·nW_t and uW_c,t. Berger et al. (2022) show that the CES approach is a special case of the discrete choice model under restrictive assumptions about individual hours responses; the paper intentionally avoids those assumptions.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-across-calibrations-is-documented-regarding-the-strength-of-the-monopsony-channel"&gt;Q4. What heterogeneity across calibrations is documented regarding the strength of the monopsony channel?&lt;/h3&gt;
&lt;p&gt;The paper documents several dimensions of heterogeneity: (1) Steady-state markdown: the relationship between the steady-state markdown and the markdown&amp;rsquo;s response to government spending is hump-shaped (inverted U-shape). At the baseline value of 0.67, the markdown rises by approximately 0.3 percentage points; at a steady-state markdown of 0.9, the response is roughly half as large. Perfect competition (markdown = 1) and maximum monopsony (markdown → 0) both imply no response. (2) Wealth effect on labor supply (χ): as χ increases from zero (baseline, near-GHH preferences) to one (strong wealth effect), the markdown response and the output amplification decline monotonically. With a near-zero wealth effect (baseline), amplification relative to the perfect-competition counterfactual is approximately 0.25 percentage points of steady-state GDP; it diminishes substantially as χ rises. (3) Profit redistribution (φd): output amplification rises from approximately 0.25 (no redistribution, baseline) to approximately 0.4 when half of profits are redistributed to workers. (4) Tax progressivity (φτ): the channel is stronger under regressive taxation (more of the burden falling on workers) and weaker under progressive taxation, in contrast to the cyclical-inequality channel. (5) Degree of tax financing (φg): higher contemporaneous tax financing strengthens the channel because it raises workers&amp;rsquo; current marginal valuation of income more directly. (6) Price stickiness (ξ): changing price adjustment costs has little effect on the markdown response and the countercyclical-monopsony amplification.&lt;/p&gt;
&lt;h3 id="q5-how-is-the-separation-elasticity-measured-and-linked-to-the-models-concept-of-monopsony-power"&gt;Q5. How is the separation elasticity measured and linked to the model&amp;rsquo;s concept of monopsony power?&lt;/h3&gt;
&lt;p&gt;The separation elasticity γ is the wage elasticity of worker-firm separations: the percentage change in a firm&amp;rsquo;s separation rate in response to a 1% change in the wage. In the model, γ is shown to be proportional to η − 1/φ (the extensive-margin component of labor supply elasticity to the firm), because firm size and separation rate are linked through a constant elasticity derived from the logit choice structure. Empirically, the paper extends Langella and Manning (2021) to quarterly frequency using SIPP data from 2000Q4 to 2019Q4. Month-by-month complementary log-log regressions of separation dummies on residualized log hourly wages (purged of demographic, industry, occupation, human capital, and time effects) yield time-varying quarterly estimates of γ. A higher γ (less negative, since separations fall with higher wages) indicates lower monopsony power. The VAR incorporates this time-varying series as the inverse monopsony measure.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-countercyclical-monopsony-channel-interact-with-the-wealth-effect-and-price-markup-channels"&gt;Q6. How does the countercyclical-monopsony channel interact with the wealth effect and price markup channels?&lt;/h3&gt;
&lt;p&gt;The three channels interact in both complementary and partially offsetting ways. The wealth effect on hours worked (χ &amp;gt; 0) independently shifts the market labor supply curve rightward when taxes rise, increasing employment. However, a larger hours response implies a smaller consumption response, which reduces the increase in workers&amp;rsquo; marginal utility of consumption. Since uW_c,t is a key driver of η, a stronger wealth effect on hours dampens the countercyclical-monopsony channel. Similarly, the countercyclical price markup channel (ξ &amp;gt; 0) raises the marginal revenue product of labor when government spending pushes up demand, boosting employment through an independent channel that also raises labor income — which in turn reinforces η. Yet changing price stickiness has quantitatively little effect on the markdown response in the calibrated model. Income redistribution between agent types mediates the interaction: when capitalists bear most of the tax burden (progressive taxation), workers&amp;rsquo; marginal utility of income rises less, weakening the monopsony channel. When workers bear the burden (regressive taxation or profit redistribution), the monopsony channel is strengthened.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-distributional-consequences-of-the-countercyclical-monopsony-channel"&gt;Q7. What are the distributional consequences of the countercyclical-monopsony channel?&lt;/h3&gt;
&lt;p&gt;When government spending rises, the reduction in employer market power forces firms to pay wages closer to the marginal revenue product, increasing labor income and decreasing profits. This redistribution from capitalists (profit recipients) to workers operates through the wage markdown declining (i.e., markup rising toward one). Under monopsonistic competition with endogenous employer market power, this redistribution is stronger than under perfect competition, where only the price markup channel operates. The VAR evidence confirms these distributional predictions: government spending shocks reduce corporate profits (after taxes) and raise labor income in U.S. data. In the model, this redistribution also feeds back into the mechanism: workers facing declining after-tax income (or receiving a portion of declining profits) place greater weight on pay in their workplace choices, further eroding employer market power.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-cantore-and-freund-2021-and-the-tank-literature-on-fiscal-multipliers"&gt;Q8. How does this paper relate to Cantore and Freund (2021) and the TANK literature on fiscal multipliers?&lt;/h3&gt;
&lt;p&gt;The paper extends the worker-capitalist TANK model of Cantore and Freund (2021), who introduced capitalists that do not participate in the labor market to avoid the criticism (Broer et al. 2019, 2021) that the Bilbiie (2008, 2020) cyclical-inequality channel relies on countercyclical profit income inducing rich households to supply more labor. The Cantore-Freund framework delivers income redistribution between high-MPC workers and low-MPC capitalists without relying on labor supply responses of the rich. This paper adds monopsonistic competition to that framework, introducing a new form of cyclical variation in inequality through time-varying wage markdowns. The interaction with the Bilbiie cyclical-inequality channel is analyzed formally: in particular, tax progressivity has opposing effects under the two channels — progressive taxation amplifies the Bilbiie effect (redistribution to high-MPC workers) but weakens the monopsony channel (capitalists bear more of the tax burden, reducing workers&amp;rsquo; marginal valuation of income).&lt;/p&gt;
&lt;h3 id="q9-what-robustness-is-discussed-or-implied-regarding-the-empirical-var"&gt;Q9. What robustness is discussed or implied regarding the empirical VAR?&lt;/h3&gt;
&lt;p&gt;The paper addresses robustness primarily through the following design choices: (1) Use of the Wu-Xia shadow federal funds rate rather than the actual federal funds rate, to capture monetary policy stance during the zero lower bound period; (2) inclusion of the spending growth forecast variable to control for anticipation effects; (3) inclusion of the average tax rate as a control for fiscal financing; (4) detrending all VAR variables as deviations from linear trends. The separation elasticity itself is shown to be robustly procyclical across three detrending methods (linear, linear-quadratic, and HP-filter with λ=1600), with R² values of 49.9%, 43.6%, and 17.1%, respectively, and regression slopes of 1.52, 1.40, and 1.51 in each case. The paper notes that standard biases in separation elasticity estimation (from unobserved heterogeneity, inadequate controls for alternative offers, absence of firm-level exogenous wage variation) are likely roughly constant over time, which validates using changes in the estimated elasticity as changes in true monopsony power, following Langella and Manning (2021, p. 2942). The sample for the monopsony series (2000Q4–2019Q4) is shorter than the macro VAR sample (1981Q3–2019Q4) due to data availability.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-analytical-results-from-the-simplified-model"&gt;Q10. What are the analytical results from the simplified model?&lt;/h3&gt;
&lt;p&gt;Under flexible prices (no price markup channel), no wealth effect on hours worked (χ = 0), no financial market access for workers (ψW → ∞), full tax financing, and no profit redistribution, the paper derives closed-form expressions for output, labor income, and profits following a government spending shock. Output and labor earnings respond positively to spending only when θ is finite (workers value both pay and non-pay characteristics, so η is endogenous). When θ = ∞ (workers only care about pay → perfect competition with constant η) or θ = 0 (workers only care about non-pay → constant η again), government spending has zero output effect. The parameter Γ = 0 in both limiting cases. For intermediate θ, Γ &amp;gt; 0, government spending raises output and redistributes income from capitalists to workers. This establishes that the countercyclical-monopsony channel is the sole mechanism at work in the simplified model and that it requires intermediate values of workers&amp;rsquo; preference for non-pay characteristics.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper implies that fiscal multipliers may be larger than standard New Keynesian models predict if labor markets exhibit significant employer monopsony power — calibrated to produce a steady-state wage markdown of 2/3 (wages at two-thirds of marginal revenue products), consistent with empirical estimates for the U.S. The countercyclical-monopsony channel provides expansionary effects of government spending even in models where the wealth effect on labor supply is negligible and price markups do not decline. The distributional consequences of fiscal expansions are also stronger under monopsony: income shifts from profit recipients (capitalists) to wage earners more substantially. Scope conditions include: the channel is weaker with stronger wealth effects on hours worked; it is stronger when government spending is financed through current taxes rather than deficit (more tax financing raises workers&amp;rsquo; marginal valuation of income more sharply); it is stronger under regressive rather than progressive taxation; and it is stronger when profit income is redistributed to workers. Progressivity of taxation affects the monopsony and cyclical-inequality channels in opposing directions, implying that the optimal tax structure from a fiscal multiplier perspective depends on which channel is quantitatively dominant.&lt;/p&gt;
&lt;h3 id="q12-what-prior-empirical-literature-on-cyclical-monopsony-power-does-this-paper-build-on-and-extend"&gt;Q12. What prior empirical literature on cyclical monopsony power does this paper build on and extend?&lt;/h3&gt;
&lt;p&gt;The paper builds on three prior empirical findings. First, substantial employer market power in U.S. labor markets (Berger et al. 2022; Langella and Manning 2021; Yeh et al. 2022). Second, unconditional countercyclicality of employer market power — Hirsch et al. (2018) for Germany, Bassier et al. (2022) for Oregon, and Webber (2022) for the U.S. all document that firms hold more monopsony power in slack labor markets. The paper&amp;rsquo;s own descriptive analysis confirms this procyclicality of the separation elasticity across multiple detrending methods. Third, Langella and Manning (2021) provide the estimation methodology for the separation elasticity using SIPP data. The paper&amp;rsquo;s extension is twofold: (a) it extends the Langella-Manning estimates to quarterly frequency and expands the sample to 2019Q4; and (b) it examines the conditional cyclicality of employer market power — specifically, how monopsony power responds to identified government spending shocks — which prior literature had not done.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Countercyclical monopsony channel&lt;/strong&gt;: The novel fiscal transmission mechanism proposed by the paper: government spending expansions endogenously reduce employer monopsony power by raising both labor income and workers&amp;rsquo; marginal valuation of income, which makes workers more responsive to relative pay differences across firms (higher η), compresses wage markdowns, and raises employment and output. The channel is &amp;lsquo;countercyclical&amp;rsquo; in that employer market power falls as spending rises.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wage markdown (µ)&lt;/strong&gt;: The ratio of the wage paid to workers to the marginal revenue product of labor, defined as µ = η/(η+1), bounded between zero and one. A smaller µ implies a larger wedge between pay and marginal product, i.e., greater monopsony power. Perfect competition corresponds to µ = 1. In the baseline calibration µ = 2/3, meaning wages equal two-thirds of the marginal revenue product.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wage elasticity of labor supply to the individual firm (η)&lt;/strong&gt;: The key measure of firms&amp;rsquo; monopsony power in the model. Defined as η = θ·uW_c,t·wt·nW_t + 1/φ, where 1/φ is the intensive-margin (hours) elasticity. The extensive-margin component θ·uW_c,t·wt·nW_t determines how strongly a firm can attract workers from competitors by raising pay. Higher η means less monopsony power (wages closer to marginal revenue product); lower η means greater power.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Separation elasticity (γ)&lt;/strong&gt;: The empirical proxy for inverse monopsony power: the wage elasticity of worker-firm separations, measuring how steeply a firm&amp;rsquo;s separation rate falls when it pays higher wages. In the model, γ is proportional to the extensive-margin component of η. Estimated from SIPP microdata via month-by-month complementary log-log regressions of separation dummies on residualized log wages, following Langella and Manning (2021).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;New classical (idiosyncrasy) monopsony&lt;/strong&gt;: The modeling approach used in the paper, following Card et al. (2018), in which monopsony power arises from workers&amp;rsquo; heterogeneous preferences over non-pay job characteristics (location, culture, flexibility) rather than from search frictions or geographic isolation. Firms differ in non-pay attributes, and because firms cannot observe individual preferences, they have wage-setting power even with frictionless worker flows between firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cyclical-inequality channel&lt;/strong&gt;: A fiscal transmission mechanism from the HANK/TANK literature (Bilbiie 2008, 2020): government spending redistributes income from low-MPC capitalists to high-MPC workers, amplifying the fiscal multiplier. The paper shows this channel interacts with the countercyclical-monopsony channel in conflicting ways — progressive taxation strengthens the cyclical-inequality channel but weakens the monopsony channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wealth effect on labor supply (χ)&lt;/strong&gt;: Parameterized via the Jaimovich-Rebelo (2009) utility function, χ governs how strongly a decline in household lifetime income (due to higher taxes) induces workers to supply more hours. The baseline calibration sets χ → 0, consistent with near-GHH preferences and estimates in Schmitt-Grohé and Uribe (2012). A higher χ dampens the countercyclical-monopsony channel by reducing the consumption response and thereby the marginal utility response.&lt;/p&gt;</description></item><item><title>Macroeconomic Effects of Public R&amp;D</title><link>https://macropaperwarehouse.com/papers/macroeconomic-effects-of-public-rd/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/macroeconomic-effects-of-public-rd/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper estimates the dynamic macroeconomic effects of US government R&amp;amp;D investment using a Structural Vector Autoregressive (SVAR) framework, with an extension to a Rational Expectations SVAR (RE-SVAR) that explicitly captures private-sector anticipation of public spending decisions. The central questions are: (1) what is the fiscal multiplier of public R&amp;amp;D spending on GDP and private R&amp;amp;D investment, and how does it compare to other government spending categories; (2) does public R&amp;amp;D crowd in or crowd out private R&amp;amp;D; and (3) how much does the private sector&amp;rsquo;s anticipation of future public R&amp;amp;D commitments amplify these effects?&lt;/p&gt;
&lt;p&gt;The dataset covers 1947Q1–2017Q3 and is drawn from the US Bureau of Economic Analysis, deflated to 2009 prices and expressed in per-capita terms. The five-variable system includes government R&amp;amp;D investment (GI), government residual spending (GG), net taxes (T), private R&amp;amp;D investment (GR), and GDP (Y), all modelled in log-levels to preserve cointegrating relationships. The lag length is set to six quarters (chosen by Hannan-Quinn criterion, consistent with the R&amp;amp;D-to-productivity lag literature). Identification rests on three mild contemporaneous restrictions: (i) government R&amp;amp;D decisions are independent of current-quarter GDP, consistent with their long-term, mission-oriented character; (ii) R&amp;amp;D spending can influence all other government expenditures in the same quarter but not vice versa; (iii) taxes affect government spending contemporaneously but not the reverse. An alternative identification (SVAR model B) reverses the within-quarter tax-spending causality and produces very similar results. The RE-SVAR extends the system by including the expected next-period public R&amp;amp;D shock, identified by assuming perfect foresight of one-quarter-ahead government R&amp;amp;D innovations and an additional restriction that public R&amp;amp;D does not respond to lagged GDP or private R&amp;amp;D.&lt;/p&gt;
&lt;p&gt;Main quantitative findings from the leading estimation (RE-SVAR model A, full sample):&lt;/p&gt;
&lt;p&gt;GDP fiscal multiplier — anticipated shock: within the quarter of implementation (one quarter after the announcement), one dollar of public R&amp;amp;D spending raises GDP by approximately 52 dollars (pure multiplier at t = 0 is 51.59; see Table 2). The multiplier peaks immediately and then declines to roughly 22–24 dollars over a six-year horizon. Critically, this GDP increase is permanent across all SVAR and RE-SVAR specifications, whereas generic government spending produces only a temporary rise.&lt;/p&gt;
&lt;p&gt;GDP fiscal multiplier — unanticipated shock: setting aside the anticipation effect, the impact-period multiplier falls to approximately 13–14 dollars (13 dollars in the scenario with no anticipation), which is still substantially larger than the peak multiplier of roughly 0.73–0.76 dollars for residual government spending (Table 1, SVAR model A).&lt;/p&gt;
&lt;p&gt;Expectations channel: at t = 0, before the actual spending increase occurs at t = 1, the news alone raises GDP by 16.48 dollars. The total peak GDP effect (55.75 dollars) is nearly double the counterfactual effect without the anticipation component (31.64 dollars). The coefficient on expected next-period public R&amp;amp;D in the private R&amp;amp;D equation is 0.58 (p-value 0.035), confirming a statistically significant anticipation channel for private R&amp;amp;D.&lt;/p&gt;
&lt;p&gt;Crowding-in of private R&amp;amp;D: public R&amp;amp;D crowds in private R&amp;amp;D at all horizons. The public-to-private R&amp;amp;D multiplier peaks at 1.81 in the quarter following the news shock (t = 0), and stabilizes at 0.75 after six years — an elasticity of 0.72, close to Moretti et al.&amp;rsquo;s (2021) estimate of 0.52 from production-function methods. At t = 0, private R&amp;amp;D rises by 0.52 in response to the announcement alone.&lt;/p&gt;
&lt;p&gt;Persistence of public spending: a one-dollar public R&amp;amp;D shock keeps GI above 2 dollars six years later, whereas residual government spending returns to baseline within four years. Cumulative total government spending over six years following a one-dollar R&amp;amp;D shock is 220 dollars, versus only 22 dollars for a generic spending increase.&lt;/p&gt;
&lt;p&gt;Output elasticity at longer horizons: the GDP multiplier expressed in elasticity terms is 0.34 one year after the anticipated shock, stabilizing between 0.23 and 0.25 over three to six years. The corresponding range for private R&amp;amp;D (GR shock) is 0.18 to 0.16, broadly consistent with cross-country evidence from Coe-Helpman (1995) and Guellec-van Pottelsberghe (2004).&lt;/p&gt;
&lt;p&gt;The paper argues that the large short-run multipliers reflect three mechanisms that can materialize quickly: (1) process-innovation cost reductions; (2) early entry of private co-investors seeking first-mover advantage; (3) embodiment of new knowledge in physical capital. At longer horizons, supply-side productivity gains and knowledge spillovers dominate. The policy conclusion is that public R&amp;amp;D is unusually effective both as a demand-side stimulus and as a long-run growth instrument, provided government credibly announces and maintains multi-year funding commitments that stabilize private-sector expectations.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The baseline SVAR identification (model A) imposes three contemporaneous exclusion restrictions: government R&amp;amp;D decisions are exogenous to same-quarter GDP and to other fiscal variables (because R&amp;amp;D budgets reflect long-term strategic priorities, not countercyclical reactions); GI can influence GG contemporaneously but not vice versa; and taxes affect spending in the same quarter but not the reverse. A key threat is non-fundamentalness: because public R&amp;amp;D programs are announced well in advance, what appears to the econometrician as a surprise shock is actually largely anticipated by the private sector, biasing the SVAR impulse responses. The paper addresses this by extending the SVAR to a Rational Expectations SVAR (RE-SVAR) that adds the expected next-period GI shock to the information set of private agents, identified by the additional assumption that GI does not respond to lagged GDP or private R&amp;amp;D. A secondary threat is the direction of same-period causality between taxes and spending; an alternative model (SVAR model B) reverses this and finds only minor quantitative differences. The Lucas Critique applies to the counterfactual simulation of an unanticipated shock since the model was estimated under a perfect-foresight assumption.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-re-svar-separate-the-anticipation-effect-from-the-effect-of-the-actual-spending-increase"&gt;Q2. How does the RE-SVAR separate the anticipation effect from the effect of the actual spending increase?&lt;/h3&gt;
&lt;p&gt;The RE-SVAR model includes E[GI_{t+1} | Omega_t] — the expectation of next-period public R&amp;amp;D — as a forward-looking right-hand-side variable in the private R&amp;amp;D and GDP equations. Under the perfect-foresight assumption, this expectation equals the realized next-period structural shock. The IRF for an anticipated GI shock therefore starts at t = 0 when the news arrives and the actual spending rise occurs at t = 1. By comparing (i) the full anticipated IRF (news at t = 0 + realization at t = 1) to (ii) a modified version where the news term is removed from the information set (unanticipated shock), the paper isolates the incremental contribution of expectations. At t = 0 the news alone raises GDP by 16.48 and private R&amp;amp;D by 0.52; the total peak GDP effect with anticipation is 55.75, versus 31.64 without it — a difference of roughly 24 dollars at the one-year horizon.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-main-mechanisms-proposed-to-explain-the-unusually-large-short-run-fiscal-multiplier"&gt;Q3. What are the main mechanisms proposed to explain the unusually large short-run fiscal multiplier?&lt;/h3&gt;
&lt;p&gt;Three channels are proposed for the large immediate GDP response. First, process innovation can reduce production costs without long lags from the start of R&amp;amp;D investment. Second, anticipatory entry of private co-investors seeking first-mover advantages intensifies investment at the very beginning of a research program, even before results are commercialized. Third, innovation embodied in new physical capital means R&amp;amp;D expenditure is accompanied by complementary investment in physical equipment, amplifying the aggregate demand stimulus. At longer horizons, supply-side productivity gains from knowledge spillovers across firms and sectors become the dominant channel. The paper also notes that public R&amp;amp;D programs are frequently accompanied by large-scale complementary government procurement (e.g., defense agency procurements), further magnifying the total mobilization of public resources.&lt;/p&gt;
&lt;h3 id="q4-what-do-the-multipliers-for-residual-government-spending-gg-look-like-and-how-do-they-compare-to-public-rd"&gt;Q4. What do the multipliers for residual government spending (GG) look like, and how do they compare to public R&amp;amp;D?&lt;/h3&gt;
&lt;p&gt;From SVAR model A (Table 1), one dollar of residual government spending raises GDP by 0.73 at t = 0 (also its peak), declining to around 0.45 after six years. The peak private R&amp;amp;D multiplier of GG spending is 0.08 (after six years), rising very slowly from near zero. Compared to the GDP multiplier of public R&amp;amp;D (13.68 at t = 0, peak 16.18), the residual spending multiplier is roughly 20 times smaller. Moreover, the GDP increase from GG spending is temporary, reverting to baseline within four years, while the GDP increase from GI spending is permanent. These contrasts hold across both SVAR models A and B and across the RE-SVAR estimations.&lt;/p&gt;
&lt;h3 id="q5-what-evidence-is-there-for-the-crowding-in-of-private-rd-by-public-rd"&gt;Q5. What evidence is there for the crowding-in of private R&amp;amp;D by public R&amp;amp;D?&lt;/h3&gt;
&lt;p&gt;The paper finds strong, statistically significant crowding-in across all specifications. In the SVAR model A (Table 1), the multiplier of GI on private R&amp;amp;D (GR) reaches its peak of 0.76 after two quarters and remains at 0.41 after six years. In the RE-SVAR model A (Table 2), the anticipated public R&amp;amp;D shock raises private R&amp;amp;D by 1.81 dollars per dollar of public R&amp;amp;D at t = 0, declining to 0.75 after six years, translating to an elasticity of 0.72. Even in the alternative identification (RE-SVAR model B), the result persists, though the peak private R&amp;amp;D multiplier from anticipated GI spending is lower (0.40 after four quarters). The response of private R&amp;amp;D to both its own shock and to public R&amp;amp;D shocks is permanent across all RE-SVAR estimations, supporting the conclusion that public R&amp;amp;D accelerates the total national innovation effort rather than displacing it.&lt;/p&gt;
&lt;h3 id="q6-what-mechanisms-explain-the-crowding-in-of-private-rd"&gt;Q6. What mechanisms explain the crowding-in of private R&amp;amp;D?&lt;/h3&gt;
&lt;p&gt;The paper identifies five complementary channels: (1) Public funding covers large fixed costs (laboratories, human capital), making private research projects profitable that would not otherwise be undertaken. (2) Public R&amp;amp;D removes credit constraints faced by private innovators. (3) Anticipated technological spillovers signal profitable investment opportunities to private firms. (4) The government funding decision itself conveys a signal about the long-run profitability and viability of a research area. (5) The public-private partnership alleviates asymmetric information and the high riskiness that typically deters private R&amp;amp;D. Additionally, transparency in public procurement and entry requirements into publicly funded programs may signal quality, further encouraging private investment.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-conducted-and-what-do-they-show"&gt;Q7. What robustness checks are conducted, and what do they show?&lt;/h3&gt;
&lt;p&gt;Three robustness checks are applied to both the SVAR and RE-SVAR estimations: (i) alternative identification (SVAR model B / RE-SVAR model B) where the contemporaneous causal direction between taxes and government spending is reversed; (ii) a shorter sample excluding the period from the 2008 financial crisis onward (1947Q1–2007Q4); (iii) a longer lag length of eight quarters. For check (i), results are very similar: the GDP multiplier for GI is slightly smaller at short horizons (10.02 vs 13.68 at t = 0 in the SVAR, and 31.19 vs 51.59 at t = 0 in the anticipated RE-SVAR) but converges to similar long-horizon values. For check (ii), the impact of GI on GDP at t = 0 is 15.5 (vs 13.54), with similar hump shape; GI&amp;rsquo;s impact on GR is slightly lower. For the RE-SVAR robustness checks, the paper reports that the shape, timing, and order of magnitude remain stable, as does the finding that the anticipated GI multiplier considerably exceeds the unanticipated one. The general conclusion is no qualitative variation and only minor quantitative differences.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-re-svars-handling-of-the-non-fundamentalness-problem-and-how-is-it-justified-specifically-for-public-rd"&gt;Q8. What is the RE-SVAR&amp;rsquo;s handling of the non-fundamentalness problem and how is it justified specifically for public R&amp;amp;D?&lt;/h3&gt;
&lt;p&gt;Non-fundamentalness arises when the VAR&amp;rsquo;s implied information set is smaller than that of private agents — i.e., what the econometrician calls a surprise is actually anticipated by the economy, so estimated structural shocks are combinations of current and future structural innovations and the fundamental VAR representation is not identified. The paper argues this problem is particularly severe for public R&amp;amp;D because: (1) R&amp;amp;D budgets are part of long-term plans with detailed technical reports and high-profile public announcements (as documented with historical episodes in Section 2); (2) established procurement links between government agencies and private firms provide early information flows. The RE-SVAR addresses this by explicitly adding E[GI_{t+1} | Omega_t] to the system (Blanchard-Perotti approach applied to a non-causal VAR) and assuming perfect foresight of next-period GI innovations. External forecast measures are unavailable for government R&amp;amp;D spending, making this the only viable route. Perfect foresight is defended as particularly appropriate given the highly public, plan-driven nature of government R&amp;amp;D decisions.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q9. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The closest precursors are Deleidi and Mazzucato (2021) and Antolin-Diaz and Surico (2022). Deleidi and Mazzucato use a recursively identified SVAR where defense R&amp;amp;D spending is ordered first and find a first-quarter GDP multiplier of 24 dollars. This paper differs by: (a) using total government R&amp;amp;D (defense + non-defense) rather than only defense R&amp;amp;D; (b) providing a more general and explicitly motivated identification that goes beyond simple recursive ordering; (c) developing the RE-SVAR extension to capture the anticipation channel, which raises the estimated multiplier substantially above 24 dollars. Antolin-Diaz and Surico (2022) study military spending news with a 125-year VAR (60 lags, Bayesian shrinkage) and find a long-run defense spending GDP multiplier of 2.08 and argue that public R&amp;amp;D specifically drives long-run productivity. The present paper uses a shorter but richer five-variable quarterly system with explicit crowding-in measurement. On the crowding-in question, the paper contrasts with earlier work (Goolsbee 1998, Wallsten 2000) finding crowding-out due to inelastic supply of scientists, and aligns with more recent evidence (Becker 2015, Moretti et al. 2021) showing crowding-in once a broader set of mechanisms is accounted for.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Three core policy implications are identified. First, public R&amp;amp;D is a highly effective instrument for stimulating long-run technological innovation and economic growth: the permanent GDP response and the strong private R&amp;amp;D crowding-in indicate that public investment substantially elevates the country&amp;rsquo;s aggregate innovation capacity. Second, fiscal multipliers are class-specific: the multiplier for public R&amp;amp;D dramatically exceeds that for generic government spending, implying that the composition of government expenditure matters greatly for both short-run stabilization and long-run growth. The absence of crowding-out and the large short-run multipliers suggest substantial untapped productive capacity due to market failures in R&amp;amp;D. Third, the anticipation channel is quantitatively important: ignoring private-sector foresight understates the true multiplier, and this implies that the credibility and advance communication of government R&amp;amp;D commitments are themselves policy instruments — long-term, publicly announced programs that stabilize expectations can effectively mobilize private co-investment that would not occur under uncertain or ad hoc spending. Scope conditions: results are estimated on US data 1947Q1–2017Q3, a country with large and heterogeneous federal R&amp;amp;D programs; extrapolation to countries with different institutional settings, R&amp;amp;D compositions, or capital market structures requires caution. The model uses a 1.5-year lag structure that may not fully capture very long-run R&amp;amp;D-to-productivity channels estimated at 5–20 years in micro studies.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-pure-fiscal-multiplier-and-why-does-the-paper-use-it-instead-of-the-standard-multiplier"&gt;Q11. What is the &amp;lsquo;pure fiscal multiplier&amp;rsquo; and why does the paper use it instead of the standard multiplier?&lt;/h3&gt;
&lt;p&gt;Standard fiscal multipliers are calculated by dividing the cumulative IRF of GDP to a unit shock in a given spending category by the cumulative IRF of total government spending to the same shock. The problem is that total spending includes other categories that dynamically respond to the initial shock (e.g., GI shocks cause GG to rise significantly via cross-equation dynamics), so the denominator conflates the effect of GI with the effect of induced GG changes, making multipliers across spending categories incomparable. The paper therefore uses &amp;lsquo;pure multipliers&amp;rsquo; (following Perotti 2004): the counterfactual total government spending is calculated from a version of the SVAR where the dynamics of GG are switched off (all coefficients in the GG equation are set to zero), so the denominator captures only the direct mechanical effect of the GI shock on aggregate spending without the induced cross-spending effects. This allows clean apples-to-apples comparison of one average dollar spent across different categories.&lt;/p&gt;
&lt;h3 id="q12-what-do-long-run-gdp-elasticities-imply-about-the-social-return-to-rd"&gt;Q12. What do long-run GDP elasticities imply about the social return to R&amp;amp;D?&lt;/h3&gt;
&lt;p&gt;Expressed in elasticity terms, the GDP multiplier from an anticipated GI shock is 0.34 one year after implementation and stabilizes at 0.23–0.25 over three to six years. For private R&amp;amp;D (GR shock), the corresponding elasticity is 0.18 after one year, stabilizing at 0.15–0.16. These are broadly consistent with existing cross-country production function estimates: Coe and Helpman (1995) obtain 0.22 for G7 economies; Guellec and van Pottelsberghe (2004) find 0.13 for private and 0.17 for public R&amp;amp;D spending; Ornaghi (2006) finds 0.24 for Spanish firms including spillovers. The paper notes that Jones and Summers (2020) calculate that the social return to innovation can easily generate a GDP effect of 20 dollars per dollar of R&amp;amp;D once the full set of spillovers is captured at the aggregate level, which is consistent with the dollar multipliers obtained here at longer horizons.&lt;/p&gt;
&lt;h3 id="q13-how-does-private-rd-gr-compare-to-public-rd-gi-as-a-gdp-stimulus"&gt;Q13. How does private R&amp;amp;D (GR) compare to public R&amp;amp;D (GI) as a GDP stimulus?&lt;/h3&gt;
&lt;p&gt;In the leading RE-SVAR model A, a unit shock to private R&amp;amp;D raises GDP by 27.65 at t = 0 and reaches a peak of 39.62 after one year, before stabilizing at around 24 dollars after six years. This is slightly below the public R&amp;amp;D effect (peak 55.75 at t = 0, declining to ~38 dollars and eventually ~22 after six years). The short-run superiority of public R&amp;amp;D over private R&amp;amp;D is attributed to: (1) breadth of goals — public programs simultaneously mobilize a wider set of industries; (2) longer planning horizon — reducing uncertainty and encouraging private co-investment; (3) the expectations channel available to public but not private R&amp;amp;D; (4) entry requirements and transparency signaling research quality; (5) government agencies as both funder and user, accelerating knowledge transfer. However, the superiority of public over private R&amp;amp;D is not confirmed in all specifications of the robustness analysis.&lt;/p&gt;
&lt;h3 id="q14-what-historical-evidence-does-the-paper-marshal-to-motivate-the-anticipation-mechanism"&gt;Q14. What historical evidence does the paper marshal to motivate the anticipation mechanism?&lt;/h3&gt;
&lt;p&gt;Section 2 documents several large defense and non-defense R&amp;amp;D programs where public announcements substantially pre-dated actual spending: the Sputnik response (DARPA and NASA created in 1958 following October 1957 Sputnik launch; spending projections published in Business Week months in advance); Nixon&amp;rsquo;s Strategic Nuclear Doctrine (January–February 1974 announcements of record defense budget of 92.6 billion, with Congress extending Pentagon research commitments in June 1975); Reagan&amp;rsquo;s Strategic Defense Initiative (publicly announced March 23, 1983; CBO published detailed multi-year cost projections by May 1984); Kennedy&amp;rsquo;s Moon Mission (announced May 25, 1961; NYT reported cost projections the following day; estimates revised multiple times through 1969); Nixon&amp;rsquo;s War on Cancer (December 1970 Senate report and May 1971 Nixon speech; National Cancer Act passed December 23, 1971 with pre-specified multi-year budget); Human Genome Initiative (DOE announcement March 1986; Department of Health endorsement April 1987; project ran 1990–2013); Obama&amp;rsquo;s Climate Action Plan (energy transition plans mooted from 2009; America COMPETES Acts 2007, 2010, 2014). These examples document both the forward-looking nature of R&amp;amp;D budgeting and the detailed public information available to private agents ahead of actual spending.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Rational Expectations SVAR (RE-SVAR)&lt;/strong&gt;: An extension of the standard SVAR framework that adds a forward-looking expectational variable — specifically the expected next-period public R&amp;amp;D structural shock E[GI_{t+1} | Omega_t] — to the system, allowing the model to capture the influence of private-sector anticipation on current economic outcomes rather than treating all fiscal shocks as surprises.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-fundamentalness&lt;/strong&gt;: A condition arising when the VAR&amp;rsquo;s implied information set is a strict subset of the actual information set of private agents, causing the reduced-form VAR residuals to be non-invertible linear combinations of current and future structural innovations. For public R&amp;amp;D, this means that what the econometrician identifies as a surprise shock to GI is in fact largely anticipated by the private sector, biasing estimated impulse responses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pure fiscal multiplier&lt;/strong&gt;: A class-specific fiscal multiplier calculated by isolating the GDP response to one dollar spent in a given category of government spending while holding other spending categories constant (switching off their dynamics). Contrasts with the standard multiplier, which conflates the direct effect of the shock with induced changes in other spending categories triggered by dynamic cross-equation correlations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mission-oriented spending&lt;/strong&gt;: Government R&amp;amp;D investment directed at achieving long-term strategic national goals (e.g., space exploration, defense superiority, cancer research, climate transition). Defined by three features that distinguish it from generic government expenditure: (i) long-term policy motivation independent of short-run macroeconomic conditions; (ii) advance public announcements that create private-sector expectations; (iii) potential for permanent productivity-level effects through knowledge spillovers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crowding-in&lt;/strong&gt;: In this paper, the phenomenon whereby an exogenous increase in public R&amp;amp;D investment triggers a statistically significant and persistent increase in private R&amp;amp;D investment — the opposite of the crowding-out (substitution) effect posited when an inelastic supply of scientists and engineers constrains total R&amp;amp;D activity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fiscal foresight&lt;/strong&gt;: The ability of private economic agents to predict future government spending decisions ahead of their actual implementation, arising from legislative lags, public announcements, procurement contracts, and established information channels between policy makers and private co-investors. Fiscal foresight makes standard SVAR fiscal shocks non-fundamental and amplifies the macroeconomic impact of spending by triggering anticipatory private responses before the actual dollar is spent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Anticipation channel (expectations effect)&lt;/strong&gt;: The component of the macroeconomic response to public R&amp;amp;D spending that is activated at the time of the public announcement rather than at the time of actual spending. In the RE-SVAR model, this channel accounts for the extra GDP boost of approximately 21 dollars at t = 1 and a peak of 24 dollars after one year, relative to the counterfactual scenario of an unanticipated shock.&lt;/p&gt;</description></item><item><title>Monetary financing produces neither high inflation nor miraculous fiscal multipliers</title><link>https://macropaperwarehouse.com/papers/monetary-financing-produces-neither-high-inflation-nor-miraculous-fiscal-multipliers/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/monetary-financing-produces-neither-high-inflation-nor-miraculous-fiscal-multipliers/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;When central banks pay interest on reserves — as the Federal Reserve has done since October 2008 and as is standard operating procedure today — does financing fiscal stimulus by permanently expanding the central bank&amp;rsquo;s balance sheet produce higher output than debt-financed stimulus? Van der Kwaak (2024) argues the answer is no in most model configurations, and only modestly yes in a specific extension.&lt;/p&gt;
&lt;p&gt;The motivation is practical: with government debt at high levels in many advanced economies, the private sector may be unable or unwilling to absorb additional bonds needed to fund fiscal stimuli. One alternative is monetary financing — the central bank permanently purchases the extra bonds issued to fund the stimulus (as proposed by Gali 2020b for COVID-era policy). A prior key paper (Gali 2020a) found money-financed stimuli to be substantially more effective than debt-financed ones, but that result was derived in a model where the central bank does not pay interest on reserves, so the policy rate becomes endogenous under money financing. Van der Kwaak shows this assumption is at odds with how modern central banks operate: post-GFC balance sheet expansions by the Federal Reserve and ECB have been financed almost entirely by interest-bearing reserves, with non-interest-paying currency showing no meaningful deviation from trend.&lt;/p&gt;
&lt;p&gt;The paper employs a New Keynesian DSGE model with labor as the sole production factor, a central bank that holds government bonds funded by non-interest-paying money and interest-paying reserves (with the composition endogenous), financial intermediaries subject to a Gertler-Kiyotaki (2010) / Gertler-Karadi (2011) incentive-compatibility leverage constraint on bond holdings, and a standard active Taylor rule bounded by the ZLB. Fiscal stimulus takes the form of either (i) a lump-sum tax cut or (ii) an increase in government spending, each equal to 1% of steady-state output. Money financing is modeled as the central bank acquiring the additionally issued bonds and retaining them permanently in nominal terms.&lt;/p&gt;
&lt;p&gt;The central analytical result (Proposition 1) is a proof of &amp;ldquo;extended Ricardian equivalence&amp;rdquo;: the consolidated government&amp;rsquo;s funding mix among money, reserves, government bonds, and lump-sum taxes has zero effect on inflation and the equilibrium allocation in the real economy. This holds whether or not the incentive-compatibility constraint of financial intermediaries is binding — that is, even when bonds and reserves are not perfect substitutes and money financing genuinely reduces the government&amp;rsquo;s funding costs. The key mechanism: because the central bank pays interest on reserves, the deposit rate equals the policy rate in equilibrium, and the policy rate is the sole endogenous variable on which households&amp;rsquo; deposit return depends. As a result, household consumption-savings decisions are completely decoupled from the financing mix; inflation and real quantities are pinned down entirely by the standard NK equilibrium conditions plus the Taylor rule. Proposition 2 further shows that net cash flows between households and the government/financial sector ultimately just finance exogenous government expenditures, so changes in bond prices and lump-sum taxes produce no net wealth effects on households.&lt;/p&gt;
&lt;p&gt;This irrelevance result is shown to extend analytically to: (i) the ZLB regime (since the central bank still controls the policy rate under money financing), (ii) any maturity structure of government debt, (iii) the ECB&amp;rsquo;s two-tiered reserve system (where minimum reserves earn zero and excess reserves earn the policy rate), (iv) ex ante sovereign default risk, (v) an alternative leverage constraint form (deposits capped relative to reserves plus a fraction of bonds), and (vi) a model with physical capital when corporate securities are held by unconstrained households.&lt;/p&gt;
&lt;p&gt;The irrelevance breaks only when balance-sheet-constrained financial intermediaries also hold corporate securities financing the physical capital stock (Section 4.2 / Sims-Wu 2021 extension). In that case, central bank bond purchases under money financing compress bond yields, which via the intermediaries&amp;rsquo; portfolio-choice condition also compresses expected returns on corporate securities, stimulating investment. The quantitative difference between money- and debt-financed stimuli, measured by the discounted cumulative fiscal multiplier over 1,000 quarters, is 0.26 — substantially smaller than the 0.50 difference found by Gali (2020a). For the spending stimulus, the debt-financed multiplier is 0.9103 and the money-financed multiplier is 1.1719, giving a money-over-debt advantage of 0.2616. For the tax cut, the debt-financed multiplier is -0.0219 and the money-financed multiplier is 0.2397, again a difference of 0.2616. The smaller advantage relative to Gali (2020a) reflects the fact that in Gali&amp;rsquo;s framework the policy rate is not controlled by the central bank under money financing, so households&amp;rsquo; saving return falls endogenously and consumption expands sharply — an effect that is entirely absent here because the central bank retains full control of the policy rate.&lt;/p&gt;
&lt;p&gt;The policy implication is that proposals to use monetary financing to achieve &amp;ldquo;miraculous&amp;rdquo; multipliers beyond the normal spending multiplier are misguided in modern institutional settings where central banks pay interest on reserves. Money financing avoids increasing private-sector-held debt but does not amplify macroeconomic stimulus relative to conventional debt financing in the baseline case, and offers only a small incremental boost in the more structured extension.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-key-analytical-result-and-what-is-the-formal-proposition-that-establishes-it"&gt;Q1. What is the key analytical result and what is the formal proposition that establishes it?&lt;/h3&gt;
&lt;p&gt;Proposition 1 proves &amp;rsquo;extended Ricardian equivalence&amp;rsquo;: the consolidated government&amp;rsquo;s funding mix among money, reserves, government bonds, and lump-sum taxes has zero impact on inflation and the equilibrium allocation in the real economy. The proof works by exhibiting a self-contained subset of equilibrium conditions — households&amp;rsquo; first-order conditions for consumption, labor, and deposits; the Taylor rule; firms&amp;rsquo; pricing conditions; and market clearing — that uniquely pins down all real quantities and inflation without including any equation governing the government&amp;rsquo;s or central bank&amp;rsquo;s financing mix. Because the deposit rate equals the policy rate in equilibrium (due to reserves not being subject to the incentive-compatibility constraint), households&amp;rsquo; saving return depends only on inflation and real variables, so the funding mix drops out entirely.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-irrelevance-result-hold-even-when-the-incentive-compatibility-constraint-of-financial-intermediaries-is-binding-and-bonds-and-reserves-are-not-perfect-substitutes"&gt;Q2. Why does the irrelevance result hold even when the incentive-compatibility constraint of financial intermediaries is binding and bonds and reserves are NOT perfect substitutes?&lt;/h3&gt;
&lt;p&gt;When the constraint binds, reserves earn a lower return than bonds, so the central bank&amp;rsquo;s bond purchases do increase bond prices and reduce government funding costs — but these price changes generate no net wealth effects on households. Proposition 2 shows formally that all cash flows between households on one side and the government and financial intermediaries on the other ultimately just finance (exogenous) government expenditures on final goods. Changes in bond prices, intermediary dividends, and households&amp;rsquo; bond and deposit returns cancel out in the household budget constraint, so W_t = g_t regardless of the financing mix. The intuition is that the financial sector and government together form a closed circuit relative to households, and because government spending is exogenous, the circuit&amp;rsquo;s net effect on household wealth is always the same.&lt;/p&gt;
&lt;h3 id="q3-how-does-this-result-differ-from-gali-2020a-and-why-is-the-multiplier-advantage-of-money-financing-larger-in-that-paper"&gt;Q3. How does this result differ from Gali (2020a), and why is the multiplier advantage of money financing larger in that paper?&lt;/h3&gt;
&lt;p&gt;Gali (2020a) assumes the monetary base consists solely of non-interest-paying money. In that setting, when the central bank permanently expands the monetary base to finance a fiscal stimulus, it cannot simultaneously control the policy rate and the money supply, so the policy rate becomes endogenous and falls relative to a debt-financed stimulus. This endogenous reduction in the rate at which households can save causes a substantial increase in consumption. In van der Kwaak&amp;rsquo;s framework, the central bank pays interest on reserves and retains full control of the policy rate regardless of whether the stimulus is debt- or money-financed, eliminating this consumption-expansion channel. As a result, Gali finds a money-over-debt multiplier advantage of 0.50, while van der Kwaak finds 0.26 in the one model extension where irrelevance is broken, and zero in the baseline.&lt;/p&gt;
&lt;h3 id="q4-in-what-model-extension-is-the-irrelevance-result-broken-and-what-is-the-mechanism"&gt;Q4. In what model extension is the irrelevance result broken, and what is the mechanism?&lt;/h3&gt;
&lt;p&gt;The irrelevance breaks when balance-sheet-constrained financial intermediaries hold both government bonds and corporate securities (financing the physical capital stock), as in Sims and Wu (2021) and van der Kwaak (2023). In this configuration, the incentive-compatibility constraint links the expected excess returns on bonds and corporate securities through a fixed ratio lambda_b / lambda_k. When money financing causes the central bank to acquire additional bonds, bond prices rise and expected bond returns fall. Via the portfolio-choice optimality condition, this also compresses expected returns on corporate securities, which encourages investment. A direct link thus emerges from the government&amp;rsquo;s financing mix to the real economy through the financial sector&amp;rsquo;s balance sheet. Without this channel — whenever corporate securities are held by unconstrained households, or the model has no physical capital — the irrelevance holds exactly.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-exact-quantitative-multiplier-results-from-the-numerical-exercise"&gt;Q5. What are the exact quantitative multiplier results from the numerical exercise?&lt;/h3&gt;
&lt;p&gt;Using the discounted cumulative multiplier formula summed over 1,000 quarters (Table 2): (i) Debt-financed tax cut: -0.0219. (ii) Money-financed tax cut: 0.2397. Difference: 0.2616. (iii) Debt-financed spending stimulus: 0.9103. (iv) Money-financed spending stimulus: 1.1719. Difference: 0.2616. The money-over-debt advantage is identical (0.2616) for both types of stimulus, though the levels differ substantially. The debt-financed tax-cut multiplier is negative because higher bond issuance generates capital losses on intermediaries&amp;rsquo; bond portfolios, tightening the incentive-compatibility constraint and reducing credit provision and investment. Money financing mitigates these losses by having the unconstrained central bank absorb the newly issued bonds, raising bond prices and net worth.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-does-the-paper-conduct-on-the-irrelevance-result"&gt;Q6. What robustness checks does the paper conduct on the irrelevance result?&lt;/h3&gt;
&lt;p&gt;The paper proves the irrelevance analytically for: (1) Both binding and slack incentive-compatibility constraints (Section 3.1). (2) Any maturity structure of government debt — the maturity parameter rho drops out of the relevant equilibrium conditions (Section 3.2.1). (3) The ZLB — since the central bank still controls the reserve rate even under money financing (Section 3.2.1). (4) An alternative leverage constraint where deposit capacity depends on reserves plus a discounted fraction of bonds rather than a fixed fraction of bond value (Appendix C.2). (5) The ECB&amp;rsquo;s two-tiered reserve system, where minimum reserves receive zero interest and excess reserves receive the policy rate; the deposit rate becomes (1-theta)*policy rate instead of the policy rate itself, but is still solely determined by the policy rate (Proposition 3, Section 3.2.2). (6) Models with physical capital when households hold the corporate securities (Proposition 4, Section 4.1). (7) Ex ante sovereign default risk following Corsetti et al. (2013) (Appendix C.1).&lt;/p&gt;
&lt;h3 id="q7-what-is-extended-ricardian-equivalence-as-defined-by-the-author-and-how-does-it-differ-from-the-original-barro-1974-result"&gt;Q7. What is &amp;rsquo;extended Ricardian equivalence&amp;rsquo; as defined by the author, and how does it differ from the original Barro (1974) result?&lt;/h3&gt;
&lt;p&gt;Barro&amp;rsquo;s (1974) Ricardian equivalence shows that the funding mix between government debt and lump-sum taxes has zero effect on the real economy. Van der Kwaak extends this to include the monetary base — the funding mix among money, reserves, government bonds, and lump-sum taxes has zero impact on inflation and the real equilibrium. This is a strictly more general result because it covers the substitution of money/reserves for bonds (i.e., monetary financing), not just the substitution of debt for taxes. Crucially, the extension holds even when bonds and reserves are not perfect substitutes (when the incentive-compatibility constraint binds), which is the nontrivial part of the contribution.&lt;/p&gt;
&lt;h3 id="q8-how-is-money-financing-modeled-in-the-paper"&gt;Q8. How is &amp;lsquo;money financing&amp;rsquo; modeled in the paper?&lt;/h3&gt;
&lt;p&gt;A money-financed stimulus is modeled as one in which the government bonds newly issued to fund the additional spending or the tax cut are acquired by the central bank and permanently retained on its balance sheet in nominal terms. For a spending stimulus, the parameter kappa_g = 1 means the central bank&amp;rsquo;s nominal assets expand by the amount of each period&amp;rsquo;s additional government purchases (g_t - g_bar). For a tax cut, kappa_tau = 1 means the central bank acquires bonds equal to the tax-cut component tau_tilde_t. Debt financing corresponds to kappa_g = 0 or kappa_tau = 0. The central bank&amp;rsquo;s dividends (profits net of interest on reserves and seigniorage on currency) are returned to the fiscal authority each period, so central bank net worth is zero. The author notes this is consistent with the legal constraints on central banks (Buiter 2014) since it takes the form of permanent QE rather than overt fiscal transfers.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-role-of-the-incentive-compatibility-constraint-in-generating-the-bond-price-spread-and-why-does-the-irrelevance-result-still-hold"&gt;Q9. What is the role of the incentive-compatibility constraint in generating the bond-price spread, and why does the irrelevance result still hold?&lt;/h3&gt;
&lt;p&gt;The Gertler-Kiyotaki constraint limits the volume of government bonds intermediaries can hold relative to their net worth (chi_t * n_t = lambda_b * q^b_t * s^{b,f}_t when binding). When binding, intermediaries cannot freely expand bond holdings in response to higher bond supply, so an increase in bond supply under a debt-financed stimulus depresses bond prices and creates capital losses. Conversely, the unconstrained central bank buying additional bonds under money financing raises bond prices. So the constraint creates a genuine price and funding-cost differential between money- and debt-financed stimuli. Yet the irrelevance still holds because, as shown in Proposition 2, these bond-price changes, together with changes in intermediary dividends, net out from the household budget constraint — the household sees the same net obligation regardless of financing mix.&lt;/p&gt;
&lt;h3 id="q10-how-does-corollary-1-relate-to-the-empirical-observation-about-the-monetary-base-composition"&gt;Q10. How does Corollary 1 relate to the empirical observation about the monetary base composition?&lt;/h3&gt;
&lt;p&gt;Corollary 1 proves analytically that any expansion of the monetary base under money financing consists entirely of an expansion in interest-paying reserves — non-interest-paying money holdings are unchanged. This is because, in equilibrium, households&amp;rsquo; demand for non-interest-paying money depends only on consumption and the nominal deposit rate (via the money-in-utility first-order condition), neither of which changes under money financing (by the irrelevance result). This directly matches the empirical evidence shown in Figures 1 and 4 for the Federal Reserve and ECB respectively: post-GFC balance-sheet expansions were almost entirely in interest-paying reserves, with currency in circulation showing no deviation from trend.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-tax-cut-mechanism-under-debt-financing-in-the-numerical-exercise-and-why-is-the-multiplier-negative"&gt;Q11. What is the tax-cut mechanism under debt financing in the numerical exercise, and why is the multiplier negative?&lt;/h3&gt;
&lt;p&gt;Under a debt-financed tax cut (kappa_tau = 0), the fiscal authority must issue more bonds to offset the revenue shortfall. Because financial intermediaries&amp;rsquo; incentive-compatibility constraint is binding, they cannot perfectly elastically absorb the additional bond supply; bond prices fall, causing capital losses on intermediaries&amp;rsquo; existing holdings. This reduces net worth, tightens the constraint further, and forces intermediaries to reduce lending to the real economy. The capital price and investment therefore fall. The trough in output is at most about 0.03% of steady-state output, but the cumulative multiplier is -0.0219 — negative because the adverse financial amplification from falling bond prices more than offsets any direct effect of the lump-sum transfer on households. This mechanism is similar to van der Kwaak and van Wijnbergen (2017).&lt;/p&gt;
&lt;h3 id="q12-what-is-the-calibration-strategy-and-how-closely-does-it-follow-gali-2020a"&gt;Q12. What is the calibration strategy, and how closely does it follow Gali (2020a)?&lt;/h3&gt;
&lt;p&gt;The calibration of the model with financial intermediaries holding corporate securities follows Gali (2020a) for most household and production parameters: discount factor beta = 0.995, risk aversion sigma_c = 1, inverse Frisch elasticity phi = 5, price semi-elasticity of money demand eta = 7, Calvo probability psi_p = 3/4, elasticity of substitution epsilon = 9, labor share = 0.75, steady-state government debt / output = 2.4 (60% of annual GDP), AR(1) for government spending rho_g = 0.5. Deviations from Gali include: government spending share of output set at g_bar/y_bar = 0.2 (consistent with advanced economy averages), steady-state investment share i_bar/y_bar = 0.2, and a monetary base equal to 1/3 of quarterly output (as in Gali) now split into non-interest-paying money (10% of quarterly output) and interest-paying reserves (1.63 times currency). For financial intermediaries: average banker tenure 24 quarters (sigma = 0.9583), adjusted leverage ratio 5, steady-state spread on corporate securities and bonds over deposits = 25 quarterly basis points (100 annual basis points), implying lambda_b = lambda_k. Capital adjustment cost gamma_k = 2.5.&lt;/p&gt;
&lt;h3 id="q13-how-does-the-paper-relate-to-wallace-1981-and-when-does-the-neutrality-argument-break-down"&gt;Q13. How does the paper relate to Wallace (1981) and when does the neutrality argument break down?&lt;/h3&gt;
&lt;p&gt;Wallace (1981) first showed that open-market operations are neutral in complete-markets models where all investors can purchase any asset at market prices without binding constraints. Woodford (2012) distills the key conditions: assets are valued only for pecuniary returns, and all investors face the same market prices with no binding position constraints. Van der Kwaak&amp;rsquo;s irrelevance extends the Wallace neutrality to incomplete markets with binding leverage constraints on bond holdings, which go beyond Woodford&amp;rsquo;s conditions. The neutrality breaks only when the binding constraint links together multiple asset classes — specifically when the same constraint covers both government bonds and corporate securities, creating a direct transmission from bond prices to the cost of capital.&lt;/p&gt;
&lt;h3 id="q14-how-does-the-paper-relate-to-reis-and-tenreyro-2022-on-helicopter-money"&gt;Q14. How does the paper relate to Reis and Tenreyro (2022) on helicopter money?&lt;/h3&gt;
&lt;p&gt;Reis and Tenreyro (2022) study helicopter drops — direct transfers of newly created central bank liabilities to households — and derive an irrelevance result that applies only when bond and reserve interest rates are equal (perfect substitutes). Van der Kwaak&amp;rsquo;s irrelevance extends to the case where the return on bonds exceeds that on reserves (binding incentive-compatibility constraint). A second difference is that Reis-Tenreyro focus on helicopter money (a liability-side transfer), while van der Kwaak models money financing as permanent QE (an asset-side expansion). Third, van der Kwaak also studies money-financed government spending stimuli, which Reis-Tenreyro do not.&lt;/p&gt;
&lt;h3 id="q15-what-are-the-implications-for-policy-proposals-to-use-monetary-financing-in-high-debt-environments"&gt;Q15. What are the implications for policy proposals to use monetary financing in high-debt environments?&lt;/h3&gt;
&lt;p&gt;The core message for policy is nuanced. On the fiscal side, monetary financing does achieve its main stated goal: it prevents private-sector-held government debt from rising, since the additional bonds are absorbed by the central bank. On the stimulus effectiveness side, however, money financing has no macroeconomic advantage over debt financing in the baseline model (and in most extensions). The one setting where there is an advantage — intermediaries holding both bonds and corporate securities — yields only a modest multiplier boost of 0.26 relative to debt financing, compared to the 0.50 suggested by Gali (2020a). This smaller number reflects the fundamental institutional difference: with interest-on-reserves, the policy rate stays fixed under money financing, eliminating the consumption-expansion channel. The paper also implies there is no inflationary danger from money financing in this setup — the irrelevance result holds for inflation as well as real variables — directly contradicting fears that monetary financing inherently produces high inflation.&lt;/p&gt;
&lt;h3 id="q16-what-happens-to-inflation-under-money-financing-compared-to-debt-financing-in-the-analytical-result"&gt;Q16. What happens to inflation under money financing compared to debt financing in the analytical result?&lt;/h3&gt;
&lt;p&gt;The extended Ricardian equivalence result covers inflation explicitly: the path of inflation is identical under money financing and debt financing in all the analytical baseline cases. This is because inflation is pinned down by the New Keynesian Phillips curve and the Taylor rule, neither of which depends on the financing mix. The central bank retains full control of the policy rate under money financing (because it pays interest on reserves), so the Taylor rule continues to govern inflation dynamics. This directly contradicts the claim that monetary financing is inherently inflationary; in the model, it is neither inflationary nor expansionary relative to debt financing.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Extended Ricardian equivalence&lt;/strong&gt;: The author&amp;rsquo;s label for the proposition that the consolidated government&amp;rsquo;s funding mix among money, reserves, government bonds, and lump-sum taxes has zero effect on both inflation and the equilibrium allocation in the real economy. It extends Barro (1974)&amp;rsquo;s original Ricardian equivalence (which covered only debt vs. taxes) to include the monetary base, and holds even when bonds and reserves are not perfect substitutes due to binding intermediary leverage constraints.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Money-financed fiscal stimulus&lt;/strong&gt;: In this paper&amp;rsquo;s modeling: a fiscal stimulus (tax cut or spending increase) in which the additional government bonds issued to fund it are acquired by the central bank and permanently retained on its balance sheet in nominal terms. This is equivalent to a permanent expansion of the monetary base equal to the size of the stimulus, and is distinct from helicopter drops (which involve direct transfers rather than bond purchases).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Incentive-compatibility constraint (binding case)&lt;/strong&gt;: A Gertler-Kiyotaki (2010) / Gertler-Karadi (2011) constraint limiting financial intermediaries&amp;rsquo; bond holdings relative to net worth: chi_t * n_t = lambda_b * q^b_t * s^{b,f}_t when binding. When binding, it creates a spread between bond and reserve returns, meaning bonds and reserves are not perfect substitutes. The paper&amp;rsquo;s irrelevance result holds whether or not this constraint binds, which is the nontrivial analytical contribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Interest-paying reserves (interest on reserves)&lt;/strong&gt;: Central bank liabilities that pay a nominal interest rate set by the central bank, distinct from non-interest-paying currency (&amp;lsquo;outside money&amp;rsquo;). The paper argues this is the empirically relevant form of modern monetary base expansion: post-GFC balance-sheet growth by the Fed and ECB was almost entirely in interest-paying reserves. Paying interest on reserves allows the central bank to simultaneously control the policy rate and the size of its balance sheet, which is the feature that drives the irrelevance result.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cumulative (discounted) fiscal multiplier&lt;/strong&gt;: As computed in the paper following Gali (2020a): the ratio of the sum of output deviations from steady state over 1,000 quarters to the sum of the fiscal instrument deviations over the same horizon. The relevant multiplier here is the difference between money- and debt-financed versions: 0.26 in the extension with corporate securities held by intermediaries, compared to 0.50 in Gali (2020a).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Two-tiered reserve system&lt;/strong&gt;: The ECB framework (in operation since July 2023) under which intermediaries must hold minimum reserves equal to a fixed fraction of deposits (currently 1%) at zero interest, while excess reserves earn the policy rate. The paper proves (Proposition 3) that extended Ricardian equivalence carries over to this system: the nominal deposit rate becomes (1-theta)*policy rate, but since the policy rate remains the sole endogenous variable determining the deposit rate, the irrelevance result is unaffected.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Source-text-origin note&lt;/strong&gt;: The working paper title reads &amp;lsquo;Monetary financing does not produce miraculous fiscal multipliers&amp;rsquo;; the published EJ title adds &amp;rsquo;neither high inflation nor&amp;rsquo; — the summary uses the published title as given in the task, which also reflects the paper&amp;rsquo;s second finding (no inflationary effect).&lt;/p&gt;</description></item><item><title>Nonlinear Monetary Policy Tradeoffs</title><link>https://macropaperwarehouse.com/papers/nonlinear-monetary-policy-tradeoffs/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/nonlinear-monetary-policy-tradeoffs/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper measures how the inflation-unemployment tradeoff associated with monetary policy varies with both the sign of the monetary intervention (easing versus tightening) and the state of the business cycle (booms versus recessions) for the US economy over 1973:M1 to 2019:M6. The motivation is that standard linear Phillips-curve estimates implicitly impose a constant tradeoff, yet a flat Phillips curve would simultaneously predict that (i) stimulating activity during a recession costs nothing in terms of inflation and (ii) reducing inflation costs very large amounts of unemployment — both empirically extreme predictions that have very different policy implications. The paper challenges both extremes.&lt;/p&gt;
&lt;p&gt;The empirical strategy extends the Proxy-SVAR approach of Mertens-Ravn (2013) and Stock-Watson (2018) to a nonlinear setting. The economy is described by a Vector Moving Average augmented with nonlinear functions of the monetary policy shock — specifically its absolute value (capturing sign dependence) and its interaction with a recession indicator (capturing state dependence). Under a finite-order VARX representation assumption and a linear monetary policy rule assumption, the paper proves (Proposition 1) that even though the underlying VARX is nonlinear, the monetary shock can be recovered as the projection of an external instrument onto residuals of a misspecified linear VAR. Once the shock is recovered, it and its nonlinear functions are used as regressors in a VARX to estimate nonlinear impulse responses. The instrument is the Degasperi-Ricco (2022) extension of Miranda-Agrippino and Ricco (2021), with a baseline span of 1991:M1-2015:M12 extrapolated to the full sample. The VAR contains five variables: the 1-year Treasury bond rate, industrial production growth, the Gilchrist-Zakrajsek excess bond premium, the unemployment rate, and CPI inflation, estimated with 7 lags. The recession indicator equals 1 when average GDP growth over the previous 12 months is negative.&lt;/p&gt;
&lt;p&gt;The monetary policy tradeoff is defined analogously to the fiscal multiplier: the ratio of the cumulative average impulse response of inflation (unemployment) to the cumulative average impulse response of unemployment (inflation) over horizons H. In a nonlinear setting the easing tradeoff and tightening tradeoff are no longer inverses of one another and must be treated separately.&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. For monetary easing during recessions, the inflation cost of reducing unemployment is small and statistically insignificant: point estimates of T+ range from -0.03 to -0.17 (in absolute value) across horizons H = 12 to H = 48 months, with 68% confidence intervals spanning from approximately -5.3 to +2.8 at H = 12 and -3.4 to +2.7 at H = 48. For monetary tightening during booms, the unemployment cost of reducing inflation is moderate and statistically significant: T- estimates range from -0.51 to -0.61 across H = 12 to H = 48, with 68% confidence intervals entirely below zero (e.g., -1.10 to -0.26 at H = 12 and -1.23 to -0.24 at H = 48). In other words, reducing inflation by 1 percentage point during a boom requires raising unemployment by roughly 0.5 to 0.6 percentage points. These results are qualitatively robust to excluding the post-2008 zero-lower-bound period (pre-2009 subsample) and to alternative specifications. By contrast, monetary tightening during recessions implies a very large and unfavorable tradeoff. Easing during booms is extremely inflationary with virtually no real effect.&lt;/p&gt;
&lt;p&gt;A Likelihood Ratio test for the null hypothesis that all nonlinear terms are zero is rejected at the 1% level, confirming the statistical importance of nonlinearities. The null hypothesis of shock invertibility (Assumption A4) is not rejected at the 5% level across all combinations of VAR lags and residual leads tested.&lt;/p&gt;
&lt;p&gt;A simple model with downward nominal wage rigidities — in which the wage floor introduces a kink in the aggregate supply curve — provides a theoretical rationale for the sign- and state-dependent tradeoff: an expansionary shock in a full-employment economy raises inflation with no output effect (the economy sits on the vertical AS segment), while a contractionary shock makes the wage rigidity binding and reduces output with no price effect (the horizontal AS segment). Monte Carlo validation using artificial data generated by the calibrated DSGE model shows that the proposed empirical procedure recovers the theoretical nonlinear impulse responses very accurately.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-assumptions-required"&gt;Q1. What is the identification strategy and what are the main assumptions required?&lt;/h3&gt;
&lt;p&gt;Identification proceeds in two steps. First, the monetary shock is recovered by projecting an external instrument (Degasperi-Ricco 2022) onto the residuals of a standard linear VAR — this is justified by Proposition 1, which shows that even though the VAR is misspecified (it omits the nonlinear terms), the shock can still be recovered as a linear combination of VAR residuals under four assumptions: (A0) a structural VMA representation in which the shock is orthogonal to past observables and to the remaining structural shocks at all leads and lags; (A1) a finite-order VARX representation; (A2) invertibility of the Wold representation; (A3) a valid instrument (relevance and exogeneity); and (A4) informational sufficiency, meaning the monetary shock can be expressed as a linear combination of current and past observables — a condition implied by a linear monetary policy rule. Second, once the estimated shock and its nonlinear functions (absolute value and interaction with the state dummy) are in hand, they are used as exogenous regressors in a VARX to estimate nonlinear impulse response functions.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-threats-to-identification"&gt;Q2. What are the main threats to identification?&lt;/h3&gt;
&lt;p&gt;Three main threats are acknowledged. (1) Instrument validity: if the instrument (Degasperi-Ricco 2022) is weak or contaminated by information shocks, the first-stage projection may recover a mislabeled shock. The authors note the first-stage F-statistic is adequate per Miranda-Agrippino and Ricco (2021) but acknowledge that the weak-instrument problem in the nonlinear context is non-trivial and left for future research. (2) Assumption A4 (informational sufficiency): if the central bank follows a nonlinear rule or the VAR variables are not sufficient to recover the shock, identification fails. The authors test this using the Forni-Gambetti-Ricco (2023) invertibility test — regressing the instrument on current and future VAR residuals and checking whether future residuals matter — and fail to reject invertibility at 5% across all lag/lead combinations. (3) Model misspecification in the nonlinear VARX: the VARX approximation may not capture all relevant nonlinearities generated by the true DSGE. The Monte Carlo validation on artificial DSGE data provides reassurance that the approach recovers the true nonlinear responses accurately.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-distinguish-sign-dependence-from-state-dependence"&gt;Q3. How does the paper distinguish sign dependence from state dependence?&lt;/h3&gt;
&lt;p&gt;The paper includes two nonlinear terms as regressors in the VARX: the absolute value of the shock |u_t^r|, which captures sign-dependent effects (i.e., whether a tightening and an easing of equal magnitude have asymmetric effects), and the product s_{t-1} * u_t^r, which captures state-dependent effects (i.e., whether the same-sign shock has different effects depending on whether the economy was in a recession before the shock arrived). The two components are estimated simultaneously, allowing their separate contributions to be read off impulse responses in Figure 3. Robustness checks in the Online Appendix report models estimated with only sign dependence and only state dependence in isolation, with results described as qualitatively similar to Barnichon-Matthes (2018) and Tenreyro-Thwaites (2016), respectively.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-key-quantitative-results-on-impulse-responses"&gt;Q4. What are the key quantitative results on impulse responses?&lt;/h3&gt;
&lt;p&gt;In the full nonlinear model, monetary tightening generates large and significant effects on real variables (unemployment, industrial production) regardless of the state, while monetary easing has more muted real effects. For prices, sign and state components operate in opposite directions: the largest inflation responses are associated with tightening during expansions. Numerically, the tradeoff estimates from Table 2 show: (a) easing during recessions — T+ point estimates of -0.03 at H=12, -0.12 at H=24, -0.17 at H=36, -0.17 at H=48 months (all statistically insignificant at 68%); (b) tightening during booms — T- point estimates of -0.51 at H=12, -0.61 at H=24, -0.59 at H=36, -0.53 at H=48 months (all statistically significant at 68%). For the pre-2009 subsample (excluding the ZLB period), tightening-in-booms estimates are somewhat larger in absolute value (-0.63 to -0.70) but confidence intervals widen to include zero at longer horizons.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-key-implication-for-pushing-on-a-string-results-in-the-prior-literature"&gt;Q5. What is the key implication for &amp;lsquo;pushing on a string&amp;rsquo; results in the prior literature?&lt;/h3&gt;
&lt;p&gt;Tenreyro-Thwaites (2016) and Barnichon-Matthes (2018) document that monetary easing is less effective at stimulating real activity, especially during recessions — an apparent &amp;lsquo;pushing on a string&amp;rsquo; result. The current paper accepts that the real effect of easing in recessions is muted, but adds a crucial dimension: price responses are also muted in the same circumstances, so the inflation-unemployment tradeoff is actually favorable even when the absolute size of real effects is small. The policy implication is that central banks can still usefully deploy monetary easing during recessions as long as interventions are sufficiently aggressive to achieve the desired stimulus, since the inflationary cost of doing so is low.&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-measure-the-tradeoff-differently-from-phillips-curve-regressions"&gt;Q6. How does this paper measure the tradeoff differently from Phillips-curve regressions?&lt;/h3&gt;
&lt;p&gt;The tradeoff is defined as the ratio of the cumulative average impulse response of inflation to the cumulative average impulse response of unemployment (or vice versa) in response to an identified monetary shock, analogous to a fiscal multiplier. This approach avoids three problems that plague standard Phillips-curve estimates: (i) it does not require specifying a structural Phillips-curve equation, reducing misspecification risk; (ii) it does not require data on inflation expectations or the natural rate of unemployment, which are unobserved and introduce measurement error; (iii) identification comes from exogenous monetary shocks rather than OLS variation in unemployment, so the endogeneity problem is avoided.&lt;/p&gt;
&lt;h3 id="q7-what-theoretical-mechanism-rationalizes-the-nonlinear-tradeoffs"&gt;Q7. What theoretical mechanism rationalizes the nonlinear tradeoffs?&lt;/h3&gt;
&lt;p&gt;A simple New-Keynesian-style model with downward nominal wage rigidities (Wt &amp;gt;= theta * W_{t-1}) generates a kink in the aggregate supply curve. When the economy operates at full employment and inflation is non-negative, an expansionary monetary shock stimulates demand but the wage rigidity is non-binding, so the economy sits on the vertical segment of the AS curve: output cannot exceed its natural level, and the only effect is higher inflation. By contrast, a contractionary shock makes the wage rigidity binding, pushing the economy onto the flat segment of the AS curve: firms cut employment rather than nominal wages, so output falls but prices are unaffected. More generally, averaging over periods of full employment and periods of involuntary unemployment, tightening has larger real effects and weaker price effects than easing — matching the empirical pattern — because a contractionary shock keeps the economy below full employment for a longer time.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-are-conducted"&gt;Q8. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;Three main robustness checks are reported in the main text, each presented with impulse-response figures (Figures 6, 7, 8): (1) replacing the authors&amp;rsquo; state dummy (based on 12-month average GDP growth) with NBER recession dates; (2) replacing the 1-year Treasury bond rate with the Federal Funds rate and with the 6-month Treasury Bill rate; (3) replacing the baseline Degasperi-Ricco instrument with the Jarocinski-Karadi (2020) instrument both raw and cleaned (regressed on six lags of VAR variables). In all cases, the qualitative result — tightening in booms produces larger real effects than easing in recessions, while price responses are more muted in recessions — is preserved, and the tradeoff pattern remains favourable for easing in recessions and tightening in booms. The Online Appendix additionally reports results using: the unemployment rate as the state variable (instead of industrial production); the VAR extended with the 10-year Treasury Bill rate and M2 monetary aggregate; models with only sign dependence; models with only state dependence; and an alternative estimation using the instrument directly in place of the estimated shock (which yields implausible results, validating the two-stage procedure).&lt;/p&gt;
&lt;h3 id="q9-what-does-the-monte-carlo-validation-using-the-dsge-model-establish"&gt;Q9. What does the Monte Carlo validation using the DSGE model establish?&lt;/h3&gt;
&lt;p&gt;The paper generates 1000 artificial realizations from a calibrated downward-nominal-wage-rigidity DSGE model (beta=0.99, sigma=1, theta=1, phi_pi=1.5, rho_m=0.5, sigma_r=0.25%, sigma_a=0.45%, solved by nonlinear global projection using Chebyshev polynomials). It then applies the nonlinear Proxy-SVAR procedure to each artificial dataset and compares average estimated impulse responses with average true (model-generated) generalized impulse responses. The two are described as &amp;lsquo;very similar&amp;rsquo; (Figure 10), demonstrating that the empirical nonlinear VARX representation accurately approximates the nonlinearities of the DSGE even though the VARX is in principle misspecified relative to the true model. This validates both the econometric procedure and the interpretive link between the empirical findings and the theoretical mechanism.&lt;/p&gt;
&lt;h3 id="q10-why-does-the-paper-estimate-the-shock-from-a-misspecified-linear-var-rather-than-the-varx-directly"&gt;Q10. Why does the paper estimate the shock from a misspecified linear VAR rather than the VARX directly?&lt;/h3&gt;
&lt;p&gt;The monetary shock is latent. Proposition 1 shows that, under the stated assumptions, the monetary shock equals (up to a scaling constant) the projection of the external instrument onto the VAR residuals of the linear VAR, even though the VAR omits the nonlinear terms. This is because the linear monetary policy rule implies the shock is a linear combination of current observables, and the VAR residuals span the same space. Using the instrument directly in the VARX instead of going through steps I and II introduces a non-proportional bias in the nonlinear case (unlike the linear case where the attenuation bias from measurement error in the instrument is proportional across units and corrects under normalization). The Online Appendix shows that bypassing the two-stage shock-estimation procedure yields implausible impulse response estimates.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-scope-of-the-empirical-findings-and-what-caveats-apply"&gt;Q11. What is the scope of the empirical findings and what caveats apply?&lt;/h3&gt;
&lt;p&gt;Three scope conditions are explicitly stated. (1) State uncertainty: the tradeoff varies significantly with the state of the economy, so if the central bank is uncertain about current economic conditions, interventions carry considerable risk — a disinflation during what turns out to be a weaker-than-anticipated economy could incur very large unemployment costs. (2) Historical average: estimates reflect the effects of average monetary interventions over 1973-2019 and may not generalize to unusually large, persistent, or unconventional policy actions. (3) Accompanying fiscal policy: the tradeoff could be influenced by fiscal policy measures that accompanied monetary interventions during the sample period. The sample also excludes the post-2019 inflation surge, so inference about that episode is not direct. The identification requires a valid external instrument, whose strength in the nonlinear context is an open question.&lt;/p&gt;
&lt;h3 id="q12-how-does-this-paper-relate-to-barnichon-mesters-2020-2021-and-gali-gambetti-2020"&gt;Q12. How does this paper relate to Barnichon-Mesters (2020, 2021) and Gali-Gambetti (2020)?&lt;/h3&gt;
&lt;p&gt;Barnichon-Mesters (2020, 2021) and Gali-Gambetti (2020) also exploit identified monetary shocks to estimate the conditional inflation-unemployment relationship (the &amp;lsquo;Phillips multiplier&amp;rsquo;) and to investigate whether the Phillips curve slope has changed over time. The main additional contribution of the present paper is to show that the relationship is not only time-varying but specifically sign- and state-dependent, driven by the direction of monetary intervention and the current phase of the business cycle. The sign- and state-dependent tradeoff framework provides a richer characterization that can explain why a flat aggregate Phillips curve is compatible with moderate costs of disinflation and low inflationary costs of stimulus — something a time-varying-slope model alone does not deliver.&lt;/p&gt;
&lt;h3 id="q13-what-does-the-paper-say-about-the-implications-for-disinflation-episodes-like-2022-23"&gt;Q13. What does the paper say about the implications for disinflation episodes like 2022-23?&lt;/h3&gt;
&lt;p&gt;The paper does not directly analyze the 2022-23 episode (the sample ends at 2019:M6 and the paper was written with November 2025 dating for the online appendix). However, the results imply that if the economy is in a boom when disinflation begins — as was broadly the case in 2022 — the unemployment cost of reducing inflation is moderate (roughly 0.5-0.6 percentage points of unemployment per percentage point of inflation at a 24-36 month horizon), substantially less than would be implied by a flat Phillips curve. The authors explicitly note that their results suggest central banks can pursue disinflation without necessarily incurring very large unemployment costs, subject to the caveats about state uncertainty and scale of the intervention.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Monetary policy tradeoff&lt;/strong&gt;: In this paper&amp;rsquo;s usage: the ratio of the cumulative average impulse response of inflation to the cumulative average impulse response of unemployment (for easing) or vice versa (for tightening), in response to an identified monetary shock, averaged over a horizon H. In a linear model easing and tightening tradeoffs are inverses; in the nonlinear model they must be estimated separately. The concept is deliberately defined without assuming a Phillips curve and without requiring inflation expectations or the natural rate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sign dependence&lt;/strong&gt;: The property that a monetary easing and a monetary tightening of equal magnitude have asymmetric effects on inflation and unemployment, not just opposite-signed effects of the same absolute magnitude. Captured in the VARX by including the absolute value of the monetary shock as an exogenous regressor.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;State dependence&lt;/strong&gt;: The property that the effects of a monetary shock of given sign and magnitude differ depending on whether the economy was in a recession or a boom in the period before the shock arrived. Captured in the VARX by including the product of the recession indicator (s_{t-1}) and the monetary shock as an exogenous regressor.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nonlinear Proxy-SVAR&lt;/strong&gt;: The paper&amp;rsquo;s proposed econometric framework: a Vector Moving Average augmented with nonlinear functions of the monetary shock, which admits a VARX representation. Identification extends the standard Proxy-SVAR by showing — via Proposition 1 — that the latent monetary shock can be recovered from the residuals of a misspecified linear VAR, using an external instrument, under a linear monetary policy rule. The estimated shock and its nonlinear functions are then used as exogenous regressors to recover nonlinear impulse response functions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Downward nominal wage rigidity&lt;/strong&gt;: A labor market friction, modeled as the constraint W_t &amp;gt;= theta * W_{t-1}, that creates a kink in the aggregate supply curve. When the constraint binds (during downturns), firms respond to contractionary shocks by cutting employment rather than nominal wages, generating unemployment without deflation. When the constraint is non-binding (during expansions), expansionary shocks raise nominal wages and prices without affecting employment beyond full-employment output. In this paper the rigidity is the key mechanism generating a sign- and state-dependent monetary tradeoff.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Informational sufficiency (Assumption A4)&lt;/strong&gt;: The identifying assumption that the monetary policy shock can be expressed as a linear combination of current and past observable variables — equivalently, that the central bank follows a linear monetary policy rule. This allows the shock to be recovered from the residuals of a standard linear VAR even when the true model is nonlinear. Tested empirically via the Forni-Gambetti-Ricco (2023) invertibility test (checking whether the instrument Granger-causes future VAR residuals); not rejected at the 5% level in the authors&amp;rsquo; data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Generalized Impulse Response Function (GIRF)&lt;/strong&gt;: In this nonlinear context, defined as E(x_{t+h} | u_t^r = u-bar) - E(x_{t+h} | u_t^r = 0) for h = 0, 1, &amp;hellip;, where u-bar is a given shock size. Unlike linear IRFs, GIRFs depend on the sign and magnitude of the shock and on the state of the economy, and are computed by summing the linear response alpha(L)*u-bar and the nonlinear response Phi(L)*g(u_t^r, &amp;hellip;).&lt;/p&gt;</description></item><item><title>Optimal Fiscal Policy in a Climate-Economy Model with Heterogeneous Households</title><link>https://macropaperwarehouse.com/papers/optimal-fiscal-policy-in-a-climate-economy-model-with-heterogeneous-households/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/optimal-fiscal-policy-in-a-climate-economy-model-with-heterogeneous-households/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether inequality and redistributive taxation should make climate policy more or less ambitious, and how optimal carbon taxes interact with optimal income taxes when households differ in productivity, wealth, and energy demand. The motivation is twofold: equity considerations belong at the center of normative climate analysis, and the distributional consequences of environmental policies are increasingly recognized as critical for their political feasibility — as illustrated by the Yellow Vests episode in France. The paper extends Barrage (2020)&amp;rsquo;s representative-agent dynamic climate-Ramsey model to a heterogeneous-agent setting, using the Werning (2007) technique to characterize the Ramsey optimum in terms of aggregate variables. The government maximizes utilitarian social welfare choosing linear taxes on labor income, capital income, energy, and pollution plus a uniform lump-sum transfer. The climate module is calibrated to DICE 2016 (Nordhaus, 2017). Household heterogeneity is calibrated to US data: ten productivity groups from SCF 2013 hourly wages ranging from $6.44 (bottom decile) to $101.35 (top decile), yielding a model consumption Gini of 0.33, very close to the empirical value of 0.32 (Heathcote et al., 2010). Tax rates are set at effective US rates from Trabandt and Uhlig (2012): capital income tax of 41.1% and labor income tax of 25.5%. The model period is five years beginning in 2015, and the discount factor follows DICE at beta = 1/(1.015) per year, with inverse IES sigma = 1.45. The main quantitative exercise compares optimal policy to a climate-skeptic planner who sets carbon taxes to zero. Key findings: (i) Tax distortions have a negligible effect on the optimal carbon tax in the heterogeneous-agent setting. The second-best carbon tax is initially only 0.5% below the social cost of carbon (SCC) and subsequently fluctuates within about 0.2% above or below it — in sharp contrast to Barrage (2020), who finds tax distortions reduce optimal carbon taxes by 8% in the representative-agent setting. The key mechanism is that, with heterogeneous agents, the government optimally levies distortionary taxes for redistributive purposes (not merely to finance public spending), so the marginal cost of public funds (MCF) averages to 1 over time and its temporal deviations are quantitatively trivial. (ii) Income inequality only slightly reduces the optimal carbon tax: residual consumption inequality after optimal income-tax redistribution lowers the SCC by 3.9% in the baseline. The mechanism is that inequality raises the average marginal utility of consumption (because the marginal utility function is convex), increasing the opportunity cost of abatement; this effect dominates when IES &amp;lt; 1 (sigma &amp;gt; 1 in the calibration). (iii) The optimal carbon tax path starts at $21.7/tCO2 in 2020 and reaches $229.2/tCO2 one century later — levels consistent with Barrage (2020) and Nordhaus (2017/2018) but insufficient to achieve the Paris +2°C target under baseline damages. (iv) Comparing optimal policy to the climate-skeptic baseline, the additional carbon tax revenue is split nearly equally: the present value of labor taxes falls by 0.7% of GDP, while transfers rise by 0.8% of GDP. This violates the weak double-dividend hypothesis, which prescribes using carbon tax revenue exclusively to cut distortionary taxes. (v) The optimal policy has progressive welfare effects in the 21st century, because increased tax progressivity benefits lower-income households. The average discounted welfare gain is 5.8% of consumption under baseline damages. In the long run, gains become regressive because richer households (with IES &amp;lt; 1) are willing to pay proportionally more in consumption to avoid temperature increases. By contrast, a representative-agent double-dividend policy — using all carbon revenue to cut labor taxes — is regressive from the outset, with low-income households bearing a net cost even in the short run. The 3.9% inequality effect on the SCC is robust to changes in fiscal pressure and damage calibration but is sensitive to sigma: with sigma = 2, inequality reduces optimal carbon taxes by 16.2% rather than 3.9%. Extensions with wealth heterogeneity, heterogeneous energy demand (calibrated to CEX), and heterogeneous environmental damage sensitivity confirm that the MCF remains negligible and the inequality effect on carbon taxes remains small in quantitative terms.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-theoretical-result-on-the-optimal-carbon-tax-and-why-does-it-differ-from-barrage-2020"&gt;Q1. What is the core theoretical result on the optimal carbon tax and why does it differ from Barrage (2020)?&lt;/h3&gt;
&lt;p&gt;The optimal carbon tax is approximately Pigouvian — set equal to the social cost of carbon — because the MCF averages to 1 over time with balanced-growth preferences when households are heterogeneous and the government can optimize a uniform lump-sum transfer. In Barrage (2020)&amp;rsquo;s representative-agent model, the government cannot choose the level of lump-sum taxes or transfers because there is no redistribution motive, so distortionary taxes are the only way to finance public spending and the MCF exceeds 1, reducing optimal carbon taxes by 8%. With heterogeneous agents, the government optimally provides lump-sum transfers for redistribution, so the constraint on transfers is barely binding and the MCF is close to 1 even when the ability to adjust transfers is removed.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-mechanism-by-which-inequality-affects-the-optimal-carbon-tax-and-what-is-the-sign"&gt;Q2. What is the mechanism by which inequality affects the optimal carbon tax, and what is the sign?&lt;/h3&gt;
&lt;p&gt;Inequality reduces the optimal carbon tax when IES &amp;lt; 1 (sigma &amp;gt; 1). The mechanism operates through the Pigouvian tax formula: pollution abatement reduces aggregate consumption, and the welfare cost of this reduction depends on the social marginal utility of consumption (Vc,t). With inequality, Vc,t is affected by two opposing forces. First, the average marginal utility of consumption is higher because of Jensen&amp;rsquo;s inequality (convex marginal utility function), increasing the opportunity cost of abatement and pushing the pollution tax down. Second, additional consumption goes disproportionately to richer households with lower marginal utilities, reducing Vc,t and pushing the tax up. When IES &amp;lt; 1, the first (higher average marginal utility) effect dominates, so inequality unambiguously reduces the SCC and hence the optimal pollution tax. When IES = 1, the two effects exactly cancel and inequality has no effect.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-mcf-and-why-does-it-average-to-1-in-the-heterogeneous-agent-setting"&gt;Q3. What is the MCF and why does it average to 1 in the heterogeneous-agent setting?&lt;/h3&gt;
&lt;p&gt;The MCF is defined as the ratio of the public (planner&amp;rsquo;s Lagrange multiplier on the resource constraint) to the private (aggregate welfare-weighted) marginal utility of consumption. It measures the social cost of transferring resources from the private to the public sector. The MCF averages to 1 because the first-order condition for the uniform lump-sum transfer implies that the sum of the Lagrange multipliers on agents&amp;rsquo; implementability constraints is zero. With balanced-growth preferences, this implies the welfare-weighted average MCF equals 1 from period 0. The temporal covariance between type-specific shadow costs (theta_i) and the type-specific implementability term (I_{c,i,t}) averages to zero over time, so while the MCF can deviate temporarily from 1, it is 1 on average.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-double-dividend-hypothesis-and-how-does-the-papers-optimal-policy-relate-to-it"&gt;Q4. What is the double-dividend hypothesis and how does the paper&amp;rsquo;s optimal policy relate to it?&lt;/h3&gt;
&lt;p&gt;The weak double-dividend hypothesis holds that it is optimal to use carbon tax revenue exclusively to reduce distortionary taxes, yielding both environmental and efficiency dividends. The paper shows this does not hold with heterogeneous agents: at the optimum, the welfare gain from a marginal reduction in tax distortions equals the welfare loss from increased inequality, so the government splits carbon revenue between cutting distortionary taxes and increasing redistribution. In the baseline quantification, the split is roughly equal: present-value labor taxes fall by 0.7% of GDP and lump-sum transfers rise by 0.8% of GDP. By contrast, following the double-dividend prescription — using all carbon revenue to reduce labor taxes without raising transfers — generates a strongly regressive policy in which low-income households bear net welfare costs even in the short run.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-calibration-strategy-and-how-does-the-model-match-us-inequality-data"&gt;Q5. What is the calibration strategy and how does the model match US inequality data?&lt;/h3&gt;
&lt;p&gt;The economic side is calibrated to the US, while the climate side uses DICE 2016. The discount factor follows DICE (beta = 1/(1.015) per year), and sigma = 1.45 (IES = 1/1.45). Household productivity is calibrated using SCF 2013 hourly wage deciles, yielding ten equal-sized groups with hourly wages from $6.44 (bottom) to $101.35 (top), normalized so that the productivity-weighted average is 1. Although productivity inequality is directly targeted rather than moments of the consumption distribution, the model correctly predicts the consumption Gini of 0.33, close to the empirical 0.32 (Heathcote et al., 2010). Capital and labor income tax rates are from Trabandt and Uhlig (2012): 41.1% and 25.5% respectively. Government debt-to-GDP is approximately 111% (average 2011-2015, IMF). The Frisch elasticity of labor supply is targeted at 0.75 (Chetty et al., 2011). Production in both sectors is Cobb-Douglas with energy share nu = 0.04 from Golosov et al. (2014).&lt;/p&gt;
&lt;h3 id="q6-what-happens-to-optimal-income-taxes-in-the-model"&gt;Q6. What happens to optimal income taxes in the model?&lt;/h3&gt;
&lt;p&gt;The optimal labor income tax roughly doubles from its calibrated level of 25% to about 50% in the first period and stabilizes there. Revenue from these taxes is rebated via the uniform lump-sum transfer, achieving most of the desired redistribution. Because optimal labor income taxes are approximately constant over time, the associated intertemporal distortions are small, and the optimal capital income tax converges to zero quickly after the second period. The mechanism is that, with access to lump-sum transfers, the only reason to tax capital income is to mitigate intertemporal distortions created by labor income taxation; when labor taxes are roughly constant, this motive is weak.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-sensitivity-analysis-reveal-about-the-robustness-of-the-39-inequality-effect"&gt;Q7. What does the sensitivity analysis reveal about the robustness of the 3.9% inequality effect?&lt;/h3&gt;
&lt;p&gt;The effect of inequality on optimal carbon taxes is robust along several dimensions but sensitive to sigma. Under the high-damage scenario (cubic rather than quadratic damage function, yielding an SCC about four times larger), the inequality effect falls to 2.6% rather than 3.9%, because higher carbon taxes reduce warming and thus the share of utility (rather than production) damages. The effect is roughly proportional to the degree of productivity inequality: half the inequality implies about half the effect on the carbon tax. The effect changes more than proportionally with sigma: with sigma = 2 (IES = 0.5), inequality reduces carbon taxes by 16.2%, versus 3.9% with the DICE value of sigma = 1.45. With sigma = 1, the effect is exactly zero. Government expenditure levels and fiscal pressure have negligible effects on the results. The share of damages entering utility directly matters: if only 10% of damages affect utility directly (versus the baseline 26%), the inequality effect falls to 1.8%; if 40% affect utility directly, it rises to 5.2%.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-role-of-initial-wealth-inequality"&gt;Q8. What is the role of initial wealth inequality?&lt;/h3&gt;
&lt;p&gt;Initial wealth inequality (studied in Section 6.1) creates an additional motive for deviating from Pigouvian taxation in period 0 only. Because the planner cannot use the period-0 capital tax to expropriate initial wealth (it is fixed at 41.1%), higher damages would reduce interest rates and thereby partially mitigate wealth inequality (a subtle indirect redistribution mechanism), calling for lower pollution taxes in period 0. Quantitatively, this produces a significant reduction in the initial-period optimal carbon tax. However, from period 1 onward, the optimal tax rules are unaffected by initial wealth heterogeneity, and the effects of MCF and income inequality remain very similar to the baseline. Welfare gains from carbon taxation in the wealth-heterogeneity extension are U-shaped with income but strictly increasing in initial wealth.&lt;/p&gt;
&lt;h3 id="q9-how-does-energy-demand-heterogeneity-stone-geary-extension-affect-the-results"&gt;Q9. How does energy-demand heterogeneity (Stone-Geary extension) affect the results?&lt;/h3&gt;
&lt;p&gt;The extension introduces a second dirty consumption good with Stone-Geary preferences, calibrated using CEX data to match the average energy expenditure share of 10.8% and the observed distribution of energy budget shares across and within income groups. Target emissions share from household energy consumption is 30%. The optimal pollution tax formula remains a modified Pigouvian rule (the MCF structure is unchanged), and the MCF effect remains negligible. The inequality effect on carbon taxes stays near 3.9%, rising marginally to 4.1% with identical energy necessity and 4.1% with heterogeneous energy necessity. Theoretically, the optimal excise tax on the energy good is zero when energy preferences are homogeneous; with heterogeneous necessity levels calibrated to the US, the optimal energy excise tax is quantitatively tiny: about -0.4% of energy prices (a small subsidy). The negative sign arises because within-income-group heterogeneity in energy needs means that energy-intensive households (who are valued more by the planner on average) can be partially targeted via a subsidy. Under the double-dividend scenario with energy inequality, regressive effects are magnified: the poorest, most energy-intensive households actually lose in welfare terms even accounting for long-run climate mitigation benefits.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-paper-establish-theoretically-about-heterogeneous-environmental-damages"&gt;Q10. What does the paper establish theoretically about heterogeneous environmental damages?&lt;/h3&gt;
&lt;p&gt;Proposition 6 (Section 6.3) shows that with additively separable environmental utility and a utilitarian planner, heterogeneous marginal utility damages from pollution have no effect on the optimal pollution tax: they enter the welfare criterion symmetrically and cancel in the aggregate. The pollution tax increases relative to the utilitarian benchmark only if the planner&amp;rsquo;s welfare weights are positively correlated with marginal utility damages — that is, if the planner cares relatively more about the households that are more exposed. A Rawlsian planner would set a higher pollution tax if and only if the least-well-off household is also more sensitive to environmental degradation.&lt;/p&gt;
&lt;h3 id="q11-what-are-third-best-policy-results-when-either-income-tax-is-fixed"&gt;Q11. What are third-best policy results when either income tax is fixed?&lt;/h3&gt;
&lt;p&gt;The paper analyzes policies where either the labor or capital income tax is fixed at its current calibrated level (studied in Appendix E, with results referenced in the main text). These constraints introduce an additional fiscal interaction effect on the optimal carbon tax — the carbon tax is pushed below its second-best Pigouvian level when the fixed tax is set at a sub-optimally low level, and above it when the fixed tax is sub-optimally high. The roles of the MCF and income inequality remain similar to the second-best baseline under these third-best constraints.&lt;/p&gt;
&lt;h3 id="q12-how-does-the-paper-relate-to-and-differ-from-the-double-dividend-and-pollution-taxation-literatures"&gt;Q12. How does the paper relate to and differ from the double-dividend and pollution taxation literatures?&lt;/h3&gt;
&lt;p&gt;The paper builds on three earlier pillars. First, Pigou (1920) established first-best Pigouvian taxation. Second, a large literature (Sandmo, 1975; Bovenberg and de Mooij, 1994; Bovenberg and Goulder, 1996) showed that in representative-agent second-best settings the MCF exceeds 1 and optimal pollution taxes fall below the Pigouvian level. Barrage (2020) is the closest dynamic general-equilibrium predecessor, finding the 8% reduction from tax distortions. Third, Jacobs and de Mooij (2015) and Jacobs and van der Ploeg (2019) showed in static models with heterogeneous agents and a uniform lump-sum transfer that the MCF equals 1. This paper extends this insight to a fully dynamic climate-economy framework with general equilibrium and a rich model of household heterogeneity. The key innovation relative to Barrage (2020) is agent heterogeneity, which both provides microfoundations for distortionary taxation and significantly changes the quantitative implications for optimal carbon taxes. Relative to Jacobs and de Mooij (2015), the contribution is the dynamic setting, the linkage to the DICE climate module, and the full quantitative characterization including distributional welfare analysis and multiple sources of heterogeneity.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q13. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary policy implication is that a carbon tax should be set approximately equal to the SCC (Pigouvian level) and the associated revenue should be split roughly equally between increasing lump-sum transfers and reducing distortionary labor taxes — rather than following the double-dividend prescription of using all revenue to reduce distortionary taxes. This combination is both more efficient (the MCF argument) and more equitable (progressive in the short run). The scope conditions are: (a) the result applies under a utilitarian welfare criterion with linear income taxes and a uniform lump-sum transfer; (b) it requires that the government can optimize the level of lump-sum transfers for redistribution; (c) the approximately Pigouvian result is quantitatively robust to alternative damage functions, fiscal pressure, and energy demand heterogeneity, but the degree to which inequality lowers the carbon tax depends sensitively on the IES/inequality aversion parameter sigma; (d) the calibration is designed to capture US conditions assuming that the US internalizes the full global impact of its emissions (strategic considerations are abstracted away); (e) heterogeneous environmental damage sensitivity does not affect the utilitarian optimum, but would increase the optimal carbon tax under a more inequality-averse social planner.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Marginal Cost of Public Funds (MCF)&lt;/strong&gt;: The ratio of the public (planner&amp;rsquo;s shadow price on the resource constraint) to the private (aggregate welfare-weighted) marginal utility of consumption. In this paper, it captures the divergence between second-best and first-best pollution taxes due to fiscal distortions. With heterogeneous agents and an optimized uniform lump-sum transfer, the MCF averages to 1 over time under balanced-growth preferences, implying that tax distortions do not systematically push the carbon tax below the Pigouvian level — unlike in the representative-agent setting where the MCF exceeds 1.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pigouvian tax (second-best)&lt;/strong&gt;: In this paper&amp;rsquo;s context, the Pigouvian tax refers to the pollution tax equal to the social cost of pollution (the discounted present value of marginal production and utility damages), evaluated at the second-best allocation rather than the first-best. When the MCF equals 1 (as it approximately does in the heterogeneous-agent setting), the second-best optimal pollution tax is equal to this second-best Pigouvian level, which may itself differ from the first-best Pigouvian level due to residual consumption inequality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Cost of Carbon (SCC)&lt;/strong&gt;: The present discounted value of marginal climate damages (both production and utility losses) from emitting one additional ton of CO2, converted into consumption units using the social marginal utility of consumption. In the paper, the SCC corresponds to the case where the MCF is set to 1 in every period, and it is affected by consumption inequality through its effect on the social marginal utility of consumption. With sigma &amp;gt; 1, residual inequality raises the opportunity cost of abatement, reducing the SCC by 3.9% in the baseline calibration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Double-dividend hypothesis (weak)&lt;/strong&gt;: The claim that it is optimal to use the entire proceeds of a carbon tax to reduce existing distortionary taxes, yielding both an environmental dividend (less pollution) and an efficiency dividend (lower tax distortions). The paper shows this does not hold with heterogeneous agents: because distortionary taxes serve a redistributive purpose, reducing them at the margin has a welfare cost (increased inequality), so the planner optimally splits revenue between tax reduction and increased transfers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ramsey problem (climate-economy)&lt;/strong&gt;: The government&amp;rsquo;s optimization problem in this paper: maximizing utilitarian social welfare over an infinite horizon by choosing paths for linear taxes on labor income, capital income, energy, and pollution, plus a uniform lump-sum transfer, subject to households&amp;rsquo; optimality conditions (implementability constraints), resource constraints, climate dynamics from DICE, and abatement technology constraints. The approach extends Werning (2007) to a dynamic climate-economy context.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implementability condition&lt;/strong&gt;: The constraint in the Ramsey problem that captures each household&amp;rsquo;s lifetime budget constraint in terms of aggregate variables and market weights. It requires that the present value of a household&amp;rsquo;s consumption minus labor income equals its initial assets plus its share of the present value of lump-sum transfers, evaluated using the social marginal utilities implied by the planner&amp;rsquo;s choice of taxes. The shadow cost of this constraint for each household type (theta_i) determines the MCF through its covariance with a fiscal externality term.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Residual inequality&lt;/strong&gt;: The level of inequality that remains after the planner has optimally set all income taxes and the lump-sum transfer — i.e., the inequality that cannot be eliminated because individualized lump-sum transfers are not feasible and only linear instruments are available. In the paper, it is this residual inequality (not total inequality) that affects the optimal carbon tax: the carbon tax responds to the inequality that income-tax policy cannot address, not to the underlying productivity or wealth dispersion per se.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Balanced-growth preferences&lt;/strong&gt;: A preference specification of the form u(c, h, Z) = [c(1 - varsigma*h)^gamma]^(1-sigma)/(1-sigma) + u_hat(Z), with 1/sigma the intertemporal elasticity of substitution. This specification ensures that the economy admits a balanced growth path and plays a key role in the paper&amp;rsquo;s theoretical results: under balanced-growth preferences, the welfare-weighted average MCF equals 1 from period 0, and when IES = 1 (sigma = 1) the MCF is exactly 1 in every period.&lt;/p&gt;</description></item><item><title>Populism and the Skill-Content of Globalization</title><link>https://macropaperwarehouse.com/papers/populism-and-the-skill-content-of-globalization/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/populism-and-the-skill-content-of-globalization/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates how the skill structure of globalization shocks — rather than globalization per se — drives the long-run evolution of populism across countries, making a unified empirical case that what gets imported or who immigrates matters as much as how much.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; The literature has documented that trade exposure and immigration fuel populist voting, but prior work has studied these channels separately, used narrow time windows, and relied on binary party classifications that cannot capture shifts in populism across the full party landscape. Rodrik&amp;rsquo;s (2018) widely-cited hypothesis holds that trade shocks drive left-wing populism (as in Latin America) and immigration drives right-wing populism (as in Europe). The authors examine whether this hypothesis survives when skill content is explicitly disaggregated and both channels are studied jointly in a unified long-panel setting.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data, sample, and empirical strategy.&lt;/strong&gt; The authors construct a new continuous, time-varying populism score for 3,860 party-election pairs covering 1,206 unique parties across 628 national elections in 55 countries from 1960 to 2018. The score is built from the Manifesto Project Database (MPD) using two dimensions identified in the political-science literature: an anti-establishment stance (AES) and a commitment-to-protect stance (CTP). A two-stage polychoric PCA extracts synthetic indices for each dimension and then combines them into a single populism score. The paper defines populist parties as those scoring more than one standard deviation above the mean (threshold validated by comparison with four external databases — Van Kessel, Swank, PopuList, GPop 1 — with ratios of accurate forecasts ranging from 80 to 91 percent). Two dependent variables are studied: (i) the volume margin of populism, the vote share of classified populist parties, estimated with PPML given many zero observations (about 60 percent of the full sample); and (ii) the mean margin of populism, the vote-weighted average populism score of all parties, estimated with OLS. Globalization regressors are skill-specific: imports of low-skill and high-skill labor-intensive goods (as shares of GDP, sourced from Feenstra et al. 2005 and UN Comtrade) and immigration inflows of low-skill and high-skill workers (from Abel 2018, skill-level imputed from dyadic migrant-stock selection ratios). To address reverse causality — populist governments restrict trade and immigration, biasing OLS downward — the authors implement a gravity-based IV strategy: a zero-stage PPML regression predicts bilateral flows using time-invariant dyadic fixed effects interacted with a post-1990 dummy and origin-country-year fixed effects, then aggregates to the destination level; these predicted flows serve as instruments. For the volume margin, a reduced-form IV approach replaces actual with predicted flows (to avoid the incidental-parameter problem in PPML with fixed effects). For the mean margin, standard 2SLS is used; the Kleibergen-Paap F-statistic is around 10–12, reasonable given four instruments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main quantitative findings.&lt;/strong&gt; (All claims below are with country and year fixed effects throughout; IV results reinforce baseline OLS/PPML results.)&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Low-skill labor-intensive imports raise total and right-wing populism along both the volume margin and the mean margin. In the OLS mean-margin specification the coefficient on low-skill imports is approximately 4, implying a 1 percentage-point increase in the import-to-GDP ratio for low-skill goods is associated with a 0.04 increase in the mean margin of populism (scaled in standard deviations of the populism score). The 2SLS coefficient on the total mean margin is approximately 5.0 (significant at 5%), and on the right-wing mean margin approximately 4.1 (significant at 5%). For the volume margin, the reduced-form IV coefficient on low-skill imports is 0.91 (significant at 10%) for total and 1.82 (significant at 5%) for right-wing populism. These effects are larger by a factor of approximately 1.3 when IV is used relative to OLS/PPML, consistent with downward bias from reverse causality. Low-skill imports do not significantly affect left-wing populism in baseline estimates; a left-wing response cannot be ruled out during severe crises, when shocks are persistent, or among EU countries specifically.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;High-skill labor-intensive imports reduce the volume of populism, especially right-wing populism. In the reduced-form IV specification the coefficient on high-skill imports is -1.22 (significant at 10%) for total volume and -2.14 (significant at 5%) for right-wing volume. The mean-margin effect of high-skill imports is insignificant.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Low-skill immigration induces a transfer of votes from left-wing to right-wing populist parties, leaving total volume and the mean margin unchanged. The baseline PPML coefficient on low-skill immigration is 1.52 (significant at 1%) for right-wing volume and -1.78 (significant at 1%) for left-wing volume. In the reduced-form IV the right-wing volume coefficient is 1.97 (significant at 1%) and the left-wing coefficient is -1.70 (significant at 10%). The mean margin of total populism is not significantly affected by low-skill immigration in any specification.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;High-skill immigration reduces the volume of right-wing populism (PPML coefficient -1.32, significant at 1%; IV coefficient -2.02, significant at 5%) and generates a weak substitution toward left-wing populism in the baseline.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Descriptive findings: populism fluctuated since the 1960s, peaking after major economic crises (the oil shocks of the 1970s, deep crises of the 1990s, and after 2008). Right-wing populism reached an all-time high in the EU after 2005. The share of elections with at least one right-wing populist party rose from about 5 percent to more than 50 percent in EU member states over the study period.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Mechanisms.&lt;/strong&gt; Decomposing the volume margin into extensive (number of populist parties) and intensive (average vote share per party) sub-margins reveals that: the trade channel operates primarily through the intensive margin (existing populist parties gaining more votes); the immigration channel operates through the extensive margin (new right-wing populist parties with moderate scores entering parliament). Low-skill trade and immigration never increase the populism score of parties that have never been classified as populist, indicating that globalization shifts the composition of the party system rather than radicalizing mainstream parties.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Amplifiers and heterogeneity.&lt;/strong&gt; The right-wing populism response to low-skill imports is amplified during periods of de-industrialization and when internet coverage is high. Diversity in the origin mix of imported goods dampens the right-wing response. The populism response to low-skill immigration is not amplified by cultural distance between natives and immigrants; if anything, high cultural distance slightly reduces the centrist and left-wing populist responses. The effects on volume margin are primarily driven by EU28 countries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions and caveats.&lt;/strong&gt; Analysis is at the country level; party-level repositioning dynamics are left for further research. The unified trade-plus-immigration framework is new, but the long panel setting, unbalanced sample, and aggregate data impose limits on identifying specific mechanisms. The finding that globalization does not affect never-populist parties&amp;rsquo; scores limits concerns about contamination through party contagion in the short run. These results only partially confirm Rodrik&amp;rsquo;s (2018) hypothesis — left-wing populism is not robustly driven by trade shocks at the aggregate level, and trade&amp;rsquo;s effects are not confined to non-European contexts.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The identification relies on a two-stage approach. In the first stage (zero-stage gravity model), the authors predict bilateral flows of low- and high-skill goods and migrants using (i) time-invariant dyadic fixed effects interacted with a post-1990 structural-break dummy and (ii) origin-country-year fixed effects capturing time-varying push factors at the source. Critically, destination-country-time characteristics are excluded from the zero-stage, so the predicted aggregated flows capture only supply-side variation and bilateral connectivity — not demand-side populism dynamics in the destination. These predicted flows are then used as instruments. For the mean margin, standard 2SLS is implemented; for the volume margin, a reduced-form IV approach replaces actual flows with predicted flows to avoid the incidental-parameter problem in a PPML model with many fixed effects. The main threats are: (1) correlated origin shocks — if a push shock in origin country j simultaneously triggers populism in destination i through channels other than trade/migration (e.g., financial contagion), the exclusion restriction is violated; the authors cannot fully rule this out but note that including year fixed effects absorbs common global shocks; (2) the post-1990 structural break is used as an additional source of variation for bilateral dyadic ties, but the Berlin Wall dummy simultaneously captures many unobserved structural changes; (3) imputation of the skill structure of migration flows from census-round selection ratios (1990, 2000, 2010) introduces measurement error, though the authors show robustness to using only the year-2000 ratio; (4) Kleibergen-Paap F-statistics are around 10–12 when all four endogenous variables are instrumented simultaneously, which is modest; the authors show values are substantially larger when instrumenting one or two variables at a time.&lt;/p&gt;
&lt;h3 id="q2-how-are-trade-and-immigration-distinguished-empirically-and-how-is-the-skill-content-measured"&gt;Q2. How are trade and immigration distinguished empirically, and how is the skill content measured?&lt;/h3&gt;
&lt;p&gt;Trade data come from Feenstra et al. (2005) for 1962–2000 and UN Comtrade for 2001–2015. Product categories at the SITC 3-digit level are classified by skill and technology intensity following the Trade and Development Report (2002), yielding five categories: primary commodities, labor-intensive/resource-based, and manufacturing with low-, medium-, and high-skill labor intensity. The baseline uses only the low-skill and high-skill manufacturing ends; medium-skill goods are tested in robustness (their inclusion causes collinearity that kills volume-margin significance while preserving mean-margin results). Migration data come from Abel (2018) — five-year bilateral migration flow estimates interpolated to annual frequency. The skill level of migration flows is imputed by applying census-round skill-selection ratios (ratio of college graduates in the dyadic migrant stock to the native pre-migration population, from the closest available census round of 1990, 2000, or 2010) to the interpolated flows. Both trade and immigration variables enter as percentages — imports as share of GDP, immigration as share of destination population — averaged over the election year and the preceding year.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-difference-between-the-volume-margin-and-the-mean-margin-of-populism-and-why-does-it-matter"&gt;Q3. What is the difference between the volume margin and the mean margin of populism, and why does it matter?&lt;/h3&gt;
&lt;p&gt;The volume margin is the aggregate vote share of parties classified as populist (using a binary threshold of one standard deviation above mean in the populism score); it equals zero in elections with no populist party (about 60 percent of observations). The mean margin is the vote-weighted average populism score of all parties — populist and non-populist alike — so it is always defined and continuous. The mean margin captures the average ideological &amp;rsquo;exposure&amp;rsquo; of voters to populist ideas in a given election, including the spillover of populist ideas into mainstream parties. The distinction matters because globalization can affect the political landscape through multiple channels: it may shift votes toward existing populist parties (intensive margin of the volume margin), it may encourage new populist parties to enter (extensive margin), or it may shift the policy positions of all parties toward more populist stances (captured by the mean margin). The paper finds that low-skill trade raises both margins, but through different mechanisms — the volume effect operates through the intensive margin while the mean-margin effect partly reflects score increases among centrist populist parties. Low-skill immigration raises only the volume margin (through extensive-margin changes, not the mean margin).&lt;/p&gt;
&lt;h3 id="q4-how-is-the-populism-score-constructed-and-how-is-it-validated"&gt;Q4. How is the populism score constructed, and how is it validated?&lt;/h3&gt;
&lt;p&gt;The score is built from the Manifesto Project Database, which counts quasi-sentences associated with specific political topics as shares of party manifestos. Six MPD variables are selected, grouped into two dimensions: anti-establishment stance (AES — political corruption mentions and anti-pluralism/political authority mentions) and commitment-to-protect stance (CTP — protectionism, internationalism, EU institutions, and nationalization). A polychoric PCA within each dimension extracts the first principal component (by Kaiser criterion — eigenvalues above one). The two synthetic indices are then combined into a single populism score by equal weighting. A party is classified as populist if its score exceeds one standard deviation above the mean. This threshold maximizes the partial correlation with three of four external databases and maximizes accurate-forecast rates across all four databases. Probit regressions of existing binary classifications (Van Kessel 2015, Swank 2018, PopuList 2019, GPop 1 2020) on the continuous score yield ratios of accurate forecasts between 80 and 91 percent. OLS correlations with continuous external measures (GPop 2 leader-speech scores, CHES expert survey) are positive and significant. Unsupervised k-means clustering on the (AES, CTP) space confirms that parties above the one-SD threshold cluster distinctly in a well-separated region of the two-dimensional space. Extended scores using more MPD variables do not improve fit, confirming parsimony.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-across-left-wing-and-right-wing-populism-is-documented"&gt;Q5. What heterogeneity across left-wing and right-wing populism is documented?&lt;/h3&gt;
&lt;p&gt;The paper systematically decomposes results by political orientation (terciles of the RILE left-right index from MPD). Key heterogeneities: (1) Low-skill imports raise total and right-wing populism but not left-wing populism along the volume margin — this holds in baseline PPML and reduced-form IV. The mean-margin result is also concentrated in total and right-wing. (2) Low-skill immigration shifts votes from left-wing to right-wing populism (with opposing-sign PPML coefficients of 1.52 and -1.78, both significant at 1%), leaving total populism unchanged. High-skill immigration reverses this — it reduces right-wing and weakly increases left-wing populism. (3) High-skill imports reduce right-wing populism particularly (PPML -1.30, IV -2.14) and weakly shift votes toward left-wing populism. (4) Descriptively, the average populism score of right-wing populist parties increased since 2005 and reached 1.7 (2.1 standard deviations) in 2018, while left-wing populist parties&amp;rsquo; average score declined to 1.4 (1.75 standard deviations) — for the first time since the 1960s, radical-right populism is more intense than radical-left. (5) The volume-margin effects of globalization are primarily driven by EU28 countries. Among non-EU countries or when Latin America is excluded, results are directionally preserved but sometimes less precisely estimated.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The authors conduct an extensive battery documented in Appendix D: (1) Lag structure — the globalization variables are redefined using flows at t, t-1, t-2, average of t and t-1 (baseline), and the sum between elections; results on immigration are robust across lags; trade significance holds except at very short (election year) or very long (between elections) windows. (2) Populism threshold — results are preserved at the lax (0.9 SD) threshold and mostly preserved at the strict (1.1 SD) threshold, though some become insignificant when well-known parties like Syriza, M5S, and La France Insoumise exit the classification. (3) Skill imputation for immigration — using only year-2000 selection ratios yields similar results; interactions with migrant-stock quartile dummies are mostly insignificant. (4) Skill content of imports — adding labor-intensive and medium-skill imports does not disturb the baseline; collinearity from medium-skill imports kills volume-margin trade significance. (5) Origin-country income level — positive populism responses are concentrated in flows from low-income countries on the volume margin, but the mean-margin positive response is more driven by North-North movements. (6) Sub-samples — results are not driven by post-1990 years alone (interaction with post-1990 dummy attenuates but does not eliminate effects), not by Latin American countries (exclusion leaves results unchanged), and not by the unbalanced panel structure (restricting to countries present since 1970 confirms results). (7) Turnout — globalization variables do not significantly predict turnout, and results are robust to controlling for turnout. (8) Electoral system — results hold when controlling for electoral system; proportional representation systems show a significant effect of low-skill imports on left-wing populism volume. (9) Exports and emigration — including skill-specific export and emigration flows does not substantially alter the main coefficients; export and emigration effects are less significant and robust than import and immigration effects. (10) Vote-share normalization — results are robust to normalizing vote shares to sum to 100 percent.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work-especially-autor-et-al-2020-and-the-immigration-literature"&gt;Q7. How does this paper relate to and differ from closely related prior work, especially Autor et al. (2020) and the immigration literature?&lt;/h3&gt;
&lt;p&gt;Autor, Dorn, Hanson, and Majlesi (2020) study the electoral consequences of the China trade shock in the US, documenting polarization effects concentrated in a specific trade shock and a narrow time frame. The present paper extends this by: (1) spanning 60 years and 55 countries (vs. US-focused short panels); (2) studying trade and immigration jointly in one specification; (3) using continuous populism scores rather than party platforms; (4) distinguishing left- vs. right-wing populism responses; (5) examining skill content rather than origin-country GDP growth. On immigration, Edo et al. (2019) and Moriconi et al. (2022, 2019) document that the skill structure of immigration matters for voting — high-skill immigration reduces far-right votes while low-skill immigration raises them. The present paper confirms these findings in a much larger multi-decade panel and adds the novel result that low-skill immigration does not affect total populism but merely shuffles votes between left-wing and right-wing populism. On Rodrik&amp;rsquo;s (2018) taxonomy, the paper only partially confirms his hypothesis: left-wing populism is not robustly driven by trade shocks in the cross-country aggregate (only under specific amplifying conditions), and trade&amp;rsquo;s effects are not confined to non-European settings. A key novelty vs. the entire prior literature is the simultaneous inclusion of skill-specific trade and immigration flows — no prior cross-country long-panel study had done this.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The skill-content result implies that globalization&amp;rsquo;s effect on populism depends critically on whether economic integration predominantly involves low-skill or high-skill goods and workers. Policies that shift the composition of globalization toward high-skill activities — skill-upgrading policies, investment in education and retraining, managed migration policies that attract high-skill workers — could mechanically reduce populist pressures. The finding that low-skill immigration transfers votes from left to right without increasing total populism has a nuanced implication: reducing low-skill immigration may primarily benefit left-wing parties at the expense of right-wing ones rather than reducing aggregate political instability. The amplification by de-industrialization and internet access suggests that the populist dividend of adverse trade shocks is largest precisely when affected regions are also losing manufacturing jobs and when social media spreads grievance discourse. The attenuation by diversity in imported goods suggests that more geographically diversified trade may reduce the cultural-threat salience of any single origin. Scope conditions: the volume-margin effects are largely driven by EU28 countries, so the quantitative magnitudes may not generalize to other institutional contexts with different electoral systems; the analysis is at the country level and abstracts from regional labor-market dynamics; party-level repositioning of mainstream parties is not modeled.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-handle-the-measurement-challenge-of-comparing-populism-scores-across-countries-and-time"&gt;Q9. How does the paper handle the measurement challenge of comparing populism scores across countries and time?&lt;/h3&gt;
&lt;p&gt;This is a central methodological concern. The authors use party manifestos, which are available consistently across the 55 countries and the full 1960–2018 period in the Manifesto Project Database, allowing a principled content-based scoring without relying on expert surveys (which are available only for limited periods) or dichotomous external classifications (which are time-invariant in some datasets and country-limited in others). The two-stage PCA with polychoric principal components ensures that the dimensions are extracted from the structure of the data without imposing cardinal interpretations on ordinal quasi-sentence counts. The populism score has zero mean by construction with a standard deviation of 0.81, making cross-country and cross-time comparisons meaningful within the sample. The authors validate cross-country comparability by showing that the GPop 1 classification (which spans 1960–2018 for 36 countries) is well predicted by the score even though the score was not calibrated to that dataset specifically. An unsupervised clustering algorithm (k-means on the two dimensions) independently recovers the same set of parties as those above the one-SD threshold, without using any external label. The authors acknowledge that deliberate exclusion of immigration and multiculturalism variables from the score construction prevents mechanical correlation between the populism measure and the globalization regressors, which is an important design choice for the causal analysis.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-trends-in-the-right-left-decomposition-of-populism-over-the-study-period"&gt;Q10. What are the trends in the right-left decomposition of populism over the study period?&lt;/h3&gt;
&lt;p&gt;Descriptively (Section 3): the number of left-wing populist parties (as counted by the extensive margin) increased more than right-wing populist parties in the most recent period, partly because centrist parties are entering the populist bucket. However, the vote share gains (intensive margin) are dominated by right-wing populist parties. The share of elections with at least one left-wing populist party rose from about 15 to 30 percent globally over the study period. The share of elections with at least one right-wing populist party rose from about 5 to more than 50 percent in the EU and from about 10 to 25 percent in the rest of the world. The average populism score of right-wing populist parties increased since 2005, reaching 1.7 (about 2.1 standard deviations) in 2018, while the average score of left-wing populist parties declined to 1.4 (about 1.75 standard deviations). This means that for the first time since the 1960s, right-wing populist parties are on average more populist (by their own score) than left-wing populist parties. The gap between populist and non-populist parties&amp;rsquo; average scores has widened since 2008, consistent with the within-country Theil inequality increase after the financial crisis.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Volume margin of populism&lt;/strong&gt;: The aggregate vote share obtained by parties classified as populist (those with a populism score exceeding one standard deviation above the mean). Estimated with PPML given the large share of zero observations (about 60 percent of the sample). Captures whether populist parties win more votes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mean margin of populism&lt;/strong&gt;: The vote-weighted average populism score of all parties that obtained at least one seat in an election, regardless of whether they are classified as populist. Captures the average ideological &amp;rsquo;exposure&amp;rsquo; of voters to populist ideas, including spillovers into mainstream parties. Estimated with OLS.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Anti-establishment stance (AES)&lt;/strong&gt;: One of two dimensions underlying the paper&amp;rsquo;s populism score. Measured from Manifesto Project Database quasi-sentences on political corruption and anti-pluralism (political authority), capturing the core populist premise that the people are virtuous and the ruling class corrupt, leaving no room for pluralism or minority protection.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Commitment-to-protect stance (CTP)&lt;/strong&gt;: The second dimension underlying the populism score. Measured from Manifesto Project Database quasi-sentences on protectionism, internationalism, EU institutions, and nationalization, capturing populists&amp;rsquo; claim to shield &amp;rsquo;the people&amp;rsquo; from external or alien economic and cultural threats.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Skill-content of globalization&lt;/strong&gt;: The decomposition of import flows into goods intensive in low-skill vs. high-skill labor (using the SITC 3-digit classification from the Trade and Development Report 2002), and of immigration inflows into low-skill and high-skill workers (using dyadic skill-selection ratios from census rounds). The key empirical innovation of the paper: it is the skill content, not the size, of globalization flows that determines the direction and ideological valence of populist responses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gravity-based IV strategy&lt;/strong&gt;: An instrumentation approach that predicts bilateral skill-specific flows of goods and migrants using a zero-stage PPML regression with time-invariant dyadic fixed effects (interacted with a post-1990 structural-break dummy) and origin-country-year fixed effects, then aggregates predicted flows to the destination level. Excludes destination-country-time characteristics to purge reverse causality (populist governments restricting trade and immigration) and omitted variable bias.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extensive vs. intensive margin of the volume margin&lt;/strong&gt;: The decomposition of the total vote share for populist parties into the number of populist parties running (extensive margin) and the average vote share per populist party (intensive margin). Low-skill imports primarily affect the intensive margin (existing populist parties gain more votes); low-skill immigration primarily affects the extensive margin (new right-wing populist parties enter parliament).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vote-transfer mechanism of low-skill immigration&lt;/strong&gt;: The paper&amp;rsquo;s finding that low-skill immigration reallocates votes between left-wing and right-wing populist parties without changing total populism. The authors interpret this as low-skill immigration enabling new right-wing populist parties with moderate populism scores to gain at least one seat in parliament (an extensive-margin effect), while simultaneously reducing the vote share and/or number of left-wing populist parties.&lt;/p&gt;</description></item><item><title>Taxation of Capital: Capital Levies and Commitment</title><link>https://macropaperwarehouse.com/papers/taxation-of-capital-capital-levies-and-commitment/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/taxation-of-capital-capital-levies-and-commitment/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Barro and Chari (2024) revisit the long-standing debate over optimal capital income taxation, unifying the Chamley-Judd zero-tax result, the Straub-Werning positive-tax amendment, and the Chari-Nicolini-Teles (2020) commitment-based framework into a single coherent analysis centered on the treatment of the &amp;ldquo;period-zero problem.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The research question is fundamental: under what commitment assumptions is the optimal long-run tax rate on capital income zero, positive, or negative, and does optimal policy require special treatment of the initial period? The paper operates entirely within a deterministic neoclassical growth model with a representative household whose preferences are time-separable, separable between consumption and labor, and homothetic — the &amp;ldquo;standard preferences&amp;rdquo; of Chari et al. (2020). The government&amp;rsquo;s tax instruments are proportional consumption tax rates (τ_t^c), proportional asset-income tax rates (τ_t^k), and possibly a one-time proportional levy on initial assets (l_0 ≤ 1). No empirical estimation is performed; the contribution is analytical and quantitative through calibrated simulation.&lt;/p&gt;
&lt;p&gt;The central theoretical finding is that the transitional dynamics of Chamley-Judd and the fully positive long-run capital taxes of Straub-Werning both derive from the same source: the period-zero Ramsey planner&amp;rsquo;s incentive to impose capital levies on assets that happen to exist at the start of the optimization. In Chamley et al., direct levies are precluded (l_0 = 0) and the capital-income tax rate is capped at 100%, so the planner engineers indirect levies via positive future τ_t^k (possibly forever, as Straub-Werning show) and time-varying consumption taxes. In the Chari-Nicolini-Teles (2020) formulation, the planner instead faces a constraint that household initial wealth in utility units (W_0) must meet a designated threshold (W̃_0). Under this constraint, the optimal policy features a one-time direct capital levy l_0 in period zero, zero asset-income taxes in all periods (τ_t^k = 0 for t ≥ 0), and a uniform consumption tax for all t ≥ 0. The level of l_0 and the consumption tax rate are jointly determined to satisfy the wealth constraint and the government budget.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s main contribution is extending the Chari et al. period-zero commitment to all periods, thereby achieving time-consistency and eliminating period zero&amp;rsquo;s special status. If each period-t policymaker faces a wealth constraint W_t ≥ W̃_t with W̃_t set high enough that the policymaker voluntarily chooses l_t = 0, the full sequence of policies is time-consistent and accords with Woodford&amp;rsquo;s (1999) &amp;ldquo;timeless perspective&amp;rdquo;: period zero is like any other period, capital-income tax rates are always zero, and consumption taxes are constant.&lt;/p&gt;
&lt;p&gt;The appendix provides quantitative validation using a U.S.-calibrated model: government consumption = 20% of output, capital-income tax rate = 38% (initial steady state, from Barro-Furman 2018), public debt = 70% of output, labor-income tax rate = 26%, discount factor β = 0.97 (implying a 3% real interest rate), capital share α = 0.34, and depreciation δ = 0.08. Welfare gains from switching to the Ramsey policy (with the wealth-in-utility constraint set to the pre-reform steady-state value) are 0.82% of steady-state consumption under standard preferences, 0.76% under balanced-growth preferences, and 0.62% under zero-wealth-effect preferences. Under balanced-growth preferences, the capital stock rises monotonically to a new steady state approximately 12% higher, government debt rises about 6 percentage points, the labor-income tax rate stays essentially constant at approximately 30% (roughly 4 percentage points above the old steady state), and the capital-income tax rate is approximately 1% in the first period and then drops quickly to zero. Under zero-wealth-effect preferences, the initial capital-income tax rate is slightly higher at approximately 7% before dropping sharply. Under an extreme scenario with the initial capital stock at half its steady-state level and public debt at twice its normal ratio, the capital-income tax rate starts at approximately 3% and gradually approaches zero. In all three cases, constraining the capital-income tax rate to zero and holding the labor-income tax rate constant yields welfare indistinguishable from the unconstrained Ramsey optimum. The paper concludes that zero taxation of capital income is approximately optimal across all three preference specifications, and that the apparent necessity of positive long-run capital taxes in existing literature is an artifact of the period-zero commitment asymmetry.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-period-zero-problem-and-why-is-it-central-to-the-papers-argument"&gt;Q1. What is the &amp;lsquo;period-zero problem&amp;rsquo; and why is it central to the paper&amp;rsquo;s argument?&lt;/h3&gt;
&lt;p&gt;The period-zero problem refers to the asymmetry in the standard Ramsey formulation whereby the period-zero policymaker can commit to all future tax rates but is not bound by any commitments made in the past. Because assets already in existence at period zero are inelastically supplied ex post, the planner has a strong incentive to expropriate them via a capital levy — directly (l_0) or indirectly through high early tax rates on asset income or non-constant consumption tax rates. Chamley-Judd and Straub-Werning results, while superficially different, both arise from this same incentive. The Barro-Chari paper argues that period zero is in reality just an arbitrary starting point for analysis, not a date on which commitment ability uniquely materializes, and that correctly accounting for this eliminates the period-zero problem.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-chari-nicolini-teles-2020-formulation-differ-from-chamley-et-al-and-what-does-it-imply"&gt;Q2. How does the Chari-Nicolini-Teles (2020) formulation differ from Chamley et al., and what does it imply?&lt;/h3&gt;
&lt;p&gt;Chamley et al. preclude direct capital levies (l_0 = 0) and cap τ_t^k ≤ 1, so the planner engineers indirect capital levies via positive future asset-income taxes and time-varying consumption taxes. Chari et al. (2020) instead constrain the household&amp;rsquo;s initial wealth in utility units (W_0) to be at least a designated threshold W̃_0, but leave all tax instruments unrestricted. Under this constraint, the optimal policy selects a one-time direct capital levy l_0, zero asset-income taxes forever, and uniform consumption taxes. The critical difference is that when l_0 = 0 is the outcome under the Chari et al. formulation, it is an optimizing response to a high W̃_0 rather than an arbitrary restriction, so there is no incentive for indirect levies.&lt;/p&gt;
&lt;h3 id="q3-how-is-time-consistency-achieved-and-what-is-the-timeless-perspective"&gt;Q3. How is time-consistency achieved, and what is the &amp;rsquo;timeless perspective&amp;rsquo;?&lt;/h3&gt;
&lt;p&gt;Time-consistency fails if future policymakers are unconstrained because they will repeat the period-zero capital levy logic for their own &amp;lsquo;initial&amp;rsquo; period. The paper shows that introducing a series of per-period wealth constraints — W_t ≥ W̃_t for all t ≥ 0, where W_t is period-t household wealth in utility units — achieves time-consistency if each W̃_t is set high enough that each policymaker voluntarily chooses l_t = 0. The required sequence of W̃_t corresponds exactly to the wealth path generated by the period-0 policymaker&amp;rsquo;s committed Ramsey plan. When this holds, the analysis conforms to Woodford&amp;rsquo;s (1999) &amp;rsquo;timeless perspective&amp;rsquo;: each policymaker adopts the program that would have been committed to far in the past, period zero is not special, capital-income taxes are always zero, and consumption taxes are constant.&lt;/p&gt;
&lt;h3 id="q4-what-role-do-restrictions-on-tax-instruments-play-and-why-does-the-paper-prefer-wealth-constraints-over-direct-instrument-restrictions"&gt;Q4. What role do restrictions on tax instruments play, and why does the paper prefer wealth constraints over direct instrument restrictions?&lt;/h3&gt;
&lt;p&gt;Direct instrument restrictions — such as banning capital levies (l_t = 0) or forcing τ_t^k = 0 and constant consumption taxes — are vulnerable to circumvention through other instruments. For example, time-varying labor-income tax rates (τ_t^n) introduce intertemporal wedges equivalent to indirect capital levies, so a prohibition on capital-income taxes can be undone by varying labor taxes. Constraints on household wealth in utility units (Eqs. 7 and 8) are robust to this vulnerability because any tax instrument that reduces household utility-unit wealth below the threshold violates the constraint, regardless of which specific instrument is used.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-partial-commitment-interpretation-of-the-per-period-wealth-constraints"&gt;Q5. What is the &amp;lsquo;partial commitment&amp;rsquo; interpretation of the per-period wealth constraints?&lt;/h3&gt;
&lt;p&gt;The paper offers two interpretations. The first is that the sequence of W̃_t was set at the founding of a country (e.g., 1789 for the United States). The more palatable &amp;lsquo;partial commitment&amp;rsquo; interpretation is that each period-t policymaker specifies the wealth commitment W̃_{t+1} for the next policymaker, in exchange for adhering to the commitment W̃_t set by the preceding policymaker. This bilateral exchange generates the same sequence of wealth constraints that would have been set arbitrarily far into the past.&lt;/p&gt;
&lt;h3 id="q6-what-happens-in-the-stochastic-extension-of-the-model"&gt;Q6. What happens in the stochastic extension of the model?&lt;/h3&gt;
&lt;p&gt;In a stochastic setting with fluctuations in government spending, technology, war and peace, etc. (as in Chari et al. 2020, proposition 3), choices of capital levies and tax rates become state-contingent rules, following the Lucas-Stokey (1983) framework. Non-zero direct capital levies are optimal under emergency conditions such as war, pandemic, or major financial crisis, and correspondingly below average during non-emergencies. Consumption and labor-income tax rates follow random-walk-like processes, analogous to the tax-rate smoothing predictions of Barro (1979, 1990) that apply when state-contingent capital levies are unavailable.&lt;/p&gt;
&lt;h3 id="q7-how-is-the-covid-inflation-episode-interpreted-within-this-framework"&gt;Q7. How is the COVID inflation episode interpreted within this framework?&lt;/h3&gt;
&lt;p&gt;The paper interprets the post-2020 rise in the U.S. price level through the fiscal theory of the price level (Cochrane 2023; Barro-Bianchi 2023; Bianchi-Faccini-Melosi 2023). The surge in &amp;lsquo;unfunded&amp;rsquo; government spending during and after the COVID pandemic was financed by the inflation that eroded the real value of nominally-denominated government bonds. This constitutes a state-contingent capital levy on bondholders. A cautionary note is added: the availability of such a mechanism may encourage excessive spending, analogous to Ricardo&amp;rsquo;s (1820) argument for balanced-budget war finance.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-role-of-heterogeneity-among-households-in-potentially-generating-commitment"&gt;Q8. What is the role of heterogeneity among households in potentially generating commitment?&lt;/h3&gt;
&lt;p&gt;The paper discusses two sources. First, drawing on Broner-Martin-Ventura (2010), if the government cares about domestic holders of its bonds but not foreign holders, and if bonds can be traded on secondary markets so the two groups cannot be separated, then default becomes unattractive ex post because it harms domestic residents. This gives the government an incentive to promote secondary markets as a commitment device against sovereign default — potentially extensible to capital taxation commitments. Second, the distinction between old and new capital (e.g., via investment tax credits) partially limits the attractiveness of high capital-income taxes by tying the tax rate on old capital to the rate on new capital, which creates investment disincentives. However, as Straub-Werning demonstrate, this commitment may be too weak to drive the optimal capital-income tax to zero.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-calibration-targets-and-preference-specifications-used-in-the-quantitative-experiments"&gt;Q9. What are the calibration targets and preference specifications used in the quantitative experiments?&lt;/h3&gt;
&lt;p&gt;The model is calibrated to represent the U.S. economy with: government consumption = 20% of output, capital-income tax rate = 38% (from Barro-Furman 2018), public debt = 70% of output, labor fraction of time endowment = 1/3, discount factor β = 0.97 (3% real interest rate), capital share α = 0.34, depreciation δ = 0.08. Three preference specifications are explored: (1) standard preferences (time-separable, separable, homothetic in c and n); (2) balanced-growth preferences with consumption-leisure Cobb-Douglas aggregator and IES = 0.5; (3) zero-wealth-effect preferences. The wealth constraint W̃_0 is set to match the pre-reform steady-state wealth in utility terms.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-detailed-quantitative-results-across-preference-specifications"&gt;Q10. What are the detailed quantitative results across preference specifications?&lt;/h3&gt;
&lt;p&gt;Under standard preferences: capital-income tax rate is always exactly zero, labor-income tax rate is constant, welfare gain = 0.82% of steady-state consumption. Under balanced-growth preferences (IES = 0.5): initial capital-income tax ≈ 1%, quickly drops to zero; capital stock rises ≈ 12% to new SS; government debt rises ≈ 6 pp; labor-income tax ≈ 30% (constant, ≈ 4 pp above old SS of 26%); welfare gain = 0.76%; steady-state public debt under zero-capital-tax policy = 33% of output; initial capital levy l_0 = 0.126; new SS labor tax = 0.297. Under zero-wealth-effect preferences: initial capital-income tax ≈ 7%, drops sharply; welfare gain = 0.62%; l_0 = 0.160; new SS labor tax = 0.301; maximum capital tax rate = 0.070. Under extreme initial conditions (balanced-growth, capital stock at half SS level, debt at twice normal ratio): capital-income tax ≈ 3% initially, approaches zero; l_0 = 0.033; new SS labor tax = 0.400. Across all cases, constraining capital-income tax to zero with constant labor tax yields welfare nearly identical to the unconstrained Ramsey optimum.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-scope-of-the-zero-capital-tax-result-and-what-preference-conditions-support-it"&gt;Q11. What is the scope of the zero-capital-tax result and what preference conditions support it?&lt;/h3&gt;
&lt;p&gt;The zero-capital-tax result holds exactly under standard preferences (time-separable, separable between consumption and labor, and homothetic in consumption and labor), which satisfy the Diamond-Mirrlees-Sandmo-Sadka conditions for uniform taxation of goods. Under balanced-growth preferences, it holds with σ = 1 but not necessarily when σ ≠ 1. Under zero-wealth-effect preferences it does not hold if V is strictly concave. However, the quantitative experiments show that deviations from zero are small and short-lived under all three specifications, so zero capital taxation is approximately optimal across the board.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-relationship-between-the-papers-results-and-tax-rate-smoothing-models"&gt;Q12. What is the relationship between the paper&amp;rsquo;s results and tax-rate smoothing models?&lt;/h3&gt;
&lt;p&gt;Barro (1979, 1990) showed that optimal income-tax rates follow a random walk when capital levies are unavailable. The present paper shows that, once state-contingent capital levies are available (the Lucas-Stokey stochastic extension), consumption and labor-income tax rates also exhibit random-walk-like behavior, as realizations of spending and technology shocks move the optimal tax rates. This provides a unified framework connecting capital levy theory and tax-rate smoothing.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-survivalinstitutional-arguments-for-why-commitment-constraints-might-exist-in-practice"&gt;Q13. What are the survival/institutional arguments for why commitment constraints might exist in practice?&lt;/h3&gt;
&lt;p&gt;The paper suggests a selection argument: societies that fail to maintain commitments of the form W_t ≥ W̃_t severely under-accumulate capital because anticipating capital levies causes households and firms not to invest, potentially causing the economy to effectively disappear. This selection pressure may explain why functioning market economies tend to develop institutions (constitutions, property rights, secondary markets) that approximate the required commitments. Major regime changes, such as the Bolshevik revolution (100% default on Czarist bonds), can destroy these commitments, but many regime changes (e.g., France after World War II) do not fully repudiate prior obligations.&lt;/p&gt;
&lt;h3 id="q14-how-does-this-paper-relate-to-and-differ-from-the-three-main-antecedents-chamley-judd-straub-werning-and-chari-et-al-2020"&gt;Q14. How does this paper relate to and differ from the three main antecedents (Chamley-Judd, Straub-Werning, and Chari et al. 2020)?&lt;/h3&gt;
&lt;p&gt;Chamley (1986) and Judd (1985, 1999) showed zero long-run capital-income tax is optimal under the Ramsey formulation with l_0 = 0 and τ_t^k ≤ 1. Straub-Werning (2020) showed that positive capital-income taxes can be optimal even in the steady state under the same constraints when the IES is below one. Chari et al. (2020) replaced instrument restrictions with a utility-wealth constraint for period zero, obtaining a direct capital levy in period zero plus zero capital-income taxes thereafter. Barro-Chari extend Chari et al.&amp;rsquo;s period-zero constraint to all periods, achieving time-consistency and removing period zero&amp;rsquo;s special status. The novel contribution is the multi-period, time-consistent version of the Chari et al. framework and the quantitative demonstration that zero capital taxation is approximately optimal across preference specifications.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Period-zero problem&lt;/strong&gt;: The asymmetry in the standard Ramsey formulation in which the period-zero policymaker can commit to all future tax rates but faces no commitments from the past, creating a strong incentive to expropriate existing assets via capital levies (direct or indirect); the paper&amp;rsquo;s central target of critique.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital levy&lt;/strong&gt;: A proportional confiscation of asset holdings (l_t), distinct from ongoing taxes on the flow of asset income; a direct capital levy takes a fraction of the stock outright, while indirect capital levies are engineered through high asset-income tax rates or time-varying consumption taxes that reduce the real value of existing wealth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wealth constraint in utility units (W_t ≥ W̃_t)&lt;/strong&gt;: A commitment device, following Chari-Nicolini-Teles (2020) and Armenter (2008), that requires each period&amp;rsquo;s policymaker to leave households with at least a threshold level of wealth measured in units of utility rather than goods; instrumental in eliminating the period-zero problem without directly restricting tax instruments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Timeless perspective&lt;/strong&gt;: Woodford&amp;rsquo;s (1999) principle that the policymaker should adopt the behavior that would have been committed to far in the past contingent on current events, rather than optimizing from the current period taking past expectations as given; the paper shows its Ramsey results conform to this principle once per-period wealth constraints are imposed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time-consistency (in optimal taxation)&lt;/strong&gt;: The property that a tax plan chosen at date 0 will be voluntarily continued by each subsequent policymaker; fails in the Chari et al. (2020) baseline formulation when future policymakers are unconstrained because each will want to re-impose a &amp;lsquo;period-zero&amp;rsquo; capital levy, achieved here only when per-period wealth constraints W_t ≥ W̃_t are sufficient to deter direct levies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Indirect capital levy&lt;/strong&gt;: The engineering of a de facto reduction in the real value of existing wealth through policy instruments other than a direct asset levy — specifically positive tax rates on future asset income (τ_t^k &amp;gt; 0) or non-constant consumption tax rates that alter the present value of after-tax consumption; the mechanism underlying both Chamley-Judd transitional dynamics and Straub-Werning permanent positive capital taxes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Standard preferences&lt;/strong&gt;: Preferences that are time-separable, separable between consumption and labor, and homothetic in consumption and labor (Eq. 1 in the paper: u(c,n) = [c^{1-σ}/(1-σ)] − η·n^{1+Ψ}); the class under which uniform taxation of consumption at all dates and zero tax rates on asset income are exactly optimal, satisfying Diamond-Mirrlees-Sandmo-Sadka conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;State-contingent capital levy&lt;/strong&gt;: In the stochastic extension (following Lucas-Stokey 1983), a capital levy whose magnitude depends on the realized state of the world (e.g., war, pandemic, financial crisis); optimal under emergencies when emergency government spending must be financed, and below average during normal times — the paper interprets post-2020 U.S. inflation as an implicit state-contingent levy on nominal government bonds via the fiscal theory of the price level.&lt;/p&gt;</description></item><item><title>The (In)effectiveness of Targeted Payroll Tax Reductions</title><link>https://macropaperwarehouse.com/papers/the-ineffectiveness-of-targeted-payroll-tax-reductions/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-ineffectiveness-of-targeted-payroll-tax-reductions/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper studies the cost-effectiveness of targeted payroll tax reductions as a tool for stimulating labor demand among marginalized workers, using a natural experiment from Italy. The motivation is policy-relevant: governments routinely deploy targeted payroll tax cuts to combat youth and low-skill unemployment, but such subsidies risk subsidizing inframarginal hiring — employment that would have occurred without the incentive — rather than creating net new jobs. Rigorous evaluation requires two features that are rarely satisfied simultaneously: (1) the subsidy must target genuinely marginalized workers so estimates pertain to the population of interest, and (2) variation in incentives across firms must be quasi-random so firm responses are causally identified. This paper exploits a policy that satisfies both.&lt;/p&gt;
&lt;p&gt;The data are confidential matched employer-employee records from the Italian Social Security Institute (INPS), covering the universe of private non-agricultural firms with at least one employee from January 2003 to December 2009. The main analysis sample comprises 1,015,619 firms with policy-relevant firm size between 3 and 15 employees — the stratum containing the policy threshold. The study period spans 84 months.&lt;/p&gt;
&lt;p&gt;The policy variation is the Italian 2007 Budget Bill (Law 296/2006), which raised employer social security contributions (SSCs) on apprenticeship contracts from a flat rate of 148 euros per year to 10 percent of annual earnings (approximately 1,200 euros per year for an average apprentice earning 12,000 euros). However, firms with at most 9 full-time-equivalent employees (excluding apprentices) received a graduated discount: 1.5 percent of earnings in the first year (180 euros) and 3 percent in the second year (360 euros). This generated a clean discontinuity in incentives at the 9-employee threshold. The discount is equivalent to roughly two months of earnings per apprentice, or about 8 percent of the cost of a typical 19-month apprenticeship.&lt;/p&gt;
&lt;p&gt;The empirical strategy is a difference-in-discontinuities design. For each calendar month, the authors estimate a regression discontinuity specification comparing firms just above and just below the 9-employee threshold, then subtract the estimated baseline discontinuity from January 2006 (before the policy existed). This normalizes away pre-existing size-related differences in outcomes, yielding reduced-form estimates of how the policy-induced difference in SSC costs between small and large firms changed over time. The policy variation is used as an instrument for actual SSC payments to compute IV estimates of jobs supported per euro of foregone revenue.&lt;/p&gt;
&lt;p&gt;The main finding is a precise zero: the SSC discount does not increase the number of apprenticeship contracts. The reduced-form estimates of the policy&amp;rsquo;s effect on apprentice hiring are not statistically different from zero and are tightly estimated. Firms below the threshold pay approximately 25 euros less per month in SSCs than firms above, confirming the policy has fiscal bite (first-stage F-statistic = 230), but this differential generates no detectable behavioral response in employment.&lt;/p&gt;
&lt;p&gt;The policy also does not increase the rate at which apprentices are converted to permanent contracts (&amp;ldquo;transformations&amp;rdquo;). Firms do not adjust apprentice wages, do not substitute toward other contract types, do not churn through more apprentices, do not re-label existing contracts, and do not lower hiring standards for apprentices.&lt;/p&gt;
&lt;p&gt;For cost-effectiveness, the IV estimates imply that each 1 million euros of foregone SSC revenue supports the employment of 29 apprentices for one year — a point estimate not statistically different from zero. The point estimate for supported permanent-contract transformations is negative (point estimate: -2), also indistinguishable from zero. By comparison, directly hiring apprentices at their prevailing wage of 1,050 euros per month would employ 79 apprentices per million euros, making direct hiring 2.7 times more cost-effective than the subsidy. The paper surveys the broader literature and finds that once existing studies&amp;rsquo; employment effects are normalized against fiscal costs, targeted subsidies rarely appear cost-effective; hiring credits that require a new hire may outperform payroll tax cuts because they are harder to claim for inframarginal employment.&lt;/p&gt;
&lt;p&gt;The underlying mechanism is inelastic labor demand for apprentices. Survey evidence from the RIL firm survey confirms that when firms do not hire apprentices, cost is rarely the stated reason — the most common answer is that they do not need more people. When firms do hire apprentices, the most common reason is to provide training before converting them to permanent employees, not to economize on labor costs.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The identification strategy is a difference-in-discontinuities design. In each month, a regression discontinuity (RD) specification compares firms just above and just below the 9-employee SSC eligibility threshold; the authors then subtract the baseline (January 2006, pre-policy) discontinuity estimate to remove pre-existing size-related level differences. The key identifying assumption is a &amp;lsquo;weak parallel trends&amp;rsquo; assumption: the curvature of the conditional expectation function of untreated potential outcomes at the threshold is time-invariant. Threats and the evidence against them: (1) Manipulation of firm size at the threshold — addressed by showing that the CDF of policy-relevant firm size is virtually identical across all 84 months with no bunching at 9 employees before or after the reform; (2) Pre-existing trends — no pre-trends are found in the estimated discontinuity in outcomes for the four years before January 2007; (3) Compositional shifts — covariate balance tests show that firm characteristics (age, type, industry, region) at the threshold do not change over time relative to baseline; the covariate index (predicted apprentice hiring based on time-invariant firm characteristics) fluctuates between -0.0005 and +0.0005 — nearly two orders of magnitude smaller than the employment estimates; (4) Imperfect compliance — handled explicitly: the design estimates an intention-to-treat effect, which is attenuated relative to the treatment on the treated; (5) Measurement error in running variable — addressed by excluding firms within one unit of the threshold in the preferred specification; null results are robust to varying the exclusion window.&lt;/p&gt;
&lt;h3 id="q2-why-is-the-difference-in-discontinuities-design-superior-to-a-standard-difference-in-differences-design-in-this-context"&gt;Q2. Why is the difference-in-discontinuities design superior to a standard difference-in-differences design in this context?&lt;/h3&gt;
&lt;p&gt;The paper provides a formal and empirical case that standard difference-in-differences applied to a continuous firm-size running variable produces spurious results. When the conditional expectation function of outcomes with respect to firm size rotates over time (i.e., the slope changes), a DiD estimator that discretizes firms into treated and control groups will detect this rotation as a treatment effect, even if the true policy effect is zero. This is because the DiD constrains the slopes of the conditional expectation function above and below the threshold to be zero, making them implicit omitted variables. In the Italian data, the conditional expectation function of apprentice hiring with respect to firm size rotates clockwise between 2007 and 2009, coinciding with a general slowdown in hiring during the Great Recession. This rotation would cause a naive DiD analysis to conclude, spuriously, that the subsidy supported hiring. The difference-in-discontinuities design controls flexibly for the running variable in each period and isolates only the variation near the threshold, where firm size cannot proxy for trends unrelated to the policy.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-main-mechanisms-considered-for-why-the-subsidy-has-no-employment-effect-and-how-does-the-paper-distinguish-among-them"&gt;Q3. What are the main mechanisms considered for why the subsidy has no employment effect, and how does the paper distinguish among them?&lt;/h3&gt;
&lt;p&gt;The paper considers and rules out seven alternative explanations before concluding that demand for apprentices is simply inelastic: (1) Measurement error — ruled out because the null holds across specifications with different exclusion windows, and measurement error does not prevent finding significant effects on fiscal outcomes; (2) Subsidy too small — ruled out because the 8% subsidy (960 euros per apprentice per year, up to 1,460 euros at the 95th percentile of earnings) is comparable in magnitude to subsidies that generate large employment effects in Cahuc et al. (2019) and Guo (2024); (3) Low awareness — ruled out because 80% of eligible firms that hire apprentices receive the discount, confirming they must claim it actively; (4) Firms restricting hiring to maintain eligibility — ruled out because apprentices are excluded from policy-relevant firm size, so hiring an apprentice does not risk crossing the threshold; the firm-size distribution also remains stable; (5) Temporary nature of subsidy — ruled out because most apprenticeships last 19 months and the subsidy covers the first two years; moreover, the literature suggests temporary subsidies should be at least as effective as permanent ones; (6) Training requirements — ruled out because training requirements are poorly enforced, and no effects are found even among firms that previously employed apprentices (lower marginal training costs) or firms that rarely cite training costs as a deterrent; (7) Great Recession — ruled out because no effects appear in the year before the recession began, and effects are not larger or smaller for liquidity-constrained firms.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-analyses-are-conducted-and-what-do-they-show"&gt;Q4. What heterogeneity analyses are conducted and what do they show?&lt;/h3&gt;
&lt;p&gt;The authors estimate pooled post-reform difference-in-discontinuities coefficients separately across multiple dimensions and find consistently null effects with no evidence of heterogeneous treatment effects: (1) by industry — estimates across manufacturing, transportation and construction, trading, services, and other sectors are all tightly centered on zero; (2) by region — null across all Italian regions; (3) by baseline apprentice earnings quartile — null across Q1 through Q4 and for firms with no apprentices at baseline; (4) by contemporaneous apprentice earnings quartile — null; (5) by three measures of liquidity constraints (liquid assets to total assets, cash flow to total assets, revenues above/below median) — null in all six groups; and (6) by prior apprenticeship training status — null for both firms that employed at least one apprentice in 2006 and those that did not. The authors note the scope condition: estimates are internally valid for firms in a neighborhood of 9 employees, and effects for substantially larger firms cannot be ruled out to differ.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-conducted-beyond-the-main-heterogeneity-analysis"&gt;Q5. What robustness checks are conducted beyond the main heterogeneity analysis?&lt;/h3&gt;
&lt;p&gt;The main robustness checks are: (1) sensitivity of apprentice hiring effects to the amount of excluded data around the threshold (the &amp;lsquo;donut bandwidth&amp;rsquo;) — the null holds across all exclusion windows (Appendix Figure A.2); (2) placebo tests using the pre-reform periods (January 2003 through December 2006) — no pre-trends in the estimated discontinuity for any outcome; (3) covariate stability tests — the discontinuity in a covariate index predicting apprentice hiring from time-invariant firm characteristics shows no change over time, with point estimates between -0.0005 and +0.0005 versus employment estimates between -0.01 and +0.01; (4) comparison of results to a standard DiD specification — the DiD produces spurious positive effects driven by rotation of the conditional expectation function, while the difference-in-discontinuities estimate remains precisely zero; (5) examination of other outcomes (contract churn, re-labeling, worker quality, contract type substitution, temporary worker stocks) — all null.&lt;/p&gt;
&lt;h3 id="q6-how-is-cost-effectiveness-formally-measured-and-what-does-the-iv-estimate-imply"&gt;Q6. How is cost-effectiveness formally measured and what does the IV estimate imply?&lt;/h3&gt;
&lt;p&gt;Cost-effectiveness is defined as the number of jobs supported per unit of foregone revenue: omega = E[L(1) - L(0)] / E[R(0) - R(1)], where L is employment and R is tax payments. Rather than back-of-the-envelope calculation, the authors estimate this with 2SLS, instrumenting for actual SSC payments with the interaction of being below the eligibility threshold and the post-2007 indicator. This allows them to compute standard errors, which back-of-the-envelope methods do not provide. The first-stage F-statistic is 230, confirming instrument strength. Point estimates from Table 4: 29 apprentice-years supported per 1 million euros of foregone SSC (standard error 58, not significant); 647,237 euros of apprentice compensation supported per 1 million euros (standard error 921,320, not significant); and -2 permanent-contract transformations per 1 million euros (standard error 21, not significant). For context, directly hiring apprentices at 1,050 euros per month would generate 79 apprentice-years per million euros — 2.7 times more than the point estimate from the subsidy.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-paper-benchmark-its-cost-effectiveness-estimates-against-the-broader-literature"&gt;Q7. How does the paper benchmark its cost-effectiveness estimates against the broader literature?&lt;/h3&gt;
&lt;p&gt;The authors normalize employment effects from nine other studies against their fiscal costs to produce a common metric of jobs or job-years per 1 million dollars of foregone revenue. The studies span payroll tax cuts (Egebark and Kaunitz 2013; Saez, Schoefer, and Seim 2021), hiring credits (Cahuc, Carcillo, and Le Barbanchon 2019; Neumark 2013), and fiscal stimulus programs (Bartik 2001; Bartik and Erickcek 2010; Dupor and Mehkari 2016; Dupor and McCrory 2018; Feyrer and Sacerdote 2011; Wilson 2012). The conclusion is that most wage subsidies, including those that generate positive reduced-form employment effects, produce very high costs per job. With two exceptions (Bartik 2001 and Cahuc et al. 2019), cost-effectiveness estimates across the literature are extremely low. The paper argues that hiring credits may be more cost-effective than payroll tax cuts because the requirement to make a new hire makes it harder to subsidize inframarginal employment. Importantly, the Italian study&amp;rsquo;s cost-effectiveness estimates — though imprecisely estimated — are broadly consistent with the cross-study pattern once fiscal costs are accounted for.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-welfare-and-public-finance-implications-of-the-null-employment-effects"&gt;Q8. What are the welfare and public finance implications of the null employment effects?&lt;/h3&gt;
&lt;p&gt;Because the behavioral response is zero and the fiscal cost is non-zero, the policy functions as a pure transfer from the government to firms. The paper invokes the framework of Hendren and Sprung-Keyser (2020) to note that the marginal value of public funds is essentially 1 — there is no distortion introduced but also no welfare gain from resource reallocation. This interpretation cuts in two directions: (1) the pre-reform apprentice SSC subsidies (which were larger than the post-2007 discount) were also essentially transfers with large fiscal costs and no employment-creation value; and (2) the SSC increase imposed on larger firms (those with more than 9 employees) effectively raised revenue without causing meaningful employment losses, since labor demand for apprentices is inelastic. The policy is thus deemed inefficient in the sense that taxpayer revenue is lost without generating the intended social return of increasing employment of marginalized workers.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-scope-conditions-and-limitations-of-the-estimates"&gt;Q9. What are the scope conditions and limitations of the estimates?&lt;/h3&gt;
&lt;p&gt;The difference-in-discontinuities design provides internally valid estimates only for firms in a neighborhood of 9 employees, which in Italy means firms with 3 to 15 employees (90% of Italian firms and 65% of all apprentices). The paper cannot rule out that larger firms respond differently to similar subsidies. The analysis is partial equilibrium: it cannot measure spillovers, general equilibrium effects on wage-setting across the firm-size distribution, or displacement effects between firms. Cost-effectiveness estimates reflect only the direct fiscal cost of foregone SSCs and do not include fiscal externalities (e.g., effects on income tax revenues or social insurance outlays) or administrative and political costs. The exclusion of workers from the public sector means the results pertain solely to private-sector apprenticeships.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-prior-studies-on-payroll-tax-cuts-and-what-distinguishes-it-methodologically"&gt;Q10. How does this paper relate to prior studies on payroll tax cuts, and what distinguishes it methodologically?&lt;/h3&gt;
&lt;p&gt;Prior national studies (e.g., Saez et al. 2019, 2012, 2021; Egebark and Kaunitz 2013; Huttunen et al. 2013; Bozio et al. 2020; Rubolino 2021) estimate labor demand responses by comparing employment of targeted versus untargeted workers, which can overstate policy effectiveness if firms substitute targeted for untargeted workers (a SUTVA violation that would not be detected by parallel pre-trend tests). Cross-regional studies (e.g., Bennmarker et al. 2009; Benzarti and Harju 2021a; Bohm and Lind 1993; Guo 2024) study firms but typically do not target genuinely marginalized workers, so estimates reflect average rather than marginal labor demand. This paper satisfies both requirements simultaneously: the discontinuity in incentives provides quasi-random variation across firms (avoiding SUTVA), and the policy specifically targets apprentices — a non-random, marginalized group — so the estimated elasticities pertain to the actual population of interest. The paper is also the first (to the authors&amp;rsquo; knowledge) to use a formal IV strategy to estimate cost-effectiveness with standard errors, enabling statistical precision comparisons across the distribution of estimates.&lt;/p&gt;
&lt;h3 id="q11-what-does-survey-evidence-from-the-ril-data-contribute-to-the-interpretation"&gt;Q11. What does survey evidence from the RIL data contribute to the interpretation?&lt;/h3&gt;
&lt;p&gt;The RIL (Rilevazione Longitudinale su Imprese e Lavoro), a representative firm survey collected in 2005, provides direct evidence on firms&amp;rsquo; stated reasons for their apprenticeship hiring decisions. Among firms that do not hire apprentices, the most common reason by far is &amp;lsquo;we don&amp;rsquo;t need more people,&amp;rsquo; with cost cited rarely. Among firms that do hire apprentices, the dominant reason is to train workers prior to hiring them as permanent employees; &amp;rsquo;lower labor costs&amp;rsquo; is a secondary consideration. This corroborates the paper&amp;rsquo;s interpretation that demand for apprentices is driven by training-for-retention motives rather than cost arbitrage, which explains why a cost reduction leaves hiring behavior unchanged.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-policy-recommendation-and-its-scope"&gt;Q12. What is the policy recommendation and its scope?&lt;/h3&gt;
&lt;p&gt;The paper urges caution in using payroll tax credits to stimulate employment, particularly for targeted groups with inherently low or inelastic labor demand. The results suggest that, for apprentices, firms hire based on training-and-conversion needs rather than cost considerations, so subsidizing cost does not expand hiring. More broadly, the cross-study cost-effectiveness comparison suggests that hiring credits — which require a new hire as a prerequisite for receiving the subsidy — may be more efficient than payroll tax cuts precisely because they screen out inframarginal firms. The paper does not rule out effectiveness for other worker types or for much larger subsidies, but the documented uniformity of null effects across industries, regions, and firm types suggests the inelasticity finding is robust within the studied population.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Inframarginal hiring&lt;/strong&gt;: Employment that would occur absent the subsidy; when a policy subsidizes inframarginal hiring, it transfers resources to firms without generating net new jobs, making it fiscally costly but behaviorally inert.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Difference-in-discontinuities&lt;/strong&gt;: An empirical design that combines regression discontinuity with difference-in-differences: in each period a discontinuity at the policy threshold is estimated, and the pre-policy baseline discontinuity is subtracted to remove pre-existing size-related level differences and time-invariant non-linearities in the conditional expectation function.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy-relevant firm size&lt;/strong&gt;: As defined by INPS under the 2007 Budget Bill: total full-time equivalent employment minus apprentices, temporary agency workers, workers on leave (unless replaced), and workers on specific on-the-job training contracts; this is the running variable determining SSC eligibility.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cost-effectiveness (jobs per foregone revenue)&lt;/strong&gt;: The number of job-years supported per unit of foregone tax revenue (here, per 1 million euros of lost SSCs), formally estimated via instrumental variables to allow statistical inference — as opposed to back-of-the-envelope calculations that provide no standard errors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inelastic labor demand for apprentices&lt;/strong&gt;: In this paper&amp;rsquo;s sense: firms&amp;rsquo; demand for apprenticeship contracts does not respond to changes in their labor cost, because hiring decisions are driven by training-and-conversion motives (hiring to eventually retain as permanent employees) rather than by cost minimization at the margin.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rotation of the conditional expectation function&lt;/strong&gt;: A change over time in the slope of the relationship between an outcome (e.g., apprentice hiring) and the running variable (firm size); when the slope changes, standard DiD specifications that discretize firms into treated/control groups will spuriously detect a treatment effect even when the true policy effect is zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Transformation (apprentice to permanent contract)&lt;/strong&gt;: The event of a firm converting an existing apprenticeship contract into an open-ended (permanent) employment contract at the end of the apprenticeship; used as an alternative outcome to evaluate whether the subsidy increased the ultimate goal of permanent employment, not just temporary apprenticeships.&lt;/p&gt;</description></item><item><title>The Unequal Costs of Carbon Pricing: Economic and Political Effects Across European Regions</title><link>https://macropaperwarehouse.com/papers/the-unequal-costs-of-carbon-pricing-economic-and-political-effects-across-european-regions/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-unequal-costs-of-carbon-pricing-economic-and-political-effects-across-european-regions/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether carbon pricing through the EU Emissions Trading System (EU ETS) imposes economic costs that are unequally distributed across European regions, and whether those economic costs translate into political costs in the form of votes for extremist and populist parties. The motivation is both practical — political opposition has blocked or rolled back climate policies in several countries — and analytical: no prior study had systematically estimated the political consequences of carbon pricing at the subnational level.&lt;/p&gt;
&lt;p&gt;The authors build a panel dataset covering 224 NUTS2 regions from 20 European countries (covering 97% of EU GDP, plus Norway) over 2000–2019. Economic data come from the European Commission&amp;rsquo;s ARDECO database; emission data from EDGAR (aggregate GHG) and the EU ETS Transaction Log (verified ETS emissions from regulated installations, mapped to NUTS2 via zip codes); voting data from the EU-NED dataset with party classifications from The PopuList. Household expectations are measured from 34 Eurobarometer survey waves (2004–2019). The dataset spans 114 elections (110 national, four European Parliament).&lt;/p&gt;
&lt;p&gt;Identification rests on the carbon policy shocks of Kanzig (2023), constructed from high-frequency movements in EU carbon allowance futures prices around 126 regulatory events between 2005 and 2019, instrumented in a monthly VAR and aggregated to annual frequency. These shocks are orthogonal to contemporaneous economic conditions by construction, and are normalized so that the on-impact effect equals a 1% rise in Euro Area HICP energy prices. The main estimator is Jorda (2005) local projections in a panel with region fixed effects, lagged controls, and Driscoll-Kraay standard errors, estimated over a four-year horizon.&lt;/p&gt;
&lt;p&gt;Main economic findings (average region): A 1%-energy-price-equivalent carbon shock reduces real GDP by approximately 0.7% — a contraction that persists for four years. Employment, real net disposable household income, real GVA, real compensation, real investment, and hours worked all decline significantly and persistently. GHG emissions fall by roughly 1% one year after the shock, confirming the policy&amp;rsquo;s effectiveness.&lt;/p&gt;
&lt;p&gt;Main political findings: The combined extremist vote share (far-left plus far-right) rises by 0.3 to 0.4 percentage points two years after the shock and remains elevated. Populist and Eurosceptic vote shares also rise significantly in the medium term. Political fragmentation (1 minus the HHI) increases persistently. The shift is primarily toward far-right parties.&lt;/p&gt;
&lt;p&gt;Survey-based expectations: The share of respondents citing environmental issues as a top concern falls by approximately 2 percentage points and remains depressed for four years. Respondents become significantly more pessimistic about national economic and employment prospects and their own financial situation.&lt;/p&gt;
&lt;p&gt;Role of the economic channel: Using the Holm-Paul-Tischbirek (2021) decomposition, up to two thirds of the total rise in the extremist vote share over the four-year horizon is attributed to the decline in GDP, employment, and household income. The first year is more dominated by non-economic attribution effects (roughly 25% of the effect is explained by the economic channel at h=1), consistent with voters initially blaming the government&amp;rsquo;s policy choice rather than responding to realized economic deterioration.&lt;/p&gt;
&lt;p&gt;Regional heterogeneity and inequality: Regions one standard deviation above mean ETS emission intensity experience a meaningfully larger output contraction and a 20–50% larger and more persistent rise in the extremist vote share relative to the average region. Regions receiving fewer free ETS allowances face analogously larger economic and political costs. The within-country 90–10 ratio of real disposable household income rises by approximately 0.05 percentage points, with widening concentrated at the lower tail (the median-to-10th-percentile gap), meaning poorer regions bear disproportionate costs. These heterogeneous effects imply that carbon pricing contributes to regional inequality within countries.&lt;/p&gt;
&lt;p&gt;Policy implication: The EU ETS lacks direct redistribution mechanisms. The authors argue that progressive revenue recycling — household rebates calibrated to income — is necessary to cushion vulnerable regions, limit inequality, and rebuild public support for climate policy. These concerns are especially pressing given the EU ETS&amp;rsquo;s scheduled expansion to buildings and transportation in 2027.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The key identifying assumption is that the carbon policy shocks of Kanzig (2023) are exogenous with respect to regional economic conditions. The shocks are constructed from high-frequency daily movements in EU carbon allowance futures prices on days of regulatory announcements, relative to wholesale electricity prices on the prior day; the narrow event window ensures that confounding macroeconomic factors are already priced in. The shocks are then instrumented in a monthly VAR to extract structural shocks with a higher signal-to-noise ratio before being aggregated to annual frequency. The main threat would be if major regulatory announcements coincidentally coincided with other economic news. The authors defend against this by showing robustness to controlling for unemployment, stock market indices, monetary policy rates, oil prices, and a global financial crisis dummy. For the heterogeneity analysis, ETS intensity and free allowance share are fixed at their pre-sample values (end of ETS pilot phase, 2008) to rule out reverse causality from carbon pricing to the exposure measures.&lt;/p&gt;
&lt;h3 id="q2-how-is-the-economic-voting-channel-distinguished-empirically-from-other-channels"&gt;Q2. How is the economic voting channel distinguished empirically from other channels?&lt;/h3&gt;
&lt;p&gt;The authors use the decomposition approach of Holm, Paul, and Tischbirek (2021). They re-estimate the extremist vote share local projection while controlling for the contemporaneous path of GDP, employment, and household income over the same h-year horizon. The residual coefficient on the carbon shock captures voting effects not attributable to economic deterioration. Comparing the controlled and uncontrolled responses shows that over the full four-year horizon, roughly two thirds of the voting increase is explained by economic variables. In the first year, the economic channel explains only about 25% of the response, consistent with non-economic attribution effects — voters blaming a government policy choice rather than an exogenous shock — being more prominent early on.&lt;/p&gt;
&lt;h3 id="q3-what-additional-evidence-distinguishes-ets-driven-political-effects-from-other-energy-price-effects"&gt;Q3. What additional evidence distinguishes ETS-driven political effects from other energy price effects?&lt;/h3&gt;
&lt;p&gt;Two benchmarks are used. First, national carbon taxes, which prior literature shows have muted economic effects, produce no statistically significant response in either real GDP or the extremist vote share (Appendix A.2), consistent with the economic channel being essential for the political response. Second, oil supply news shocks (Kanzig, 2021), constructed with a comparable high-frequency methodology and producing a similarly sized GDP decline, generate a statistically significantly smaller increase in the extremist vote share over the first two years (Appendix A.3). The excess political response to carbon shocks over oil shocks is interpreted as reflecting voters attributing policy-driven economic pain to the government, analogously to Gabriel, Klein, and Pessoa (2023) finding that austerity-induced recessions elicit stronger political responses than general downturns.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-across-regions-is-documented-and-how-is-it-measured"&gt;Q4. What heterogeneity across regions is documented and how is it measured?&lt;/h3&gt;
&lt;p&gt;Two exposure dimensions are explored. First, ETS emission intensity (verified ETS emissions scaled by GDP) captures direct agglomeration of installations covered by the carbon market. Second, the share of freely allocated ETS allowances relative to verified emissions captures the effective carbon price faced by firms in the region. Regions one standard deviation above mean ETS intensity experience meaningfully larger output and employment contractions, and 20–50% larger and more persistent increases in the extremist vote share. Regions with fewer free allowances bear analogously larger costs. Results hold when GHG intensity (covering non-ETS sectors) replaces ETS intensity, and when sectoral composition is controlled in the free allowance analysis. A country-level inequality analysis using local projections on the 90–10 ratio of regional household income shows that carbon pricing raises within-country dispersion by approximately 0.05 percentage points, driven primarily by widening of the lower tail (50th to 10th percentile gap), indicating that poorer regions suffer most.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Vote share results are robust to: (a) excluding parties coded as borderline by The PopuList; (b) excluding European Parliament elections and using only national elections; (c) averaging national and European election outcomes in years when both occur; (d) a minimal control set of only lagged dependent variable and region fixed effects; (e) an expanded control set adding country-level unemployment rate, stock market index, monetary policy rate, Brent oil price, and a GFC dummy variable. The inequality results are robust to using the 75–25 ratio and the Gini coefficient in addition to the 90–10 ratio. The heterogeneity results are robust to including time fixed effects, which absorb the aggregate carbon shock but preserve cross-sectional variation, confirming that heterogeneous responses are not driven by aggregate confounders. Driscoll-Kraay standard errors are used throughout to allow for cross-sectional and serial dependence; clustering at region-year level delivers nearly identical results.&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q6. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Most directly related is Mangiante (2024), which documents that regions in poorer Euro Area countries are more exposed to carbon policy shocks. The present paper complements this by identifying within-country variation driven by ETS intensity and free allowance allocation, and by adding the political dimension. Kanzig and Konradt (2024) establish country-level economic effects of EU ETS shocks; this paper confirms those findings carry to the regional level and confirms comparable magnitudes. Gabriel, Klein, and Pessoa (2023) use the same econometric approach to study the political costs of austerity in European regions; the present paper finds analogous results for carbon pricing and attributes the political response similarly to economic deterioration. The finding that national carbon taxes lack economic or political bite echoes Metcalf and Stock (2023) and Konradt and Weder di Mauro (2023). The paper adds to the globalization-and-populism literature (Funke et al., 2016; Pastor and Veronesi, 2021; Colantone and Stanig, 2018) by identifying carbon pricing as another channel through which economic shocks drive extremist voting.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-direction-of-the-political-shift--toward-far-right-or-far-left"&gt;Q7. What is the direction of the political shift — toward far right or far left?&lt;/h3&gt;
&lt;p&gt;The decomposition in Appendix A.2 shows the increase in the combined extremist vote share is driven primarily by far-right parties. The far-right vote share rises significantly, while the far-left vote share shows a smaller and less precisely estimated increase. This is consistent with prior literature (Funke, Schularick, and Trebesch, 2016) documenting that far-right parties disproportionately benefit from recessions. A small decline in voter turnout is also documented, which may amplify measured increases in extremist vote shares by reducing the denominator (valid votes).&lt;/p&gt;
&lt;h3 id="q8-what-do-the-results-imply-for-environmental-concern-and-the-political-sustainability-of-climate-policy"&gt;Q8. What do the results imply for environmental concern and the political sustainability of climate policy?&lt;/h3&gt;
&lt;p&gt;Eurobarometer data show that the share of respondents ranking environmental issues among the two most important problems facing their country falls by approximately 2 percentage points following a carbon policy shock, a persistent decline lasting four years. The authors interpret this as a self-interest crowding-out effect: when carbon pricing imposes economic costs, concern for the environment is displaced by concern for living standards, consistent with Douenne and Fabre (2022). This creates a potential self-undermining dynamic: carbon pricing erodes the popular support needed to sustain and strengthen climate policy over time, particularly given that carbon-intensive regions — which suffer most economically — also see the largest decline in public support for environmental issues.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-scope-conditions-on-the-policy-implications"&gt;Q9. What are the scope conditions on the policy implications?&lt;/h3&gt;
&lt;p&gt;The findings pertain to ETS-style cap-and-trade pricing based on regulatory-driven supply restriction, not to national carbon taxes, which the paper shows have much smaller economic and political footprints. The sample covers 20 European countries with NUTS2 regional data over 2000–2019. The carbon policy shocks are derived from EU ETS regulatory events and are specific to that institutional context; generalization outside the EU ETS requires caution. Political effects operate primarily over a two-to-four-year horizon coinciding with electoral cycles. The paper&amp;rsquo;s redistribution prescription (progressive revenue recycling) presupposes a policy instrument capable of targeting household income; the EU ETS currently lacks such a mechanism, which is precisely the gap the authors flag as most urgent given the ETS expansion to buildings and transportation scheduled for 2027.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Carbon policy shock&lt;/strong&gt;: A series of exogenous regulatory surprises in EU ETS carbon allowance markets, constructed by Kanzig (2023) from high-frequency futures price movements around 126 regulatory events (2005–2019), instrumented in a monthly VAR, and normalized to produce a 1% on-impact increase in Euro Area HICP energy prices. Distinct from carbon price levels or oil shocks; isolates policy-driven changes in the supply of emission allowances, orthogonal to contemporaneous economic conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;ETS emission intensity&lt;/strong&gt;: Verified ETS emissions from regulated industrial installations in a NUTS2 region, scaled by regional GDP. The primary measure of a region&amp;rsquo;s direct exposure to EU carbon pricing; regions with higher ETS intensity experience larger economic contractions and larger shifts toward extremist parties when carbon prices rise.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Share of free allowances&lt;/strong&gt;: The ratio of freely allocated ETS emission permits to a region&amp;rsquo;s verified ETS emissions, used as a second regional exposure measure. A higher share implies a lower effective carbon price faced by firms; regions with fewer free allowances bear larger economic and political costs from carbon policy shocks. Free allowances were originally granted to protect energy- and trade-intensive sectors from rapid cost increases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extremist vote share&lt;/strong&gt;: The combined vote share of far-left and far-right parties in a region-election observation, using party classifications from The PopuList expert-coding database. The primary political outcome variable in the paper; empirically driven mainly by the far-right component in response to carbon policy shocks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Political fragmentation&lt;/strong&gt;: Defined in the paper as one minus the Herfindahl-Hirschman Index computed over all parties&amp;rsquo; vote shares in an election (1 − sum of squared vote shares). Captures the dispersion of votes across parties beyond the extremist vote share; used as a summary indicator of political polarization.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Economic voting channel&lt;/strong&gt;: The mechanism by which voters respond to carbon-pricing-induced economic deterioration — falling GDP, employment, and household income — by shifting support away from mainstream parties toward extremist alternatives. Isolated empirically via the Holm-Paul-Tischbirek (2021) decomposition; accounts for approximately two thirds of the total extremist voting response over the four-year impulse response horizon.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Regional inequality (90–10 ratio)&lt;/strong&gt;: Within-country dispersion of regional real disposable household income (or employee compensation) measured as the difference between the 90th and 10th percentile NUTS2 regions. Carbon pricing raises this measure persistently, with widening concentrated at the lower tail (the median-to-10th-percentile gap), indicating that poorer regions bear disproportionate economic costs.&lt;/p&gt;</description></item><item><title>The Winners and Losers of Climate Policies: A Sufficient Statistics Approach</title><link>https://macropaperwarehouse.com/papers/the-winners-and-losers-of-climate-policies-a-sufficient-statistics-approach/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-winners-and-losers-of-climate-policies-a-sufficient-statistics-approach/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks who wins and loses from climate policies — carbon taxes, renewable subsidies, and carbon tariffs — across 193 heterogeneous countries, and by how much. The motivation is that the standard IAM literature aggregates welfare into a global number, obscuring the distributional structure that determines political feasibility. Without knowing which countries gain and lose, and through which channels, it is impossible to understand why international cooperation is so difficult or which club structures can sustain themselves.&lt;/p&gt;
&lt;p&gt;The authors build a static Integrated Assessment Model (IAM) with heterogeneous countries, international trade in goods (Armington CES), international trade in fluid fossil (oil and gas), locally traded coal, and locally supplied renewables. Production uses a nested CES combining labour with a composite of three energy types. A reduced-form climate system maps world emissions linearly to global temperature, then to country-specific local temperatures, which damage TFP through a quadratic damage function. The key methodological contribution is a first-order (log-linear) decomposition of welfare around the current equilibrium, which expresses welfare changes analytically as a function of five observable sufficient statistics: (i) direct TFP damage, (ii) export terms-of-trade, (iii) import price index, (iv) energy cost effects (change in energy prices faced by producers), and (v) energy rent effects (change in profits of domestic fossil and renewable producers). This decomposition requires no model simulation; it reads off welfare directly from observables and a small set of elasticities.&lt;/p&gt;
&lt;p&gt;Two sets of structural parameters are estimated. First, a structural damage function is estimated using bilateral trade data from the ITPD-E dataset (2000–2016, 169 countries) via a Poisson pseudo-maximum-likelihood gravity regression that instruments temperature shocks against within-trading-partner variation in import penetration, controlling for energy market effects. The preferred specification recovers a global peak temperature of T* = 14.02°C and a damage slope parameter γ = 0.012. This strategy is designed to be robust to the Lucas critique: unlike reduced-form GDP regressions, it nets out general-equilibrium spillovers through trade and energy channels. Second, country-specific energy supply elasticities for oil-gas and coal are estimated from time-series variation in fossil rent shares and international prices (1985–2019 data), using OLS country-by-country and then an empirical Bayes shrinkage procedure with a truncated-normal prior that enforces positive elasticities. Coal is found to be substantially more elastically supplied than oil-gas; OPEC nations (e.g., Saudi Arabia) have near-inelastic oil-gas supply, while the US has relatively elastic supply.&lt;/p&gt;
&lt;p&gt;Key quantitative results from the policy experiments follow. (1) Business-as-usual: a 3°C warming by 2100 generates a 17% loss in consumption-equivalent world welfare under utilitarian weights, implying a Social Cost of Carbon of $203/tCO₂ at the current equilibrium point-of-approximation, rising to $302/tCO₂ if computed at 3°C of warming. Under Negishi (income-proportional) weights, the SCC falls to $3.31, reflecting that damages are concentrated in low-income countries with high marginal utility. Winners include Canada and Russia; losers are concentrated in Africa, Latin America, and South-East Asia. (2) Unilateral carbon tax (China, $50/tonne): global emissions rise by less than 0.07% (not fall) because China&amp;rsquo;s carbon tax shifts its energy mix from coal toward oil-gas (coal is ~1.44× dirtier per unit of energy), raising the international oil-gas price by approximately 5%, which boosts fossil exporters&amp;rsquo; rents and induces other countries to substitute back to coal. Global utilitarian welfare falls by 0.2%. China itself gains on net through falling coal prices and improved terms of trade. EU nations lose from higher energy import costs. (3) Unilateral carbon tax (USA, $50/tonne): global emissions fall by 0.8%; US welfare effects are small but positive (energy cost increases largely offset by terms-of-trade gains with Canada and Europe). (4) Renewable subsidies (42.6%, calibrated to produce the same average relative-price shift as a $50 carbon tax): on average substantially less effective than carbon taxation and more harmful to welfare because subsidies push countries up their upward-sloping domestic renewable supply curves, wasting resources on costly domestic generation (especially in countries with high baseline renewable shares such as France). (5) EU climate club ($50 carbon tax + CBAM tariffs): global emissions fall by 3%; global utilitarian welfare rises by around 5% (1% under Negishi weights), but the EU itself is a net loser — only Southern Europe (Spain, Portugal, Italy) gains; Germany and Scandinavian nations lose both from direct policy costs and from cooling that harms countries that benefit from warming. Oil-gas price falls by 4.6% within the club. (6) ASEAN climate club (same structure): global emissions fall by 0.5%; global utilitarian welfare rises by about 0.8% (0.2% Negishi); ASEAN members broadly benefit because they are already losers from climate change and the carbon-reduction benefit outweighs policy costs. Oil-gas price falls by 0.6%. (7) Global $50 carbon tax (all 193 countries): global emissions fall by 3.82%; global oil-gas price rises by 0.96% (substitution from coal toward oil-gas under a global carbon tax); global utilitarian welfare rises by about 6% (1% Negishi). Most of the utilitarian gain reflects reduced international inequality, since benefits concentrate in low-income tropical countries. Fossil exporters such as Saudi Arabia and Nigeria see energy rents rise as coal is substituted for by oil-gas globally.&lt;/p&gt;
&lt;p&gt;The central mechanism finding is that leakage operates primarily through energy trade, not goods trade: energy market effects are consistently larger than goods-market terms-of-trade effects across all policy experiments. This quantifies why unilateral climate policy is so limited in effectiveness. International coordination through climate clubs overcomes leakage but creates winners and losers within member coalitions depending on each member&amp;rsquo;s energy mix, trade exposure, and baseline climate damage.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-structural-damage-function-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy for the structural damage function and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The authors estimate the damage function using a Poisson pseudo-maximum-likelihood gravity regression on bilateral import penetration ratios (Xij/Xii) as a function of temperature differences between exporters and importers (and their squares), with country-pair fixed effects and year fixed effects. Controls for GDP/capita (polynomial), oil rent share, and renewable energy share proxy for the time-varying component of factory-gate prices driven by energy prices and wages. The key identifying assumption is that conditional on these controls and fixed effects, temperature shocks are uncorrelated with time-varying bilateral preference or cost shifters. Threats include: (1) confounding time-varying bilateral shocks correlated with temperature, such as ENSO events or specific geopolitical shocks; (2) the possibility that global (rather than local) temperature drives damages, which the paper cannot address given limited time-series variation and potential spurious correlation concerns (following Goulet Coulombe and Klieber, 2025); (3) the treatment of θ = 5 as a known parameter in computing γ from the regression coefficient, which propagates calibration error. The authors argue their strategy is robust to the Lucas critique because it nets out general-equilibrium effects on GDP that would contaminate GDP-based damage regressions.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-papers-welfare-decomposition-work-and-what-are-its-five-channels"&gt;Q2. How does the paper&amp;rsquo;s welfare decomposition work and what are its five channels?&lt;/h3&gt;
&lt;p&gt;The welfare decomposition is a first-order log-linearisation of the indirect utility around the current equilibrium. Changes in consumption-equivalent welfare for country i decompose into: (i) direct climate TFP damage (change in Dy_i); (ii) export terms-of-trade effect (change in domestic good price p_i); (iii) import price-index effect (change in price index P_i); (iv) energy cost effects (changes in oil-gas price q^f, coal price q^c_i, and renewable price q^r_i weighted by their shares in production); and (v) energy rent effects (changes in profits from fossil, coal, and renewable extraction weighted by their shares in household income). The key insight is that none of these five terms requires solving the full model; each can be computed from observable data moments (energy mix, energy rent shares, trade shares) and a small number of estimated or calibrated elasticities.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-in-climate-damages-is-documented-and-what-drives-it"&gt;Q3. What heterogeneity in climate damages is documented and what drives it?&lt;/h3&gt;
&lt;p&gt;Winners from climate change (3°C warming) are primarily cold countries: Canada, Russia, Scandinavian nations. Losers are concentrated in Africa (Djibouti, Niger, Burkina Faso, Sudan), Latin America, and South-East Asia. The heterogeneity arises from: (1) differences in baseline temperature relative to the estimated global peak productivity temperature T* = 14.02°C; countries hotter than T* lose productivity with further warming, while colder countries gain; (2) partial local adaptation (αT = 0.5) so each country&amp;rsquo;s effective peak temperature is halfway between T* and its current local temperature; (3) indirect effects through trade networks — cold, open economies can lose if major trading partners are damaged; (4) energy rent effects — fossil exporters lose energy rents as warming reduces global energy demand, partially offsetting their direct productivity gains.&lt;/p&gt;
&lt;h3 id="q4-why-does-chinas-unilateral-carbon-tax-at-50tonne-raise-global-emissions-rather-than-lower-them"&gt;Q4. Why does China&amp;rsquo;s unilateral carbon tax at $50/tonne raise global emissions rather than lower them?&lt;/h3&gt;
&lt;p&gt;China relies heavily on coal, which has a carbon concentration ratio of approximately ξc/ξf ≈ 1.44 (coal is ~44% dirtier per unit energy than oil-gas). A carbon tax on both fuels raises the effective cost of coal more than oil-gas, inducing China to substitute toward oil-gas imports. This raises the international oil-gas price by approximately 5%, which: (1) increases energy rents for fossil exporters (Gulf states, Russia) and (2) makes oil-gas costlier for other countries, incentivising them to substitute back toward coal. The net effect on global emissions is a slight increase of less than 0.07%, rather than a decline. This is the carbon leakage effect operating through energy trade.&lt;/p&gt;
&lt;h3 id="q5-why-are-renewable-subsidies-substantially-less-effective-than-carbon-taxes"&gt;Q5. Why are renewable subsidies substantially less effective than carbon taxes?&lt;/h3&gt;
&lt;p&gt;Several mechanisms distinguish the two policies. First, a carbon tax directly raises the relative price of all fossil fuels versus renewables and pushes production up the upward-sloping renewable supply curve only modestly. A renewable subsidy instead directly subsidises a reduction in the cost of renewables, which expands renewable supply — but this requires moving up the domestic renewable supply curve, wasting real resources in countries where the marginal renewable site is expensive (e.g., France with over 40% baseline renewable share). Second, a carbon tax creates a reallocation from coal to oil-gas (since the tax raises the coal price more per unit of energy), which can inadvertently raise oil-gas prices and redistribute income to exporters. A renewable subsidy does not have this feature in the same way. Third, the lump-sum financing of subsidies has a direct income cost, while carbon tax revenues are rebated, so only general equilibrium price effects matter for welfare. On average across countries, renewable subsidies cause more harm and generate smaller emission reductions per dollar.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-distinction-between-the-eu-and-asean-climate-clubs-and-why-do-outcomes-differ-so-substantially"&gt;Q6. What is the distinction between the EU and ASEAN climate clubs, and why do outcomes differ so substantially?&lt;/h3&gt;
&lt;p&gt;The EU club ($50 carbon tax + CBAM on imports from non-members) reduces global emissions by 3%, raises global utilitarian welfare by about 5%, but makes EU members net losers on average. The reason is that EU countries include many cold nations (Germany, Scandinavia) that benefit from warming; by cooling the climate, the policy harms them. Additionally, energy cost effects within the EU are heterogeneous — energy costs rise in France but fall in Poland and Germany — and Ireland is harmed through goods trade with Great Britain. The ASEAN club reduces global emissions by only 0.5% (ASEAN is smaller and less fossil-intensive in global terms), raises global utilitarian welfare by 0.8%, and ASEAN members broadly benefit because: (1) all ASEAN members are in the tropical/sub-tropical zone and thus lose from warming; (2) reducing global temperature yields direct productivity gains for members; (3) the energy rent loss for fossil exporters within ASEAN (Brunei, Indonesia) is outweighed by the climate benefit for others. The key structural difference is that the ASEAN club&amp;rsquo;s members are already losers from warming and hence have aligned incentives for carbon reduction.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-social-cost-of-carbon-computed-in-this-framework-and-how-does-it-vary-with-assumptions"&gt;Q7. What is the Social Cost of Carbon computed in this framework and how does it vary with assumptions?&lt;/h3&gt;
&lt;p&gt;Under utilitarian Pareto weights (ωi = 1, equal weight per person) and a 3°C warming by 2100, the global consumption-equivalent welfare loss is 17%, implying SCC = $203/tCO₂ at the current baseline temperature. Changing the point of linearisation to the 3°C warmer world raises the SCC to $302/tCO₂, indicating that damages accelerate as warming progresses and that the baseline approximation understates future costs. Under Negishi weights (proportional to income, ωi ∝ 1/u&amp;rsquo;(ci)), the SCC falls dramatically to $3.31/tCO₂, because damages are concentrated in low-income countries which receive little weight under income-proportional welfare aggregation. The authors note their static, log-linearised model provides a lower bound: fully dynamic IAMs with nonlinearities, uncertainty, or catastrophic-tail risks would further raise the SCC.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-estimate-energy-supply-elasticities-and-what-are-the-key-findings"&gt;Q8. How does the paper estimate energy supply elasticities and what are the key findings?&lt;/h3&gt;
&lt;p&gt;The authors regress changes in the oil-gas rent share of GDP on changes in the international oil-gas price (and changes in GDP as a control) country-by-country using first differences, recovering country-specific supply elasticities. Because some OLS estimates are noisy, negative, or below 1 (implying negative supply elasticity, inconsistent with theory), the authors apply an empirical Bayes shrinkage procedure: they impose a truncated-normal prior (truncated below 1) whose hyperparameters come from a pooled regression, and compute the posterior mean for each country. Key findings: oil-gas supply is nearly inelastic in OPEC nations (Saudi Arabia) and Russia and China, consistent with market power compressing effective supply elasticity; the US has relatively elastic oil-gas supply. Coal supply is substantially more elastic on average than oil-gas; the US and India have relatively inelastic coal supply; Russia and China have more elastic coal supply. Coal rents never exceed 1% of GDP even in the largest producers, consistent with near-competitive flat supply curves. These spatial patterns matter significantly for which countries gain or lose from energy price changes induced by climate policy.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-main-mechanism-through-which-leakage-operates--energy-trade-or-goods-trade--and-how-is-this-established"&gt;Q9. What is the main mechanism through which leakage operates — energy trade or goods trade — and how is this established?&lt;/h3&gt;
&lt;p&gt;The paper establishes that energy market effects are consistently larger in magnitude than goods-market terms-of-trade effects across all policy experiments (see Appendix Table A3). Leakage through energy trade operates because: (1) a domestic carbon tax reduces domestic demand for fossil fuels, lowering the international price of oil-gas (for small countries) or shifting demand between fuels; (2) lower oil-gas prices benefit importing countries and encourage them to use more fossil fuels, partially offsetting the original emission reduction. Goods-market leakage (productivity and competitiveness effects through the trade network) exists but is secondary. This finding has implications for policy: carbon border adjustment mechanisms (CBAMs) target goods trade leakage, but the model suggests the larger channel — energy trade leakage — is not addressed by CBAM alone.&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-or-sensitivity-analyses-does-the-paper-report"&gt;Q10. What robustness checks or sensitivity analyses does the paper report?&lt;/h3&gt;
&lt;p&gt;The paper reports several robustness exercises: (1) The damage function estimation reports results under OLS (Columns 1-2) and Poisson (Columns 3-4), with separate or restricted coefficients on importer and exporter temperatures; the preferred Poisson specification with restricted coefficients yields T* = 14.02 and γ = 0.012, and the separate-coefficient specification yields statistically indistinguishable estimates. (2) The SCC is computed at two points of approximation — the current baseline and a 3°C warmer world — yielding $203 and $302/tCO₂ respectively, giving a sense of nonlinearity bias from log-linearisation. (3) Welfare is reported under both utilitarian (ωi = 1) and Negishi (ωi ∝ 1/u&amp;rsquo;(ci)) weights throughout, and the results differ sharply, highlighting how inequality weighting matters. (4) The partial local adaptation parameter αT = 0.5 nests pure global peak (αT = 1) and pure local baseline (αT = 0) damage specifications. (5) Appendix Table A3 provides a comprehensive decomposition of welfare into climate, energy, and trade effects for all six policy scenarios (BAU, global carbon tax, China tax, US tax, EU club, ASEAN club), enabling consistency checks across experiments.&lt;/p&gt;
&lt;h3 id="q11-how-does-this-paper-relate-to-the-broader-literature-on-iams-and-sufficient-statistics"&gt;Q11. How does this paper relate to the broader literature on IAMs and sufficient statistics?&lt;/h3&gt;
&lt;p&gt;The paper makes three connections. First, it is related to the large IAM literature (Nordhaus and Yang 1996; Barrage and Nordhaus 2024; Cruz and Rossi-Hansberg 2024) but differs by explicitly decomposing welfare into observable sufficient statistics, avoiding the need to solve a large dynamic system. Second, it is related to the sufficient statistics literature in trade (Lashkaripour 2021 on trade wars; Baqaee and Farhi 2024 on trade barriers; Kleinman, Liu, and Redding 2024 on productivity shocks in trade models) — the paper extends this approach to a broad set of climate instruments in a model with detailed energy markets. Third, it differs from Bourany (2025) — a companion paper by one author — which solves for optimal climate agreement design; the present paper instead uses sufficient statistics to evaluate many given policies, trading optimality for analytical tractability and decomposability. The paper also distinguishes from Krusell and Smith (2022), which does not allow cross-border energy trade, and from Cruz and Rossi-Hansberg (2024), which does not model heterogeneous energy rents across space.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-scope-conditions-and-limitations-of-the-approach"&gt;Q12. What are the scope conditions and limitations of the approach?&lt;/h3&gt;
&lt;p&gt;Scope conditions and limitations are significant. (1) The model is static, so it cannot capture dynamic considerations: optimal intertemporal extraction paths, green paradox effects (whether carbon taxes accelerate fossil extraction), directed innovation toward renewables, adaptation capital accumulation, or dynamic leakage in energy markets. (2) The first-order log-linearisation abstracts from nonlinearities in the climate system, making the results most relevant as marginal effects near the current equilibrium rather than for large climate-policy changes or for evaluating policies at future, warmer states of the world. (3) The paper does not model market power in international energy markets (OPEC behaviour), abstracting from strategic behaviour by fossil exporters. (4) Labour is internationally immobile, so migration as a margin of adaptation is excluded. (5) Utility damages from climate change (mortality, amenity loss) are excluded — only productivity (TFP) damages are modelled; including utility damages would amplify gains and losses proportionally. (6) The framework cannot evaluate dynamic policy environments such as climate coordination with commitment problems or intergenerational redistribution from carbon taxation.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-policy-implications-of-the-papers-findings"&gt;Q13. What are the policy implications of the paper&amp;rsquo;s findings?&lt;/h3&gt;
&lt;p&gt;Several policy implications follow from the paper&amp;rsquo;s results, with important scope conditions. (1) Unilateral climate policy is largely ineffective for reducing global emissions and can even increase them (as in China&amp;rsquo;s carbon tax case); the standard free-rider analysis understates the problem because energy-market leakage can reverse the direction of emissions. (2) Renewable energy subsidies are generally a worse policy instrument than carbon taxes, because they push countries up costly domestic supply curves rather than reallocating away from fossil fuels through price signals; policy prescriptions that favour subsidies (such as the US Inflation Reduction Act) should account for this comparative inefficiency. (3) Climate clubs with both a domestic carbon tax and carbon tariffs (CBAMs) can overcome leakage effects and yield positive global welfare gains, but impose net costs on members whose composition makes them net losers from cooling (cold, energy-exporting member nations). This suggests club membership incentives are heterogeneous even within a bloc and require side payments or complementary redistribution to be stable. (4) ASEAN-style clubs where all members are hot-country losers from warming can achieve a Pareto-improvement for members while also improving global welfare, making them potentially more robust to free-riding than clubs like the EU where some members prefer a warmer climate. (5) The SCC estimated under utilitarian weights ($203/tCO₂) is substantially higher than under Negishi weights ($3.31/tCO₂), implying that the appropriate SCC for policy depends critically on how inequality across countries is weighted in the social welfare function.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Sufficient statistics (for climate policy)&lt;/strong&gt;: In this paper&amp;rsquo;s sense, a set of observable data moments and estimable elasticities — specifically nations&amp;rsquo; energy mix (shares of oil-gas, coal, renewables), energy rent shares of GDP, bilateral trade shares, energy supply and demand elasticities, and damage parameters — that fully characterise, to the first order, the welfare impact of a climate policy change without requiring the full model to be solved. The approach follows Chetty (2009) and extends it from tax incidence to climate policy in an IAM with trade.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Carbon leakage&lt;/strong&gt;: In this paper&amp;rsquo;s framework, the phenomenon by which a unilateral domestic carbon tax reduces domestic fossil demand and lowers the international price of oil-gas, inducing countries outside the policy to increase their fossil fuel consumption, partly or fully offsetting the original emission reduction. The paper shows leakage operates primarily through energy trade (oil-gas price channel) rather than through goods trade competitiveness effects, with energy effects consistently dominating in magnitude across all policy experiments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Local Cost of Carbon (LCC)&lt;/strong&gt;: The country-specific welfare cost of an additional unit of global carbon emissions, measured in monetary units as the negative of the partial derivative of country i&amp;rsquo;s welfare with respect to aggregate emissions, divided by the marginal utility of consumption. Distinct from the global Social Cost of Carbon (SCC), which aggregates LCCs across countries with Pareto weights. Countries whose productivity is harmed more by warming have a higher LCC; cold countries may have a negative LCC (they benefit from marginal warming).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural damage function&lt;/strong&gt;: The function Dy_i(E) mapping world cumulative emissions E to country i&amp;rsquo;s TFP via a quadratic temperature-productivity relationship with peak temperature T* and slope parameter γ, estimated in this paper from bilateral trade data (import penetration ratios and temperature differences) rather than from GDP-temperature regressions. The estimation is designed to be robust to the Lucas critique by netting out general-equilibrium propagation through trade and energy markets that would bias GDP-based estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Climate club&lt;/strong&gt;: In this paper&amp;rsquo;s usage (following Nordhaus 2015), a coalition of countries that jointly impose a domestic carbon tax on their own emissions and levy carbon tariffs (carbon border adjustment mechanism, CBAM) on imports from non-member countries scaled by the carbon intensity of those imports. The paper studies EU and ASEAN climate clubs and finds they differ sharply in welfare distribution: the EU club creates net losers among members (because some EU countries benefit from warming), while the ASEAN club delivers welfare gains for all members because all are hot-country losers from climate change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Energy rent effect&lt;/strong&gt;: The component of the welfare decomposition arising from changes in profits of domestic energy producers (fossil extractors, coal producers, renewable firms) due to changes in energy prices. Captured in the sufficient statistics formula as the profit share of GDP weighted by the relevant price change. Fossil-fuel-exporting countries have large positive exposure to oil-gas price increases (gains from price rises) and are harmed when global carbon policy reduces the fossil price — this is a key redistribution channel distinct from both climate damages and goods trade.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Bayes shrinkage (energy supply elasticities)&lt;/strong&gt;: In this paper, a procedure that estimates country-specific fossil and coal supply elasticities by first running OLS regressions of rent share changes on price changes country-by-country, then shrinking noisy or negative estimates toward a pooled mean by imposing a truncated-normal prior (truncated below 1 to enforce positive elasticities) and computing posterior means. Used because country-level time series are short and noisy, while the prior encodes the theoretical constraint that supply must be upward-sloping.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Negishi weights vs. utilitarian weights&lt;/strong&gt;: Two distinct social welfare aggregation methods used throughout the paper to aggregate country-level welfare changes into global welfare. Utilitarian weights (ωi = 1 per person) put equal importance on each person globally, so welfare gains in low-income tropical countries count fully; this yields high SCCs ($203/tCO₂) and large global welfare gains from carbon taxation. Negishi weights (ωi ∝ 1/u&amp;rsquo;(ci), proportional to income) downweight poor countries and upweight rich ones, yielding dramatically lower SCCs ($3.31/tCO₂) and smaller measured global welfare gains because damages concentrate in low-income countries that receive little weight.&lt;/p&gt;</description></item><item><title>Train to Opportunity: the Effect of Infrastructure on Intergenerational Mobility</title><link>https://macropaperwarehouse.com/papers/train-to-opportunity-the-effect-of-infrastructure-on-intergenerational-mobility/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/train-to-opportunity-the-effect-of-infrastructure-on-intergenerational-mobility/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether proximity to transport infrastructure can sever the occupational tie between parents and children — a question with direct bearing on the debate over place-based versus people-based policies. The authors exploit the nineteenth-century expansion of the railroad network across England and Wales, a setting where the First and Second Industrial Revolutions were remaking the occupational structure at the same time that the railroad was knitting together local labor markets and enabling geographic mobility.&lt;/p&gt;
&lt;p&gt;The empirical strategy centers on a novel dataset of close to 980,848 father-son pairs constructed from the full digitized population censuses of England and Wales in 1851, 1881, and 1911 (I-CeM project). Individuals are tracked across consecutive censuses using the Abramitzky-Mill-Perez (2019) linking procedure, which achieves match rates of 43–50% for men aged 40–52. Crucially, each individual is geolocated to the street level by matching census addresses to the GB1900 gazetteer, allowing railroad access to be measured as the straight-line distance from the childhood residence to the nearest train station — a finer measure than the district-level presence indicators used in prior work. Sons&amp;rsquo; occupations are observed at ages 40–52; fathers&amp;rsquo; occupations are measured 30 years earlier when sons were aged 10–22. Occupational mobility uses two complementary scales: HISCO categories (farming, laborer, services, sales, clerical, managerial, professional) and the continuous HISCAM social-interaction-distance ranking (scores 28–99, mean 50, SD 10).&lt;/p&gt;
&lt;p&gt;The key endogeneity problem is that railroad companies targeted low-density, cheap land, and that wealthy landowners and local politicians influenced station placement. To isolate exogenous variation, the authors construct a dynamic least-cost path (DLCP) network connecting 53 major towns identified by their 1801 populations (top 10% of the population distribution, threshold 9,172 inhabitants). The DLCP assigns slope costs to 50x50 meter grid cells and finds the minimum-cost path between every town pair. Lines are ranked by betweenness centrality to separate &amp;ldquo;early&amp;rdquo; 1851 lines from &amp;ldquo;late&amp;rdquo; 1881 lines, giving a time-varying instrument. Proximity to the nearest DLCP line is used as the instrument for proximity to the nearest actual train station, with standard errors clustered at the parish level. Controls include county and census-year fixed effects, distance to the nearest 1801 major town and its population, distance to Roman roads, ancient ports, and navigable waterways, plus household characteristics (number of servants as a wealth proxy, household size, and father&amp;rsquo;s foreign birth).&lt;/p&gt;
&lt;p&gt;Main results (preferred IV specification with full controls): sons who grew up one standard deviation — approximately 5 km, or about one hour&amp;rsquo;s walk — closer to a train station were 11 percentage points more likely to work in an occupation category different from their father&amp;rsquo;s. They were 5 percentage points more likely to be upwardly mobile, defined as a son&amp;rsquo;s HISCAM score exceeding his father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s distribution. The downward mobility estimate is 3 percentage points — positive but smaller in magnitude — indicating that railroad access raises occupational churn asymmetrically, predominantly upward. First-stage F-statistics exceed the Staiger-Stock threshold comfortably (135–414 across specifications). OLS estimates are uniformly smaller than IV estimates, consistent with historical evidence that the railroad targeted areas with weaker growth trajectories.&lt;/p&gt;
&lt;p&gt;The occupational transitions underlying these results run strongly out of farming and into professional, clerical, sales, and services categories, regardless of the father&amp;rsquo;s own occupation (Table IV). Sons growing up closer to the railroad were 19 percentage points less likely to work in a declining occupation and 16 percentage points more likely to work in a growing occupation. The distributional pattern shows an inverted-U relationship with father&amp;rsquo;s occupational decile for occupation-category switching and rank divergence, with the greatest gains concentrated among sons of middle-ranking fathers. For upward mobility specifically, the benefits diminish monotonically as father&amp;rsquo;s rank rises — sons from blue-collar backgrounds gained more (upward mobility coefficient 0.064) than sons from white-collar backgrounds (0.031).&lt;/p&gt;
&lt;p&gt;The authors decompose the total railroad effect on intergenerational mobility into three channels using a structural decomposition applied to a sample of 342,715 brothers: (1) changes in local labor-market opportunities, estimated as the effect on mobility for stayers; (2) changes in the returns to spatial mobility, estimated via a within-family comparison of brothers who moved versus stayed; and (3) changes in the rate of spatial mobility itself. Better railroad access raised the probability of moving away from the birth county by 15 percentage points. However, the estimated return to spatial mobility — the extra boost from actually moving — was reduced by railroad access (negative interaction between proximity and mover status), meaning the railroad decreased the relative advantage of leaving. The decomposition (Table C.6) shows that changes in local opportunities account for the great majority of the total mobility effect. Parish-level evidence confirms the local opportunity mechanism: better-connected parishes saw population growth, more industrial chimneys, more entrepreneurs, higher shares of skilled and literate workers, higher Gini coefficients, and higher median occupational ranks — consistent with agglomeration, industrialization, and skill-biased structural change.&lt;/p&gt;
&lt;p&gt;The policy implication is that transport infrastructure investment can reduce intergenerational persistence in occupational status, primarily by restructuring the local labor market rather than by enabling workers to exit. The caveat is that these gains were unevenly distributed — middle- and lower-ranking families benefited most, and the railroad simultaneously raised local inequality alongside local mobility.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-and-what-are-the-main-threats-it-addresses"&gt;Q1. What is the core identification strategy and what are the main threats it addresses?&lt;/h3&gt;
&lt;p&gt;The authors use a &amp;lsquo;dynamic least-cost path&amp;rsquo; (DLCP) instrument. They connect 53 major English and Welsh towns (defined as the top 10% of the 1801 population distribution, with at least 9,172 inhabitants) via least-cost routes computed over a 50×50 meter terrain grid that assigns slope-based costs to each cell. The instrument is proximity from the childhood residence to the nearest line in this DLCP network. The logic is that individuals incidentally located near the geographic route between major historical towns are more likely to be near an actual railroad — but the DLCP route is based purely on terrain costs, not on local demand, local resources, or the political lobbying that shaped where stations were actually placed. The strategy addresses: (a) reverse causality from high-growth areas attracting railroad placement; (b) sorting of ambitious or wealthy households toward connected parishes; (c) railroad companies&amp;rsquo; demand-driven routing choices. The exclusion restriction could be violated if location along least-cost paths between 1801 major towns is directly correlated with intergenerational mobility for reasons other than the railroad. The paper addresses this by controlling for distance to the nearest 1801 major town and its population (proximity to nodes), proximity to Roman roads, ancient ports, and navigable waterways (pre-existing trade routes), and household wealth proxies.&lt;/p&gt;
&lt;h3 id="q2-how-is-the-instrument-made-dynamic-and-why-does-this-matter"&gt;Q2. How is the instrument made dynamic, and why does this matter?&lt;/h3&gt;
&lt;p&gt;The authors divide the hypothetical network into &amp;rsquo;early&amp;rsquo; (1851) and &amp;rsquo;late&amp;rsquo; (1881) lines by ranking lines in decreasing order of betweenness centrality — the number of times a line connects major towns via shortest paths — until the total cost of the 1851 observed network is exhausted. This dynamic structure means the instrument varies across both space and census cohorts (sons measured in 1851-1881 versus 1881-1911). Without the dynamic feature, the instrument could conflate the effects of lines that were built early (and thus had decades to affect local economies) with lines built later. The temporal variation bolsters the plausibility of the exclusion restriction and is shown to be robust in alternative specifications using static least-cost paths and slope-free least-cost paths.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-four-dependent-variables-and-how-is-intergenerational-mobility-defined"&gt;Q3. What are the four dependent variables and how is intergenerational mobility defined?&lt;/h3&gt;
&lt;p&gt;The paper uses four measures: (1) an indicator equal to one if the son works in a different HISCO occupation category than his father; (2) the absolute value of the difference in HISCAM scores between son and father; (3) &amp;lsquo;upward mobility,&amp;rsquo; an indicator equal to one if the son&amp;rsquo;s HISCAM score exceeds his father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s score distribution; (4) &amp;lsquo;downward mobility,&amp;rsquo; the symmetric indicator for a decline greater than one standard deviation. Sons&amp;rsquo; occupations are observed when sons are 40–52 years old; fathers&amp;rsquo; occupations are measured 30 years earlier when sons were 10–22. The HISCAM scale is held constant over the period (national GB scale, 1800–1938) so that rankings reflect fixed social stratification positions rather than period-specific prestige. The paper also uses time-varying HISCAM, HISCLASS, Woollard, and Armstrong classifications as robustness checks.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-first-stage-performance-of-the-instrument"&gt;Q4. What is the first-stage performance of the instrument?&lt;/h3&gt;
&lt;p&gt;The first-stage relationship between proximity to the nearest DLCP line and proximity to the nearest actual train station is positive and statistically significant across all specifications. The Sanderson-Windmeijer F-statistic is 414 in the specification without controls, 136 with county and year fixed effects and full controls, and remains well above the conventional threshold of 10. The first-stage coefficient drops from 0.640 to 0.339 when full controls are added, indicating that a portion of the geographic correlation between the DLCP and the actual network reflects the pre-existing economic importance of towns and travel routes — which is precisely what the controls absorb.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q5. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The paper decomposes the total IV effect on intergenerational mobility using a three-part decomposition: (1) Changes in local opportunities, measured as the effect of proximity on mobility for sons who stayed in their birth county (stayers); (2) Changes in the returns to spatial mobility, estimated by comparing brothers who moved with brothers who stayed (using family fixed effects), and interacting this comparison with railroad proximity; (3) Changes in the rate of spatial mobility itself, estimated from the effect of proximity on the probability of county-to-county migration. Table C.6 shows that local opportunities account for the dominant share of the total effect. The railroad raised the migration probability by 15 percentage points (Table VI), so spatial mobility channels exist — but the railroad decreased the relative advantage of actually moving (negative interaction term in Table V), meaning the local opportunity channel more than offsets the spatial channel. Supporting evidence from parish-level regressions (Table VII) shows that better-connected parishes experienced significantly higher population growth, more industrial chimneys, more entrepreneurs per 100 square meters, higher shares of skilled and literate workers, higher Gini coefficients, and higher median occupational ranks — consistent with agglomeration and skill-biased industrialization.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-is-documented-by-fathers-occupation-and-position-in-the-distribution"&gt;Q6. What heterogeneity is documented by father&amp;rsquo;s occupation and position in the distribution?&lt;/h3&gt;
&lt;p&gt;The effects are heterogeneous by the father&amp;rsquo;s occupational position. Figure 6 shows an inverted-U pattern for occupation-category switching and absolute rank divergence: sons of middle-ranking fathers benefit most from railroad access. For upward mobility (Figure 6c), the benefits diminish monotonically from the lower end of the father&amp;rsquo;s distribution — sons of low-ranking fathers are most likely to move up. Sons of white-collar fathers see smaller (and sometimes statistically insignificant) upward mobility gains (0.031) compared with sons of blue-collar fathers (0.064), while the occupation-category switching benefit is also larger for blue-collar sons (0.108 vs. 0.057) (Table C.1). Separate transition matrices by HISCO category (Table IV) show that railroad access reduces the probability of farming for sons of all father types, and raises probabilities of clerical, sales, and services occupations. Effects on becoming a laborer are heterogeneous: for sons of farmers, proximity raises the probability of becoming a laborer; for sons in service occupations, it decreases it.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-run"&gt;Q7. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper performs an extensive battery. (1) Alternative connectivity measures: distance to the nearest railroad line, indicator variables for train station within 5, 10, and 15 km, and parish-level station presence. (2) Alternative mobility thresholds: 0.5, 1.5, and 2 standard deviations for upward and downward mobility; time-varying HISCAM to account for changing occupational prestige. (3) Removing railroad-specific occupations (train conductors, controllers) to check for mechanical effects. (4) Alternative specifications: second-order polynomials, parish fixed effects (10,419 parishes), and fully nonparametric covariate controls via k-means clustering (500 clusters). (5) Alternative instruments: a slope-free DLCP and a static (non-dynamic) least-cost path. (6) Geolocation robustness: using parish centroids instead of street-level addresses. (7) Linking bias: controlling for the individual probability of being linked using cubic polynomials on linkage probability and surname-frequency dummies; also checking that the railroad network explains little of the share of linked individuals at the parish level. (8) Subsamples: by census year (1851-1881 vs. 1881-1911), by county (leave-one-out), by rural/urban status, by father&amp;rsquo;s age, by son&amp;rsquo;s age, by birth order, by native/first-/second-generation immigrant status, by whether the son was born in the same county he grew up in, and by whether the father was in farming. (9) Causal response weighting: the Loken-Mogstad-Wiswall decomposition shows positive IV weights across the entire proximity distribution, consistent with a LATE interpretation. Results are stable across all checks.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-handle-the-selection-into-migration-problem-in-estimating-returns-to-spatial-mobility"&gt;Q8. How does the paper handle the selection-into-migration problem in estimating returns to spatial mobility?&lt;/h3&gt;
&lt;p&gt;The authors follow Abramitzky, Boustan, and Eriksson (2012) and use a within-family comparison of brothers — a subsample of 342,715 sons from 157,369 households who grew up in the same household but one moved county while the other stayed. Family fixed effects absorb the shared household characteristics (wealth, motivation, family networks, financial constraints) that jointly determine the propensity to migrate and the baseline mobility trajectory. The railroad-proximity interaction with mover status is instrumented using the interaction of the DLCP instrument with the mover indicator, via a control function approach. The estimated baseline return to spatial mobility (the mover premium) is positive and significant — movers have higher occupation-category divergence and shift more in both directions — but the railroad-induced change in return to mobility is negative, meaning that proximity to the railroad reduced the additional mobility benefit of actually migrating. This finding is the core of the conclusion that local opportunities, not spatial mobility, dominate.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-paper-document-about-local-labor-market-changes-induced-by-the-railroad"&gt;Q9. What does the paper document about local labor market changes induced by the railroad?&lt;/h3&gt;
&lt;p&gt;Parish-level IV regressions (Table VII) show that better proximity to the 1851 network (instrumented by the DLCP) is associated with: significantly higher population growth between 1851 and 1881; a significantly larger number of industrial chimneys (proxying factory concentration, sourced from Heblich-Trew-Zylberberg (2021)); more entrepreneurs per 100 square meters (from the British Business Census of Entrepreneurs); higher shares of high-skilled and literate workers; a higher Gini coefficient over occupational ranks; and a higher median occupational rank. Additionally, sons in better-connected parishes were 19 percentage points less likely to work in a declining occupation and 16 percentage points more likely to work in a growing occupation (Table C.3). Sons were also 3 percentage points more likely to be literate and 7 percentage points more likely to work in a non-manual occupation (Table C.5). These findings collectively point to agglomeration, industrialization, skill-biased technological change, and the creation of a new entrepreneur class as the mechanisms by which the railroad transformed local labor market structure.&lt;/p&gt;
&lt;h3 id="q10-what-prior-work-does-this-paper-relate-to-most-closely-and-what-distinguishes-it"&gt;Q10. What prior work does this paper relate to most closely, and what distinguishes it?&lt;/h3&gt;
&lt;p&gt;The paper sits at the intersection of the railroad-infrastructure and intergenerational-mobility literatures. In the infrastructure tradition, it relates closely to Donaldson (2018, AER) on railroads in India, Donaldson and Hornbeck (2016, QJE) on US market access, Bogart et al. (2022, JUE) on population and structural change in England and Wales, and Heblich-Redding-Sturm (2020, QJE) on London commuting and urban growth. The closest prior paper is Perez (2017) on nineteenth-century Argentina, who finds railroad access shifted children from agricultural into white-collar and skilled blue-collar occupations; this paper provides similar evidence for England and Wales at individual level and adds a full mechanism decomposition. In the intergenerational mobility tradition it relates to Long and Ferrie (2013, AER) and Long (2013, ERH) on census-based occupational mobility in Victorian Britain. The key methodological advantages of the current paper are: (a) use of the full (not 2%) census for all three years, yielding close to 1 million father-son pairs with match rates of 43–50% versus 15–33% in prior work; (b) street-level geolocation enabling individual-level rather than district-level measurement of railroad access; (c) the explicit three-way mechanism decomposition separating local opportunities, returns to migration, and migration rates; and (d) documenting rich heterogeneity by father&amp;rsquo;s occupational rank and occupation category.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-what-scope-conditions-limit-their-external-validity"&gt;Q11. What are the policy implications and what scope conditions limit their external validity?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s core policy message is that transport infrastructure investment can be an effective mechanism for reducing intergenerational occupational persistence — primarily by creating new local labor market opportunities rather than by enabling low-income workers to reach distant job centers. This provides historical support for place-based policies of the sort embodied in the Biden &amp;lsquo;Build Back Better&amp;rsquo; infrastructure proposals or the UK HS2 high-speed railway project (mentioned in the paper). The main scope conditions limiting generalizability are: (1) The setting is nineteenth-century England and Wales during the Industrial Revolution, when the occupational structure was shifting rapidly from farming to industry and commerce — the railroads arrived at a moment of latent demand for new labor market structures; (2) The benefits were not evenly distributed: middle-ranking families (by father&amp;rsquo;s occupational rank) gained most in absolute occupational switching and rank divergence, while the lowest-ranked families gained most specifically in upward mobility; (3) The railroad simultaneously raised local inequality alongside local mobility, suggesting infrastructure investment can be inequality-increasing in the cross-sectional distribution of wages even as it reduces intergenerational persistence; (4) The effects are highly localized — even 5 km of additional distance matters — implying that the placement of stations relative to where low-income families actually live is crucial for achieving distributional goals.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-paper-document-about-the-baseline-patterns-of-intergenerational-mobility-in-the-sample"&gt;Q12. What does the paper document about the baseline patterns of intergenerational mobility in the sample?&lt;/h3&gt;
&lt;p&gt;In the full sample of 980,848 father-son pairs covering 1851-1881 and 1881-1911, 80% of sons do not remain in the same HISCO occupation category as their father. The correlation between father&amp;rsquo;s and son&amp;rsquo;s HISCAM ranks is 0.28. Among sons, 18% experienced upward mobility (son&amp;rsquo;s HISCAM rank more than one SD higher than father&amp;rsquo;s) and 15% experienced downward mobility (more than one SD lower). About 31% of sons moved to a different county from where they grew up, settling on average 100 km away. Sons grew up on average 3.28 km from the nearest train station (SD 5.45 km). These descriptives reveal strong spatial clustering in intergenerational mobility patterns at the parish level.&lt;/p&gt;
&lt;h3 id="q13-does-the-late-interpretation-hold-and-what-does-the-weighting-function-show"&gt;Q13. Does the LATE interpretation hold and what does the weighting function show?&lt;/h3&gt;
&lt;p&gt;The authors verify the LATE interpretation via two approaches. First, following Loken-Mogstad-Wiswall (2012), they compute the causal response weighting function as the covariance between each discrete proximity indicator and the DLCP instrument, divided by the covariance between the proximity measure and the DLCP instrument. They find positive weights across the entire distribution of proximity to the nearest train station, concentrated most heavily for individuals residing 0.5 to 1.5 proximity units (approximately 2.7 to 8.1 km) from a train station — these are the individuals whose proximity is most affected by incidental location along the DLCP. The absence of negative weights indicates the IV estimate does not mix complier and never/always-taker effects in a sign-reversing way. Second, following Blandhol et al. (2022), a fully nonparametric specification using 500 k-means clusters for covariates yields estimates very close to the parametric baseline, consistent with a LATE interpretation of the linear IV estimator.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Dynamic Least-Cost Path (DLCP) Network&lt;/strong&gt;: The paper&amp;rsquo;s instrument for railroad access. A hypothetical railroad network connecting England and Wales&amp;rsquo;s 53 largest towns in 1801 via routes that minimize geographic cost (distance plus slope-based terrain costs), ignoring all demand-side factors. Lines are classified as &amp;rsquo;early&amp;rsquo; (1851) or &amp;rsquo;late&amp;rsquo; (1881) by betweenness centrality until the cost budget of the actual 1851 network is exhausted. Proximity from childhood residence to the nearest DLCP line instruments proximity to the nearest actual train station.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intergenerational Occupational Mobility&lt;/strong&gt;: In this paper, the degree to which a son&amp;rsquo;s adult occupation differs from his father&amp;rsquo;s, measured both categorically (same versus different HISCO category) and cardinally (difference in HISCAM scores). Upward (downward) mobility is specifically defined as the son&amp;rsquo;s HISCAM score exceeding (falling below) the father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s HISCAM distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HISCAM Score&lt;/strong&gt;: A continuous occupational ranking (range 28–99, mean 50, SD 10) derived from the frequency of social interactions — marriages, friendships, parent-child links — between occupations in historical data. Higher scores indicate a more advantageous position in the social stratification structure. The paper uses the national Great Britain scale, held constant for 1800–1938, to make rankings comparable across census years.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Local Opportunities Channel&lt;/strong&gt;: The mechanism by which railroad access improved intergenerational mobility through restructuring the local labor market — enabling commuting, attracting factories and entrepreneurs, spurring urbanization and industrialization, and creating new occupations requiring new skills — without requiring sons to migrate away from their birth county. Identified empirically as the effect of railroad proximity on mobility outcomes for sons who stayed in their birth county (stayers).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Returns to Spatial Mobility&lt;/strong&gt;: The additional intergenerational mobility benefit (or penalty) associated with actually migrating to another county, estimated using within-family variation among brothers — one who moved and one who stayed — to net out shared household-level determinants of mobility. The paper finds that railroad access reduced (made more negative) the returns to spatial mobility, meaning that the relative advantage of leaving shrank as local opportunities expanded.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inconsequential Place IV Approach&lt;/strong&gt;: An identification strategy (following Chandra-Thompson 2000 and Michaels 2008) in which the instrument for infrastructure access is constructed from the geographic convenience of locations lying between endpoints of a planned network, rather than from demand-side factors at those locations. The DLCP instrument in this paper is a specific implementation: individuals living between 1801 major towns incidentally receive railroad access because the low-cost route between towns passes near their residence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Occupational Tie (Father-Son)&lt;/strong&gt;: The tendency for sons to remain in the same occupation category or same position in the occupational ranking as their father. In this paper, severing the occupational tie means a son moves to a different HISCO category and/or achieves a HISCAM score meaningfully different from his father&amp;rsquo;s. The railroad&amp;rsquo;s main effect is framed as reducing this tie, with upward mobility being the dominant direction of change.&lt;/p&gt;</description></item><item><title>Universal Daycare and Mothers' Working Lifetime</title><link>https://macropaperwarehouse.com/papers/universal-daycare-and-mothers-working-lifetime/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/universal-daycare-and-mothers-working-lifetime/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper estimates the causal effects of universal daycare access on mothers&amp;rsquo; labor force participation, full-time employment, hours worked, and earnings across 34 years after the birth of their first child — the longest window examined in this literature. The motivation is twofold: the existing evidence base is overwhelmingly short-run, and the human capital channel (reduced depreciation of skills, accumulation of experience) implies that early labor market attachment during child-rearing years could compound over decades in ways that short-run estimates miss entirely.&lt;/p&gt;
&lt;p&gt;The identification exploits Denmark&amp;rsquo;s 1964 reform that converted a targeted (means-tested) childcare system into a universal one, which triggered a staggered geographic roll-out of daycare centers from 1966 onward across the country&amp;rsquo;s 2,033 neighborhoods nested in 277 municipalities. The paper combines digitized historical daycare yearbooks (1964–1975), the 1970 census, and administrative registers from Statistics Denmark covering 370,602 mothers who had their first child between 1964 and 1975. Employment is measured via annual contributions to the Supplementary Pension Fund (ATP); earnings from tax records are available from 1980 through 2015, adjusted to 2016 USD. The empirical strategy is a difference-in-differences design comparing mothers in neighborhoods with versus without daycare within the same municipality over time. Daycare availability when the first-born child turns four is used as the fixed treatment indicator for the long-run regressions. Municipality fixed effects absorb cross-sectional confounders; year-of-first-birth dummies capture macro trends.&lt;/p&gt;
&lt;p&gt;The contemporaneous effects are already substantial. Once year and municipality fixed effects and covariates are included, daycare availability raises the probability of participation by 1.5 percentage points when the child is two, rising to 5.3–5.7 percentage points for years three through six — translating to roughly 9 percent more likely to participate relative to the mean. Full-time employment rises by 9–12 percent relative to the mean for years three through six; hours worked increase by 0.27 hours per week (1.8 percent) when the child is four.&lt;/p&gt;
&lt;p&gt;The long-run effects persist throughout the entire working life. Relative to the sample mean, mothers with daycare access are 9.7 percent more likely to participate when the first child turns four, declining to 5.7 percent at child age 14, 3.1 percent at child age 22, and still 1.2 percent at child age 34 (when the average mother is approximately 57.7 years old). Full-time employment effects follow a parallel trajectory: 11 percent higher at child age four, 8.2 percent at child age 14, and 4.4 percent at child age 34. Log earnings (conditional on employment) range between 3 and 6 percent higher throughout the observation window; mothers earn 5.3 percent more when the child is 16 and 4.2 percent more when the child is 34.&lt;/p&gt;
&lt;p&gt;Heterogeneity by education is a central finding. For low-educated mothers (no post-secondary education, 50 percent of the sample), participation effects are 10.1 percent at child age 10, 5.1 percent at child age 17, and remain statistically significant through 32 years. For higher-educated mothers, participation effects are 3.9 percent at child age 10, fall below 1 percent by child age 17, and become statistically indistinguishable from zero by child age 23. Employment effects are thus larger and more persistent for low-educated mothers. Earnings effects, however, are more closely aligned across education groups and show a distinctive pattern for higher-educated mothers: earnings effects persist and remain significant long after employment effects have faded, suggesting that sustained attachment during child-rearing years translates into qualitative career advancement (not just more years worked) for the more educated group.&lt;/p&gt;
&lt;p&gt;Potential mediators include reduced secondary fertility and increased parental separation. Daycare for children aged three to six reduces the total number of children by 0.036 (1.6 percent relative to the mean of 2.2), reduces the probability of having more than two children by 1.8 percentage points (6.0 percent), and increases birth spacing by 0.137 years, making mothers 2.2 percentage points less likely to have a second child within two years. Additionally, mothers with daycare access are 2 percentage points more likely to live apart from the first-born child&amp;rsquo;s father when that child turns 16 — consistent with greater female economic independence. These mediator effects do not vary systematically by education level. Daycare access does not affect additional educational attainment after first birth, ruling out re-skilling as a channel.&lt;/p&gt;
&lt;p&gt;The policy implication is that subsidized universal daycare is not merely a short-run labor supply intervention but a persistent investment in female human capital accumulation, with effects that compound over careers and remain economically meaningful into near-retirement ages.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-key-threats-to-it"&gt;Q1. What is the identification strategy and what are the key threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses a staggered difference-in-differences design. The key variation is the timing of daycare center openings across neighborhoods within municipalities following the 1964/1966 Danish reform. Daycare availability in the year the first-born child turns four is the fixed treatment indicator for long-run regressions; current-year daycare availability is used for contemporaneous regressions. Municipality fixed effects absorb time-invariant local differences; year-of-first-birth dummies absorb aggregate time trends. The main threat is non-random placement of daycare centers — if centers opened in areas where female labor force participation was already rising, the estimates would be upward biased. The paper addresses this with (1) an event study at the neighborhood level using data from 1960 through 2003 showing no pre-reform differential trends between neighborhoods that later received daycare and those that did not (compared against placebo neighborhoods assigned fictitious opening dates mimicking the actual distribution), and (2) a selective migration check showing that mothers who moved longer distances from their birthplace were no more likely to reside in a neighborhood with daycare once the full conditioning set is included. A residual concern is that for mothers having their first child before 1970, neighborhood assignment is measured post-birth (1970 census), which is addressed by a robustness check excluding the pre-1970 first-birth cohort.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-deal-with-heterogeneous-treatment-effects-and-two-way-fixed-effects-bias"&gt;Q2. How does the paper deal with heterogeneous treatment effects and two-way fixed effects bias?&lt;/h3&gt;
&lt;p&gt;The paper acknowledges the recent literature on TWFE bias under treatment effect heterogeneity (De Chaisemartin and d&amp;rsquo;Haultfoeuille 2020; Callaway and Sant&amp;rsquo;Anna 2021; Sun and Abraham 2021; Borusyak et al. 2024). It replicates the pre-reform event study using the Borusyak et al. (2024) imputation estimator, which is robust to heterogeneous treatment effects and allows for covariates, and finds similar results to the standard TWFE event study (Appendix Figure A.2). The main long-run regressions fix the treatment indicator to daycare availability when the child is four, so there is no variation in treatment timing within a regression, limiting but not eliminating TWFE concerns for the long-run estimates.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-main-mechanism-behind-the-persistent-effects"&gt;Q3. What is the main mechanism behind the persistent effects?&lt;/h3&gt;
&lt;p&gt;The paper attributes the persistence to human capital dynamics: labor force participation during the child-rearing years reduces depreciation of previously accumulated human capital (from education and prior work experience) and enables new on-the-job human capital accumulation through the current job. For low-educated mothers, the primary channel appears to be the extensive margin — daycare moves mothers who would otherwise become homemakers into paid employment, and the employment effects persist because once labor market attachment is established, it is durable. For higher-educated mothers, the earnings-employment gap is the key signal: employment effects fade within roughly 23 years (consistent with convergence once children are no longer preschool age and informal care becomes feasible), yet earnings remain elevated for decades, suggesting that the women who maintained employment during child-rearing years accrued qualitatively better positions — more experience, better job-match, more promotions — compared to those who did not.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-main-mediators-and-how-are-they-distinguished"&gt;Q4. What are the main mediators and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;Three mediators are examined. First, secondary fertility: daycare for children aged 3–6 reduces number of children by 0.036, probability of a third child by 1.8 percentage points, and probability of a fourth child by 0.5 percentage points. The effect operates through daycare for children 3–6 (not 0–2), consistent with the main employment effects operating when the child is three or older. The fertility reduction increases the opportunity cost interpretation — daycare raises the effective wage, making additional children more costly in terms of foregone earnings. Second, birth spacing: mothers with daycare access wait 0.137 more years between first and second child, and are 2.2 percentage points less likely to have the second child within two years, allowing longer uninterrupted work spells. Third, parental separation: mothers with daycare access are 2 percentage points more likely to live apart from the child&amp;rsquo;s father at child age 16, consistent with greater economic independence from labor market participation reducing barriers to separation. Additional educational attainment after first birth is tested and found to be an insignificant channel (no significant effect overall, a marginal effect only for low-educated mothers), ruling out re-skilling as a mediator.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented-beyond-the-education-split"&gt;Q5. What heterogeneity is documented beyond the education split?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s primary heterogeneity analysis is by maternal education level (low: no post-secondary education versus higher: any post-secondary education including vocational training, college, or university). The education split produces the most substantive finding: employment effects are larger and more persistent for low-educated mothers, while the earnings-employment divergence is the distinctive feature for higher-educated mothers. No other dimensions of heterogeneity (by birth cohort, by municipality type beyond the urban indicator, by parity) are formally reported in the main results, though geographic robustness checks (exclusion of three largest cities, exclusion of suburbs) implicitly test whether effects are concentrated in particular settings and find they are not.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Four main sets of robustness checks are reported. First, selective migration: regressions of daycare availability on distance moved from birthplace (linear, quadratic, and IHST-transformed) with the full conditioning set show no significant relationship, ruling out systematic sorting into daycare neighborhoods. Second, pre-1970 cohort exclusion: restricting to mothers with first birth after 1970 (for whom the 1970 census address is predetermined relative to birth) yields qualitatively similar results, though participation effect sizes are somewhat smaller. Third, urban geography: excluding the three largest municipalities (Copenhagen, Frederiksberg, Aarhus, Odense) and separately excluding suburbs of Copenhagen and Aarhus both leave the main results intact. Fourth, differential time trends: allowing the most populous neighborhood within each municipality to have its own set of time dummies (to capture potentially faster urban trend evolution) does not change the finding that participation and earnings effects persist beyond 30 years. The paper also shows that results are robust to an alternative participation definition based solely on ATP contributions for all years (versus mixing ATP pre-1980 and earnings post-1980).&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-prior-work-and-what-is-its-main-contribution"&gt;Q7. How does this paper relate to prior work and what is its main contribution?&lt;/h3&gt;
&lt;p&gt;The prior literature falls into two camps. The short-run camp (Havnes and Mogstad 2011 for Norway; Carta and Rizzica 2018 for Italy; Bettendorf et al. 2015 for Netherlands; Cascio 2009 and Fitzpatrick 2012 for the US) documents modest to moderate employment effects during the preschool years. The medium-run camp (Lefebvre et al. 2009 and Haeck et al. 2015 for Quebec; Nollenberger and Rodriguez-Planas 2015 for Spain; Herbst 2017 for the US Lanham Act) tracks effects up to about 11–17 years. This paper&amp;rsquo;s first contribution is extending the window to 34 years — covering the majority of the working life — using Danish administrative data that allow continuous observation rather than decennial census snapshots. The second contribution is documenting the earnings-employment divergence for higher-educated mothers specifically, which was not visible in shorter windows. The third contribution is the simultaneous analysis of fertility, spacing, and parental separation as mediators using the same administrative data and identification strategy, rather than treating these as separate exercises in different papers.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-scope-conditions-and-policy-implications"&gt;Q8. What are the scope conditions and policy implications?&lt;/h3&gt;
&lt;p&gt;Several scope conditions qualify the policy implications. First, the context is a universal reform in a Nordic welfare state with strong labor market institutions and universal access; the results may not directly generalize to settings with low baseline female employment or weak formal sector employment. Second, the relevant margin for the 1960s–70s cohorts was daycare for children aged three to six; the paper notes that by recent decades the relevant margin has shifted to children under two (consistent with Simonsen 2010 finding effects for younger children in 2001 data), possibly reflecting changing cultural norms or the fact that 1960s–70s mothers had multiple children before returning to work. Third, the employment effects are larger for low-educated mothers, so the labor market attachment argument applies most forcefully to this group. Fourth, the negative fertility effects mean that the total welfare calculation must weigh labor market gains against reductions in desired family size. The policy implication the paper emphasizes is that universal daycare is an investment in long-run economic output, not merely a short-run participation subsidy, because the labor market attachment it induces during child-rearing years compounds over careers through human capital accumulation.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-sample-and-data-structure"&gt;Q9. What is the sample and data structure?&lt;/h3&gt;
&lt;p&gt;The sample consists of 370,602 mothers who had their first child between 1964 and 1975 and were resident in Denmark in 1970 (from the census), after excluding women with immigrant backgrounds (2.2 percent) and those who died or emigrated before the first child turned 16 (0.6 percent). Employment is observed from the birth of the first child through 34 years after (1964–2009 approximately); earnings from 1980 through 2015. The daycare panel is constructed from historical yearbooks (1964–1975) and administrative registers (1976–1993) and provides yearly neighborhood-level data on daycare availability. The average mother in the sample was born in 1945, was 23.7 years old at first birth, had 10.8 years of education, and had 2.2 children total. The sample is split roughly 50/50 between low-educated and higher-educated mothers.&lt;/p&gt;
&lt;h3 id="q10-why-do-effects-appear-only-when-the-child-is-three-not-earlier"&gt;Q10. Why do effects appear only when the child is three, not earlier?&lt;/h3&gt;
&lt;p&gt;The paper finds that contemporary participation effects are small and statistically insignificant for years zero through two, then jump sharply at year three. The paper attributes this to two factors: (1) the universal daycare reform primarily expanded slots for children aged three to six, with nurseries for children under three expanding much more slowly through the 1980s and 1990s (Figure A.1 in the paper); and (2) cultural norms and the multi-child fertility pattern of this cohort — mothers in the 1960s–70s were more likely to have multiple children before returning to work, implying that the eldest child often reached age three or four before the mother re-entered employment. This contrasts with more recent periods (Simonsen 2010 uses 2001 data) where the relevant margin has shifted to children under two.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Universal daycare&lt;/strong&gt;: In the paper&amp;rsquo;s sense, daycare centers open to children from all socioeconomic backgrounds (not means-tested), with building costs fully publicly funded and operating costs split among state, municipality, and parents (with parents paying 30 percent), following the 1964 Danish reform. Contrasted with the pre-reform &amp;rsquo;targeted&amp;rsquo; system that only subsidized institutions where two-thirds of children came from low-income families.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Working lifetime effects&lt;/strong&gt;: The paper&amp;rsquo;s central object of analysis: the causal impact of early daycare access on maternal labor outcomes measured annually across 34 years after the birth of the first child, covering the majority of the working life. Distinguished from short-run (0–7 year) and medium-run (up to 11–17 year) effects documented in prior work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor market attachment&lt;/strong&gt;: As used in the paper, the sustained connection to paid employment during the child-rearing years (when children are of preschool age). The paper argues that attachment during this period is the mechanism for long-run effects because it reduces human capital depreciation and enables on-the-job accumulation of experience and job-specific skills.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;ATP (Supplementary Pension Fund) contributions&lt;/strong&gt;: The paper&amp;rsquo;s primary employment measure for years before 1980. Annual ATP contributions are proportional to hours worked: one-third contribution corresponds to 10–19 hours/week, two-thirds to 20–29 hours/week, and full contribution to 30 or more hours/week. Used to construct both a participation dummy and a full-time employment dummy (full ATP contribution = at least 30 hours/week). Crucially, the unemployed, self-employed, and those outside the labor force made no ATP contributions during this period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Human capital depreciation channel&lt;/strong&gt;: The mechanism by which absence from the labor market during child-rearing years erodes previously accumulated skills (from education and prior work). The paper uses this concept, following Adda et al. (2017) and Lefebvre et al. (2009), to explain why participation effects on earnings can persist long after direct employment effects have diminished: mothers who worked during preschool years entered subsequent career phases with a larger, less-depreciated human capital stock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Secondary fertility decisions&lt;/strong&gt;: The paper&amp;rsquo;s term for fertility choices conditional on already having a first child, i.e., the decision to have additional children. Examined on the intensive margin (number of additional children, spacing between births) rather than extensive margin (whether to have any children), because the sample consists entirely of women who already have at least one child.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Daycare for 3–6 year olds vs. 0–2 year olds&lt;/strong&gt;: The paper distinguishes between two types of daycare that expanded at different speeds: daycare for children aged 3–6 expanded rapidly from 1966, while nurseries for children under 3 (crèches) expanded only from the 1980s–1990s. All significant effects in the paper — on employment, fertility, and parental separation — load onto access to daycare for children aged 3–6, not 0–2, consistent with the historical timing of the expansion.&lt;/p&gt;</description></item><item><title>Within-Firm Pay Inequality and Productivity</title><link>https://macropaperwarehouse.com/papers/within-firm-pay-inequality-and-productivity/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/within-firm-pay-inequality-and-productivity/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates how within-firm pay inequality relates to firm-level labor productivity, using a novel linkage of three confidential U.S. Census Bureau datasets covering millions of workers at hundreds of thousands of firms from 2003 to 2015.&lt;/p&gt;
&lt;p&gt;The motivating puzzle is that the dramatic rise in U.S. wage inequality since the 1970s is well documented, but the firm-side determinants of within-firm pay dispersion have been difficult to study due to the absence of comprehensive matched employer-employee data in the United States. The paper asks whether firms&amp;rsquo; own productivity levels can explain the structure of pay inequality within firms, and whether rising aggregate productivity can account for the secular increase in the CEO-to-median-worker pay gap.&lt;/p&gt;
&lt;p&gt;The data come from three linked sources. The Longitudinal Employer-Household Dynamics (LEHD) program provides quarterly earnings for essentially all UI-covered workers from 2003 to 2015, covering all 50 states and Washington, D.C. These earnings encompass salaries, wages, bonuses, and exercised stock options, making them comprehensive for top earners. The Longitudinal Business Database (LBD) supplies annual firm-level revenue and employment, from which the key productivity measure — real revenue per worker, deflated to 2010 dollars using the PCE deflator — is constructed. The Management and Organizational Practices Survey (MOPS), a supplement to the Annual Survey of Manufactures conducted in 2010 and 2015, provides structured management scores (scaled 0 to 1) measuring the intensity of performance monitoring, target-setting, and incentive use across manufacturing firms. The main analysis sample restricts to firms with at least 100 full-year &amp;ldquo;6-quarter sandwich&amp;rdquo; workers to ensure clean measurement of annual earnings; it covers approximately 443,000 firm-year observations and 73,000 unique firms. A supplementary Execucomp sample (4,681 firms, 2006–2016) validates results for large publicly traded firms.&lt;/p&gt;
&lt;p&gt;Three main findings are reported. First, employees at more productive firms earn more across the entire within-firm pay distribution — from the 1st to the 99th percentile. A 10 percent increase in productivity is associated with a 0.7 percent increase in average worker pay (elasticity 0.068). Moving from the 10th to the 90th percentile of the firm productivity distribution projects an 18 percent increase in average pay.&lt;/p&gt;
&lt;p&gt;Second, the pay-productivity relationship is steeper at higher pay ranks — it strengthens monotonically with seniority. For a given doubling of firm productivity, the top-paid employee (likely the CEO) sees approximately 15 percent more pay, while the median-paid employee sees approximately 7 percent more. Equivalently, the pay-productivity elasticity is 0.15 for the top earner and 0.07 for the median earner. At the percentile level, a 10 percent productivity increase predicts a 0.86 percent pay increase at the 90th percentile but only 0.53 percent at the 10th percentile. Consequently, more productive firms have higher within-firm inequality: a 10 percent productivity increase widens the top-earner-to-median-worker log pay gap by 0.9 percent, and moving from the 10th to the 90th percentile of productivity projects a 23.1 percent increase in this gap. These cross-sectional results survive firm fixed effects, demographic controls (sex, education, age), industry fixed effects at the 6-digit NAICS level, and 2SLS instrumentation with industry exposures to seven major currencies, oil prices, and economic policy uncertainty (Alfaro, Bloom, and Lin 2024). Within-worker, within-firm estimates confirm the pattern dynamically: when a firm&amp;rsquo;s productivity doubles, workers earning $45,000–$65,000 expect roughly a 1 percent pay increase while workers earning above $300,000 expect nearly a 2 percent increase. The pay-productivity relationship is roughly twice as strong for top earners at publicly traded firms as at private firms (coefficient of 0.22 vs. 0.13 for rank-1 earners), while workers outside the top 50 ranks show similar coefficients across ownership types.&lt;/p&gt;
&lt;p&gt;Third, the mechanism is traced to performance-based pay. More productive firms exhibit higher within-year pay volatility (measured as the standard deviation of quarterly log earnings within a year), particularly for top earners, consistent with larger bonus payments. Firms with higher structured management scores — capturing more intensive performance monitoring, goal-setting, and incentive pay — also show higher pay levels and higher pay volatility for top earners, with the gradient across ranks matching the productivity results.&lt;/p&gt;
&lt;p&gt;Finally, a back-of-the-envelope calculation applies the estimated pay-productivity elasticities to observed aggregate productivity growth. Aggregate U.S. labor productivity roughly doubled (96 percent compounded growth) from 1980 to 2013. The top-earner-to-median-worker pay ratio at firms with at least 100 employees rose from 7.55 in 1980 to 8.69 in 2013 (an increase of 1.14). Applying the paper&amp;rsquo;s elasticities for rank-1 (0.1534) and rank-50 (0.0657) earners to the observed productivity doubling predicts a ratio of 8.01 in 2013 — accounting for 40 percent of the actual increase. The authors interpret this as evidence that rising productivity, channeled through differential performance pay, is a quantitatively important driver of rising within-firm inequality.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-primary-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the primary identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The core cross-sectional estimates in models (1) and (2) regress percentile- or rank-specific pay on log revenue per worker, controlling for a quadratic expansion of firm-level worker demographic composition (sex, education, age and their interactions), year fixed effects, and 6-digit NAICS industry fixed effects. The main threat is omitted variable bias: unobserved firm characteristics correlated with both productivity and pay (e.g., high-skill worker sorting into high-productivity firms) could inflate estimates. The paper addresses this in three ways. First, specifications with firm fixed effects (Appendix Figure A.1) use only within-firm changes in productivity and pay, producing similar convex-across-ranks patterns. Second, the within-worker, within-firm change specification (model 4, Figure 2) holds individual workers fixed and relates earnings growth to productivity growth. Third, a 2SLS approach instruments log productivity (and its interaction with rank) using industry-level exposures to seven currency pairs, oil prices, and economic policy uncertainty constructed from rolling 10-year daily stock-return regressions by Alfaro, Bloom, and Lin (2024); the logic is that industries have idiosyncratic exposure to these aggregate shocks, so productivity movements attributable to the instruments are exogenous to individual pay-setting. The 2SLS results are broadly similar to OLS in sign and pattern, though first-stage F-statistics are approximately 3, which is weak by conventional standards. Additional tests using lagged productivity (Appendix Table A.3) show if anything stronger relationships, consistent with productivity causally passing through to pay rather than pay determining past productivity.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-proposed-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms proposed and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The primary mechanism proposed is performance-based pay (bonuses and incentive compensation) that is disproportionately concentrated among senior managers at more productive firms. The paper cannot directly observe bonus pay in the LEHD, which reports total quarterly earnings. Instead, it uses within-year pay volatility — the standard deviation of log quarterly earnings within a calendar year — as a proxy for bonus income (most visibly fourth-quarter bonus payments). Figure 4 shows that top earners at more productive firms have significantly higher pay volatility, and this relationship is steeper at higher ranks, exactly paralleling the pay-level results. The management channel is examined separately: Figure 5 shows that firms with higher MOPS structured management scores (capturing explicit monitoring, target-setting, and incentive-pay practices) display higher pay levels and higher pay volatility for top earners, again with the gradient increasing at the top. The public-vs.-private ownership comparison is a further diagnostic: if performance-based executive compensation is the mechanism, it should be stronger at publicly traded firms, where stock grants, option awards, and formal incentive contracts are more prevalent. Panel a of Figure 3 confirms the top-earner pay-productivity coefficient is 0.22 at public firms and 0.13 at private firms, while workers outside the top 50 show similar coefficients across ownership type. This asymmetry is robust to reweighting public firms to match the employment distribution of private firms (panel b of Figure 3), ruling out pure size effects as the explanation.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented-across-sectors-firm-age-and-ownership-type"&gt;Q3. What heterogeneity is documented across sectors, firm age, and ownership type?&lt;/h3&gt;
&lt;p&gt;Across sectors (Appendix Figure A.2), the positive and convex pay-productivity gradient across earnings ranks is present in nearly all 18 two-digit NAICS sectors. Shallower (less convex) patterns appear in utilities, finance and insurance, and health, which the authors attribute to heavy regulation limiting scope for differential performance pay across ranks. Across firm age groups (Appendix Figure A.3), the pattern holds across firms younger than 10 years, between 10 and 25 years, and 25 or more years. Across ownership, the pay-productivity relationship for top earners is roughly twice as large in publicly traded firms as in privately held firms, while the relationship for workers outside the top 50 is similar. Within publicly traded firms, the LEHD top-earner coefficients closely match those for named executives in the Compustat Execucomp data (Figure 3, panel a), validating both the LEHD measure of top earnings and the Execucomp-based executive pay literature.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper runs the following robustness checks: (1) Full demographic controls — a quadratic expansion of firm-level shares by sex, education category, and age group, plus interactions — included in all baseline regressions to account for worker sorting. (2) 6-digit NAICS industry fixed effects to net out cross-industry pay and productivity variation. (3) Firm fixed effects (Appendix Figure A.1): the convex pattern across ranks survives when only within-firm variation in productivity and pay is used. (4) Sector heterogeneity analysis (Appendix Figure A.2): the main pattern holds across nearly all 18 two-digit NAICS sectors. (5) Firm age heterogeneity (Appendix Figure A.3): results hold across all age groups. (6) Reweighting public firms to match private firms&amp;rsquo; employment distribution (Figure 3, panel b): the stronger pay-productivity gradient for top earners at public firms is not explained by their greater average size. (7) Size controls: including log total LEHD employment does not eliminate the pattern. (8) 2SLS with macroeconomic instruments: similar signs and pattern to OLS, supporting causal interpretation despite weak first stages. (9) Lagged productivity (Appendix Table A.3): if anything, the pay-productivity relationship by rank is slightly stronger when using prior-year productivity, reducing reverse-causality concerns. (10) Comparison to Execucomp: the LEHD public-firm top-earner coefficients align with those from Execucomp named executives. (11) Analysis of sandwich-worker selection (Appendix Table A.1): workers at more productive firms are marginally more likely to remain sandwich workers the following year, with this pattern slightly stronger at lower earnings ranks; the paper discusses this selection and argues it does not drive the main results.&lt;/p&gt;
&lt;h3 id="q5-what-exactly-is-the-lehd-earnings-measure-and-how-does-it-capture-bonuses"&gt;Q5. What exactly is the LEHD earnings measure and how does it capture bonuses?&lt;/h3&gt;
&lt;p&gt;The LEHD is based on state unemployment insurance (UI) wage records submitted by employers. It captures total quarterly earnings, including salaries, wages, bonuses, stock option exercises, and restricted stock awards when vested. Qualified (incentive) stock options are not subject to UI tax and are excluded, but these are capped and the paper judges them immaterial for top earners. The quarterly frequency of the data allows the paper to construct within-year pay volatility (the standard deviation of log quarterly earnings in a year) as a proxy for bonus income, since bonus payments typically appear as spikes in Q4. The paper uses only non-imputed demographic characteristics from ancillary LEHD sources; imputed values (e.g., education, which is imputed for 88 percent of individuals) are replaced with a constant and flagged with a missing-value indicator.&lt;/p&gt;
&lt;h3 id="q6-how-exactly-is-firm-productivity-measured-and-what-are-its-limitations"&gt;Q6. How exactly is firm productivity measured and what are its limitations?&lt;/h3&gt;
&lt;p&gt;Productivity is measured as real revenue per worker (log scale), with nominal revenue deflated to 2010 dollars using the PCE deflator. Revenue and employment come from the LBD, which covers all non-farm sectors from 1997 onward. This is a revenue-based labor productivity measure, not total factor productivity, and no industry-level price deflators are used beyond the economy-wide PCE; instead, 6-digit NAICS industry fixed effects control for cross-industry differences in revenue-per-worker levels. The LBD&amp;rsquo;s revenue coverage may be biased toward older, more stable firms, but the paper argues this has minimal impact because its sample is already restricted to large firms (at least 100 full-year workers). The paper explicitly contrasts its broad economy-wide measure with more granular TFP measures available only for manufacturing and in Economic Census years.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-structured-management-score-and-what-does-it-measure"&gt;Q7. What is the structured management score and what does it measure?&lt;/h3&gt;
&lt;p&gt;The structured management score is derived from 16 core questions in the MOPS asking plant managers about practices in three domains: performance monitoring, target setting, and incentivization of workers. Each question is scored 0 to 1, where 0 reflects least structured (less explicit, formal, frequent, or specific) and 1 reflects most structured (more explicit, formal, frequent, or specific). The firm-level score is an employment-weighted average of establishment-level scores (requiring at least 10 non-missing responses per establishment). It ranges from 0 to 1 and follows the methodology of Bloom et al. (2019), who establish that higher scores predict higher establishment-level productivity. Because MOPS targets manufacturing establishments surveyed in the ASM, the management sample is a 2.5 percent subset of the main sample, resulting in wider standard errors for management-related estimates. The paper treats this score as an indirect proxy for the adoption of performance-based incentive systems.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-song-et-al-2019-and-the-broader-between-firm-vs-within-firm-inequality-literature"&gt;Q8. How does this paper relate to and differ from Song et al. (2019) and the broader between-firm vs. within-firm inequality literature?&lt;/h3&gt;
&lt;p&gt;Song et al. (2019), also using linked LEHD-LBD data, document that the rise in U.S. earnings inequality between 1978 and 2013 was driven predominantly by increases in between-firm pay dispersion, with within-firm inequality rising more modestly. This paper takes the within-firm inequality result as a starting point and asks what firm characteristics predict cross-sectional and dynamic variation in within-firm inequality. The key addition is connecting within-firm pay dispersion to revenue labor productivity and to management practices, neither of which Song et al. (2019) directly analyze. The paper uses Song et al.&amp;rsquo;s published aggregate statistics on top-earner and median-earner pay (from their Figure VI) as the benchmark for the back-of-the-envelope calculation linking rising productivity to rising inequality. More broadly, the paper contributes to a cross-country literature (Barth et al. (2016), Card, Heining, and Kline (2013), Faggio, Salvanes, and Van Reenen (2010), Mueller, Ouimet, and Simintzi (2017)) that documents firms as the locus of increasing wage dispersion, by providing a specific firm-level mechanism — productivity and performance-pay practices.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-the-ceo-pay-literature"&gt;Q9. How does this paper relate to and differ from the CEO pay literature?&lt;/h3&gt;
&lt;p&gt;The CEO pay literature (Gabaix and Landier (2008), Frydman and Jenter (2010), Kaplan (2013), Edmans and Gabaix (2016)) debates whether rising CEO pay reflects performance, firm size, or rent extraction, but typically studies only the named top executives at large publicly traded firms covered by Execucomp. This paper&amp;rsquo;s key innovation is extending the analysis to all workers across the full within-firm pay distribution, for millions of U.S. workers at firms of all sizes and ownership types. It finds that the pay-productivity gradient is present across all earnings ranks, not only at the CEO level, though it is steeper at the top. The paper validates its LEHD-based top-earner results against Execucomp, finding close agreement for publicly traded firms, and interprets the public-vs.-private differential as consistent with formal performance-based executive contracts being more prevalent at public firms — a finding consistent with Gao and Li (2015), who show CEO pay-performance sensitivity is greater at public firms.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-aggregate-inequality-implications-and-how-robust-is-the-40-percent-estimate"&gt;Q10. What are the aggregate inequality implications and how robust is the 40 percent estimate?&lt;/h3&gt;
&lt;p&gt;The 40 percent figure comes from a back-of-the-envelope calculation in Table 4. Using Song et al.&amp;rsquo;s (2019) data, the top-earner-to-median-worker pay ratio rose from 7.55 in 1980 to 8.69 in 2013 (a change of 1.14). Aggregate U.S. labor productivity grew 96 percent compounded over this period (sourced from FRED series PRS85006092). The paper applies the pay-productivity elasticities for rank-1 (0.1534) and rank-50 (0.0657) earners from Figure 1 to this productivity growth to predict earnings levels in 2013. The predicted top-earner mean earnings is $224,357 (versus actual $301,614) and predicted median mean is $28,013 (versus actual $34,702), yielding a predicted ratio of 8.01 and an explained change of 0.46, which is 40.13 percent of the actual change of 1.14. The authors label this a &amp;lsquo;simple back-of-the-envelope&amp;rsquo; calculation and do not claim it as a structural decomposition. Key caveats: (i) the cross-sectional elasticities from 2003–2015 are applied to a 1980–2013 trend, assuming stability of these relationships over time; (ii) aggregate productivity growth may also shift the productivity distribution of firms, which the calculation does not fully model; (iii) the calculation attributes none of the remaining 60 percent, which could include technology, globalization, changing labor market institutions, or other forces.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-role-of-firm-size-in-explaining-the-results"&gt;Q11. What is the role of firm size in explaining the results?&lt;/h3&gt;
&lt;p&gt;Publicly traded firms in the sample are substantially larger than private firms on average (mean 7,763 versus 491.7 full-year employees). To ensure the stronger pay-productivity gradient at public firms is not simply a size artifact, the paper reweights public firms to match the employment distribution of private firms (using ventile-based inverse-probability weights) and finds the differential persists (panel b of Figure 3). The paper also includes log total LEHD employment as a control in additional specifications and reports similar results. The large-firm pay premium literature (Brown and Medoff (1989), Oi and Idson (1999)) posits that large firms pay more due to compensating differentials, monitoring difficulties, or rent-sharing. The paper&amp;rsquo;s finding that pay is higher at more productive firms across the entire earnings distribution is interpreted as more supportive of the rent-sharing explanation, since compensation-based and monitoring-based explanations would not apply uniformly to all workers.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The main policy-relevant implication is that rising productivity — itself associated with technology adoption and innovation — contributes substantially (estimated 40 percent) to the CEO-to-median-worker pay gap that the Dodd-Frank Act requires publicly traded firms to disclose annually from 2018. This implies that policies targeting within-firm pay inequality may need to grapple with the fact that a significant share of observed inequality is tied to real productivity differences and performance-pay practices, not purely to governance failures or rent extraction. However, several scope conditions limit this implication: the 40 percent figure is an economy-wide back-of-the-envelope estimate with caveats about stability of elasticities over time; the paper does not assess whether performance pay practices are optimally structured or reflect rent-seeking; the mechanism analysis uses pay volatility and management scores as proxies rather than direct observation of bonus contracts; and the remaining 60 percent of the inequality increase is left unaccounted for, potentially reflecting factors outside the paper&amp;rsquo;s framework.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-key-data-limitations-and-potential-measurement-concerns"&gt;Q13. What are the key data limitations and potential measurement concerns?&lt;/h3&gt;
&lt;p&gt;Several limitations are acknowledged or implicit. (1) Revenue labor productivity is used rather than TFP; the measure conflates product demand and productivity shocks and does not adjust for industry-specific output price variation. (2) LEHD earnings exclude qualified (incentive) stock options not subject to UI tax; the paper argues these are capped and immaterial for top earners, but this may understate total compensation for senior executives, especially at technology firms. (3) Within-year pay volatility is used as a proxy for bonus income rather than direct bonus data. (4) The management sample is confined to firms with at least one manufacturing establishment in the MOPS, covering only 2.5 percent of main-sample firm-year observations, limiting precision. (5) Education is imputed for 88 percent of individuals in the LEHD; the paper uses only non-imputed values and controls for missingness, but this reduces demographic control precision. (6) The IV first-stage F-statistics are approximately 3, suggesting weak instruments, so 2SLS standard errors are wide and the causal interpretation should be taken cautiously. (7) The sample is restricted to firms with at least 100 full-year workers, so results do not speak to smaller firms, which employ a large share of the U.S. workforce.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Revenue labor productivity&lt;/strong&gt;: Real revenue per worker at the firm level, computed from LBD annual revenue deflated to 2010 dollars using the PCE deflator and divided by total firm employment; the paper&amp;rsquo;s primary measure of firm performance, entered in log form in all regressions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pay-productivity elasticity (by rank)&lt;/strong&gt;: The regression coefficient on log firm productivity in a regression of mean log annual earnings for a given within-firm earnings rank or percentile; the paper documents that this elasticity rises monotonically from approximately 0.07 for the median earner to 0.15 for the top earner (rank 1), producing a convex schedule across ranks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Within-firm earnings inequality&lt;/strong&gt;: Dispersion in annual earnings among full-year workers within a single firm in a given year; measured variously as the 90th-10th percentile log earnings gap, the 99th-10th gap, the top-earner-to-50th-percentile gap, and the top-earner-to-10th-percentile gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Within-year pay volatility&lt;/strong&gt;: The standard deviation of log quarterly earnings within a calendar year for a given worker rank; used as a proxy for variable (bonus) compensation since it captures deviations from a constant salary path, particularly fourth-quarter bonus payments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structured management score (MOPS)&lt;/strong&gt;: A continuous index bounded between 0 and 1 derived from 16 MOPS survey questions on performance monitoring, target-setting, and worker incentivization practices; higher values indicate more explicit, formal, frequent, and specific management practices, following the scoring methodology of Bloom et al. (2019).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;6-quarter sandwich worker&lt;/strong&gt;: An individual who is employed at and earns above the minimum wage at the same firm in all four quarters of the current year, the fourth quarter of the prior year, and the first quarter of the following year; the restriction ensures that measured annual earnings reflect genuine full-year employment rather than partial-year spells or job transitions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DHS (Davis-Haltiwanger-Schuh) growth rate&lt;/strong&gt;: A symmetric growth rate measure defined as (x_t - x_{t-1}) / (0.5 * (x_t + x_{t-1})), bounded between -2 and 2; used in the within-worker, within-firm change analysis to measure both earnings growth and productivity growth while accommodating entry and exit.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Top-earner-to-median-worker pay ratio&lt;/strong&gt;: The ratio of mean annual earnings of the highest-paid worker to mean annual earnings of the median-paid worker within firms, aggregated across firms of different sizes using employment weights; the Dodd-Frank Act metric that publicly traded firms have been required to disclose annually since 2018, and the paper&amp;rsquo;s primary metric for the aggregate inequality calculation.&lt;/p&gt;</description></item><item><title>Zero-hours Contracts in a Frictional Labour Market</title><link>https://macropaperwarehouse.com/papers/zero-hours-contracts-in-a-frictional-labour-market/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/zero-hours-contracts-in-a-frictional-labour-market/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Dolado, Lalé, and Turon build a structural equilibrium model of the U.K. low-wage labour market to evaluate zero-hours contracts (ZHCs), employment agreements under which firms are not required to guarantee any minimum working hours and workers may decline any hours offered. The paper&amp;rsquo;s central question is whether ZHCs raise or lower welfare in general equilibrium, and through which channels. The model features two-sided heterogeneity in a random-search-and-matching environment: firms differ in the volatility of their labour demand, workers differ in their relative preferences for flexible versus regular employment, and wages are fixed at or near the statutory minimum wage. Three mechanisms operate simultaneously. First, a job-creation effect: firms facing highly volatile demand that cannot profitably hire under regular terms enter the market only because ZHCs exist. Second, a substitution effect: some firms that could hire under regular contracts instead post ZHC vacancies, crowding out regular employment. Third, a labour-force-participation effect: workers with a strong preference for flexible schedules join the labour force specifically because ZHCs exist and would withdraw if ZHCs were banned.&lt;/p&gt;
&lt;p&gt;The model is calibrated to U.K. Labour Force Survey data for the low-pay segment (roughly 16 percent of total employment), covering September 2018 through March 2020, with a sample of 9,342 individuals aged 16 to 69. A mixture-of-exponentials approach due to Karlis and Xekalaki (1999) applied to job-tenure and unemployment-duration distributions reveals statistically exactly two worker types in both ZHC employment and unemployment, and only one in regular employment, consistent with the presence of R-best workers (who prefer regular employment but accept ZHCs as a stepping stone) and Z-only workers (who would exit the labour force without ZHCs) but not R-only or Z-best workers. Calibrated parameters include a biweekly job-finding rate of λ(θ) = 0.051, a job-destruction probability of δ = 0.005, an on-the-job search efficiency of x = 0.352, and a share of R-best workers of ζ_{R-best} = 0.969. The matching function elasticity ψ is estimated to be 0.65 from U.K. occupation-level hiring and vacancy data (range 0.60–0.70 across specifications). ZHC employment accounts for 6.5 percent of the low-wage employment stock but 19.4 percent of vacancies, because higher turnover in ZHC jobs causes them to be re-advertised more frequently.&lt;/p&gt;
&lt;p&gt;A ban on ZHCs — simulated as an extreme tightening of flexible-work regulation — raises the unemployment rate by 2.0 to 2.7 percentage points depending on the assumed volatility of ZHC firms&amp;rsquo; demand. When ZHC workers have a low enough disutility of labour that they remain in the workforce after a ban (accepting regular jobs instead), the employment rate falls by the same 2.0 to 2.7 p.p., and sectoral GDP falls by only 0.02 to 0.14 percent, because higher average hours per employed worker partially offset the employment decline. When ZHC workers&amp;rsquo; disutility is high enough that they withdraw from the labour force, the employment-rate fall is larger — 4.8 to 5.4 p.p. — and sectoral GDP falls by 2.9 to 3.2 percent. Decomposing via the model&amp;rsquo;s analytical formula (Proposition 4a), lower job creation alone would reduce regular employment by almost 30 percent in isolation (λ(tilde-θ)/λ(θ) = 0.71), but this is partially offset by reduced vacancy competition (+24 percent, ceteris paribus) and improved search efficiency for regular jobs (+15 percent, ceteris paribus) after the ban.&lt;/p&gt;
&lt;p&gt;Welfare effects are measured in consumption-equivalent variation units. In general equilibrium, R-best workers (those who prefer regular jobs but sometimes hold ZHCs as a stepping stone) suffer welfare losses of −0.5 to −0.6 percent of consumption from a ZHC ban, driven primarily by longer expected unemployment spells. Yet in a partial equilibrium experiment that converts their ZHC jobs to regular jobs while holding all other equilibrium objects fixed, these same workers gain approximately +0.2 percent: the substitution effect is genuinely welfare-improving for them in isolation, but the job-creation channel dominates in general equilibrium and more than reverses that gain. Z-only workers — those who would exit the labour force if ZHCs were banned — suffer general-equilibrium welfare losses of −1.7 to −2.0 percent (low-disutility scenario) or approximately −1.8 to −2.1 percent (high-disutility scenario). These losses exceed the losses to R-best workers because Z-only workers are also forced into a type of employment they strictly prefer to avoid. The paper concludes that a ZHC ban is welfare-reducing for all workers in general equilibrium, and proposes that policy instead target ZHC use toward matches where workers voluntarily choose flexibility (Recommendation P1) and toward small firms that cannot diversify demand volatility across many positions (Recommendation P2).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-models-core-structure-and-what-frictions-drive-the-results"&gt;Q1. What is the model&amp;rsquo;s core structure and what frictions drive the results?&lt;/h3&gt;
&lt;p&gt;The model is a discrete-time steady-state random-search-and-matching model with two-sided heterogeneity. Workers are heterogeneous in their flow payoffs from regular employment (ω^i_R), flexible ZHC employment (ω^i_Z), and non-employment (ω^i_N), with these payoffs shaped by CRRA utility over consumption and a type-specific disutility of hours worked (α^i). Firms are heterogeneous in the volatility of their demand shock (σ_j), which determines the expected profit flow under each contract type. Flow profits depend on how actual hours h deviate from a stochastic target h-tilde via a quadratic loss specification. Market tightness θ is determined endogenously by free entry. The key friction is random search: workers cannot direct their search to their preferred contract type, so R-best workers sometimes end up in ZHCs and must search on-the-job to move to regular employment.&lt;/p&gt;
&lt;h3 id="q2-how-are-worker-types-identified-empirically-and-why-only-two-types"&gt;Q2. How are worker types identified empirically, and why only two types?&lt;/h3&gt;
&lt;p&gt;The paper adapts a mixture-of-exponential distributions procedure from Karlis and Xekalaki (1999), applied separately to the duration distribution of ZHC employment, regular employment, and unemployment in LFS data. A bootstrapped sequential hypothesis test determines the number of latent classes M* that best fits the survival function. For ZHC employment, two exponential components are needed (p-value for M=1 vs. M≥2 is 0.01; for M=2 vs. M≥3 it is 0.74). For regular employment, one component suffices (p-value for M=1 vs. M≥2 is 0.99). For unemployment, again two components (p-values 0.01 and 0.93 respectively). Cross-referencing which types are present in which states using the model&amp;rsquo;s theoretical exit-rate table rules out R-only and Z-best workers, leaving only R-best and Z-only workers as consistent with all three distributions simultaneously.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q3. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Identification rests on three steps. First, the mixture-of-exponentials procedure identifies the number of worker types from shape of duration distributions; this step relies on recalled job tenure and unemployment duration, which the authors acknowledge may suffer from recall bias and heaping (rounding to salient durations). Second, the turnover parameters are calibrated by minimizing distance between model-implied and empirical transition matrices across U, Z, and R states from the longitudinal LFS; the main limitation noted is that the two moments (transitions and durations) are not jointly consistent because they come from different measurement processes. Third, flow profits and payoffs are calibrated to external moments (minimum wage, replacement rate, business creation costs) and the preference for ZHC hours; the hours volatility parameter σ_Z has no direct empirical counterpart and is varied across scenarios. The model abstracts from wage bargaining, treating wages as fixed at the minimum wage, which reduces scope for confounding but is an approximation even in the low-wage sector.&lt;/p&gt;
&lt;h3 id="q4-how-are-the-three-channels--job-creation-substitution-and-labour-force-participation--distinguished-in-the-quantitative-analysis"&gt;Q4. How are the three channels — job creation, substitution, and labour-force participation — distinguished in the quantitative analysis?&lt;/h3&gt;
&lt;p&gt;The job-creation channel is captured by Z-only firms (firms with σ_Z = 6 such that regular employment is not profitable): removing ZHCs forces these firms out of the market entirely, reducing labour market tightness θ and hence the aggregate job-finding rate λ(θ). The substitution channel is captured by Z-best firms (σ_Z = 3): these firms could profitably hire under regular contracts but choose ZHCs, and after a ban they convert vacancies to regular posts, with incomplete crowd-out due to general equilibrium adjustment. The labour-force-participation channel is captured by Z-only workers: those with disutility α^i above the threshold (WTP &amp;gt; £7.9 per week to avoid regular work) withdraw from the labour force when ZHCs are banned, while those below the threshold remain and take regular jobs. The paper runs scenarios that vary both the firm side (low vs. high volatility) and the worker side (low vs. high disutility) to disentangle the magnitude of each channel.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-decomposition-of-the-effect-on-regular-employment-proposition-4a"&gt;Q5. What is the decomposition of the effect on regular employment (Proposition 4a)?&lt;/h3&gt;
&lt;p&gt;Under the calibrated parameters (no Z-best workers), regular employment in the baseline relative to the no-ZHC counterfactual equals the product of three multiplicative terms. The job-creation term is λ(θ)/λ(tilde-θ) = 1/0.71 ≈ 1.41, meaning that ZHCs raise the job-finding rate by about 41 percent relative to the no-ZHC counterfactual. The vacancy-competition term vR/v ≈ 0.81 (80.6 percent of vacancies are for regular jobs, while the remaining 19.4 percent for ZHC jobs dilute the pool). The search-efficiency term captures the fact that some R-best workers are in ZHC employment and search on-the-job at reduced intensity x &amp;lt; 1. The ceteris paribus decomposition at the ban scenario indicates: job creation alone would cut regular employment by 29 percent; competition reduction adds 24 percent; and search-efficiency gains add 15 percent — so the post-ban equilibrium has higher regular employment despite worse job creation overall.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-handle-the-partial-versus-general-equilibrium-distinction-for-welfare"&gt;Q6. How does the paper handle the partial versus general equilibrium distinction for welfare?&lt;/h3&gt;
&lt;p&gt;For R-best workers, the PE experiment replaces their ZHC jobs with regular jobs while keeping all other equilibrium objects (tightness θ, vacancy composition, etc.) fixed. This isolates the substitution effect and yields a welfare gain of approximately +0.15 to +0.18 percent for R-best workers. In general equilibrium, the full ban requires θ to fall (less job creation), which extends unemployment spells, and the net welfare effect is −0.50 to −0.62 percent. The difference between GE and PE therefore quantifies the job-creation externality that ZHCs provide — approximately 0.65 to 0.80 percentage points of consumption equivalent variation for R-best workers. For Z-only workers, the PE experiment replaces ZHC jobs with non-employment (their next-best option in the baseline), yielding PE welfare changes of −2.94 to −3.28 percent, which overstates the GE loss (−1.65 to −2.0 percent) because GE adjustment allows some Z-only workers to take regular jobs, partially compensating for the loss of ZHC access.&lt;/p&gt;
&lt;h3 id="q7-what-heterogeneity-is-documented-in-the-data-for-uk-zhc-workers"&gt;Q7. What heterogeneity is documented in the data for U.K. ZHC workers?&lt;/h3&gt;
&lt;p&gt;ZHC employment is concentrated at both ends of the age distribution: workers aged 16–29 are over-represented, as are workers aged 55–69, relative to regular employment. Mean age is 40.8 years for ZHC workers vs. 46.3 for regular workers. Gender composition is similar: 56.5 percent female in ZHCs vs. 60.4 percent female in regular employment, a difference that is not statistically significant. Educational attainment distributions are similar: 21.9 percent of ZHC workers hold a degree vs. 18.0 percent of regular workers. By industry, ZHC employment is heavily concentrated in Accommodation and food services (19.9 percent), Health and social work (20.5 percent), and Arts, entertainment and recreation (6.7 percent). Average hours worked are 18.4 per week for continuously employed ZHC workers vs. 28.1 for regular contract workers; the standard deviation of hours is 7.8 vs. 7.2. 16.6 percent of ZHC workers report wanting more hours vs. 10.1 percent in regular contracts, and 18.2 percent of ZHC workers are looking for another/additional job vs. 5.0 percent of regular workers, suggesting a minority are in involuntary underemployment while a majority are not actively seeking to change.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-key-calibrated-parameter-values-and-how-do-they-compare-to-the-broader-literature"&gt;Q8. What are the key calibrated parameter values and how do they compare to the broader literature?&lt;/h3&gt;
&lt;p&gt;The biweekly job-finding rate λ(θ) = 0.051; the biweekly job-destruction probability δ = 0.005; on-the-job search efficiency x = 0.352 (authors note this is on the high end but consistent with estimates accounting for flexible work); share of R-best workers ζ_{R-best} = 0.969; share of type-R vacancy-posting firms γ_R = 0.950. The matching function elasticity ψ = 0.65 (estimated from U.K. data, range 0.60–0.70, higher than the commonly used 0.50 but consistent with bias-corrected estimates from Borowczyk-Martins et al. 2013). The job-filling rate is 0.21 per biweek, consistent with Kuhn et al. (2021) U.K. estimates of 0.35–0.38 per month. The vacancy posting cost κ = £36.3 per week and startup cost K = £4,376, the latter close to the £4,500 implied by U.K. business creation data. Non-employment income b = £148.8 per week (replacement ratio 80 percent). The minimum wage is set to £7.50 per hour (2017 U.K. National Living Wage); labour productivity p = £8.25, implying a 10 percent productivity premium over the minimum wage.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-run-and-do-the-main-results-change"&gt;Q9. What robustness checks are run, and do the main results change?&lt;/h3&gt;
&lt;p&gt;The authors run three main robustness analyses. First, they vary the hours parameters: an alternative calibration uses σ_Z = 4.5 for both firm types but differentiates by mean hours (µ_Z = 20 for Z-best, µ_Z = 16 for Z-only); employment and unemployment effects are modestly smaller than the baseline but welfare effects are nearly identical. Second, they hold µ_Z = 18 and vary σ_Z to 1.0 (low) and 8.0 (high); results move in the expected direction and remain broadly consistent. Third, they vary the targeted job-filling rate: at λ(θ)/θ = 0.16 (25 percent lower than baseline), the unemployment response to a ZHC ban is only 0.33–0.51 p.p. and GDP effects are positive in the low-disutility case; at λ(θ)/θ = 0.26 (25 percent higher), unemployment rises by 4.1–5.5 p.p. and sectoral GDP falls by up to 6 percent. The authors conclude that the baseline calibration of 0.21 is the most plausible. The qualitative conclusions — that GE welfare effects are negative for all workers — are robust across specifications.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q10. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The closest model-based study is Scarfe (2019) on casual work in Australia. Scarfe&amp;rsquo;s model features homogeneous agents ex ante, with contract choice driven by luck (stochastic match productivity), while Dolado et al. emphasise ex ante heterogeneity in preferences/profitability as the primary source of variation. The empirical study of Datta et al. (2019) documents U.K. ZHC characteristics using LFS, online survey, and matched employer-employee data from the social care sector; Dolado et al. use the LFS but impose structural discipline to recover preference parameters and conduct GE welfare analysis. The paper differs from the dual labour market literature (Cahuc et al. 2016, 2020; Créchet 2022) in that temporary jobs in that literature have a fixed expiration date, whereas ZHCs are jobs with potentially long tenure but endogenously lower expected duration due to on-the-job search quit-outs, not contractual termination. Mas and Pallais (2017) and Angelici and Profeta (2020) use field experiments to estimate workers&amp;rsquo; valuation of flexibility; Dolado et al. instead recover this from duration distributions, allowing for general equilibrium job-creation and participation effects that field experiments cannot capture.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-sorting-patterns-in-the-equilibrium-and-what-sustains-zhc-jobs"&gt;Q11. What are the sorting patterns in the equilibrium, and what sustains ZHC jobs?&lt;/h3&gt;
&lt;p&gt;In the baseline equilibrium, 66.8 percent of filled ZHC jobs are held by R-best workers (workers who prefer regular employment but accept ZHCs as a stepping stone). Only 4.8 percent of employed R-best workers are in ZHCs at any point in time, because most vacancies are for regular jobs (80.6 percent of vacancies). This sorting has a crucial implication: ZHC vacancies would not be viable without the presence of R-best workers, because Z-only workers alone are too few to sustain the ZHC sector in equilibrium. A firm posting a ZHC vacancy accepts a higher worker-turnover risk (R-best workers quit on-the-job once they find a regular vacancy) in exchange for the profit advantage of hours flexibility; the trade-off is viable only because the random search pool contains enough R-best workers willing to take ZHC jobs temporarily.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper identifies four recommendations. P1: restrict ZHCs to matches where the worker voluntarily chooses the flexible contract when offered a choice; this would protect R-best workers who currently end up in ZHCs due to search frictions from the substitution effect without eliminating the job-creation channel. P2: prioritise access to ZHCs for small firms (as a proxy for inability to diversify demand shocks), limiting substitution by large firms while preserving genuine job creation by high-volatility operators. P3: recognise that the allocation of hours-flexibility between firms and workers is often an implicit and incomplete contract rather than an explicit one. P4: regulate the sharing of hours flexibility — specifically, who controls the timing and quantity of work — to reduce the income uncertainty that generates the main political objections to ZHCs. The scope conditions for all recommendations are: the low-wage sector of the U.K. labour market; the results do not directly apply to higher-wage workers with more bargaining power, or to markets where exclusivity clauses remain common.&lt;/p&gt;
&lt;h3 id="q13-what-key-empirical-facts-about-zhc-flows-does-the-paper-document"&gt;Q13. What key empirical facts about ZHC flows does the paper document?&lt;/h3&gt;
&lt;p&gt;From the transition matrix estimated from LFS data: 11 percent of exits from unemployment are to ZHC employment. The rate of transition to unemployment is almost 50 percent larger in ZHC employment than in regular employment (6.2 percent vs. 4.4 percent semi-annually). Job-to-job transitions from ZHC to regular employment are 6.5 percent semi-annually; the reverse (regular to ZHC) is only 0.5 percent. Nearly half of ZHC workers report job tenures longer than two years. 9.2 percent of ZHC workers were recruited in the last three months vs. 3.4 percent of regular workers; 30.3 percent of ZHC workers have been with their employer less than one year vs. 14.3 percent in regular contracts. The non-employment rate for this low-pay segment is 11.2 percent; ZHCs account for 4.6 percent of the overall sample (5.2 percent of employees), about 1.5 times the aggregate U.K. incidence rate.&lt;/p&gt;
&lt;h3 id="q14-what-does-the-model-say-about-time-spent-out-of-regular-employment-following-a-zhc-ban"&gt;Q14. What does the model say about time spent out of regular employment following a ZHC ban?&lt;/h3&gt;
&lt;p&gt;Despite higher aggregate unemployment rates after the ban, R-best workers spend less total time out of regular employment: the duration of non-regular-employment spells decreases by 7 weeks. This is because ZHCs, by acting as a stepping stone, expose workers to more frequent labour market transitions — they cycle through unemployment, ZHC employment, and regular employment rather than simply unemployment and regular employment. The ban removes the ZHC stepping stone, so workers face longer individual unemployment spells but avoid the ZHC-employment phase, and on net spend more time in regular employment. However, this does not translate into a welfare gain because (a) ZHC employment, even if imperfect, provides utility above the unemployment level, and (b) the longer unemployment spells that do occur under a ban are more costly than the shorter ZHC spells they replace.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Zero-hours contract (ZHC)&lt;/strong&gt;: In the paper&amp;rsquo;s sense, an employment arrangement under which the employer is not obligated to provide any minimum guaranteed hours of paid work, and the worker is not required to accept any hours offered. Workers on ZHCs in the U.K. hold &amp;lsquo;worker&amp;rsquo; status (between employee and self-employed), entitling them to holiday pay, minimum wage protections, and Universal Credit, but not redundancy pay. The key feature for the model is that actual hours worked equal the firm&amp;rsquo;s demand realisation, eliminating the quadratic deviation costs that arise under fixed-hours regular contracts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;R-best workers&lt;/strong&gt;: In the paper&amp;rsquo;s worker taxonomy, individuals for whom the asset value of regular employment strictly exceeds that of ZHC employment, which in turn exceeds the asset value of non-employment (W^i_R &amp;gt; W^i_Z &amp;gt; N^i). These workers accept ZHCs as a stepping stone when regular jobs are unavailable, and search on-the-job (at reduced efficiency x) for regular vacancies. They constitute 96.9 percent of the low-wage sector in the calibration and account for two-thirds of filled ZHC jobs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Z-only workers&lt;/strong&gt;: Workers for whom the asset value of ZHC employment exceeds both the value of regular employment and non-employment (W^i_Z &amp;gt; N^i &amp;gt; W^i_R, or W^i_Z &amp;gt; W^i_R &amp;gt; N^i), and who prefer non-employment to regular work. Without ZHCs, these workers&amp;rsquo; participation in the labour market depends on whether their disutility parameter α^i implies ω^i_R &amp;gt; ω^i_N. A subset — those with high disutility (WTP &amp;gt; £7.9 per week to avoid regular work) — exit the labour force if ZHCs are banned, generating the participation effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Z-only firms&lt;/strong&gt;: In the paper&amp;rsquo;s firm taxonomy, firms with high demand volatility (σ_Z = 6 in the calibration) for which regular employment is not profitable (V^j_R &amp;lt; 0 &amp;lt; V^j_Z). These firms can only operate and post vacancies because ZHCs allow them to set actual hours equal to realised demand. A ban on ZHCs causes Z-only firms to exit entirely, generating the pure job-creation loss.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Z-best firms&lt;/strong&gt;: Firms with moderate demand volatility (σ_Z = 3 in the calibration) that could profitably post regular vacancies (V^j_R &amp;gt; 0) but prefer ZHC vacancies because the hours-flexibility profit advantage outweighs the higher quit risk from R-best workers. A ban redirects these firms to regular contracts, constituting the substitution effect on the firm side.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stepping-stone effect&lt;/strong&gt;: The mechanism by which R-best workers accept ZHC employment when unemployed, using it as a bridge to search on-the-job for regular employment. ZHCs therefore simultaneously reduce unemployment duration and extend the time workers spend out of regular employment. The paper documents that a ZHC ban reduces total time out of regular employment by 7 weeks for R-best workers despite raising the unemployment rate, precisely because the stepping-stone pathway — which adds a ZHC phase before reaching regular employment — is eliminated.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption equivalent variation (welfare measure)&lt;/strong&gt;: The percentage permanent change in consumption that would make a worker indifferent between the baseline equilibrium (with ZHCs) and the counterfactual (ZHC ban). The paper uses this metric to express welfare effects: R-best workers suffer losses of −0.50 to −0.62 percent, and Z-only workers suffer losses of −1.65 to −2.0 percent, in general equilibrium following a ZHC ban.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mixture-of-exponentials identification of worker types&lt;/strong&gt;: A statistical procedure adapted from Karlis and Xekalaki (1999) that fits the empirical distribution of job tenure or unemployment duration as a mixture of M exponential distributions. Each component corresponds to a latent class of workers exiting the labour market state at a distinct rate. The optimal number of components M* is chosen via a bootstrapped sequential hypothesis test. Applied to U.K. LFS data, the procedure identifies M* = 2 for ZHC employment and unemployment, and M* = 1 for regular employment, which the model interprets as evidence for R-best and Z-only worker types.&lt;/p&gt;</description></item><item><title>Aggregate Implications of Heterogeneous Inflation Expectations: The Role of Individual Experience</title><link>https://macropaperwarehouse.com/papers/aggregate-implications-of-heterogeneous-inflation-expectations-the-role-of-individual-experience/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/aggregate-implications-of-heterogeneous-inflation-expectations-the-role-of-individual-experience/</guid><description>&lt;p&gt;Consumers&amp;rsquo; inflation expectations are heterogeneous across birth cohorts and history-dependent: using panel data from the Survey of Consumer Expectations (SCE), the paper documents that each cohort&amp;rsquo;s inflation forecast is anchored to its cumulative inflation history, with the degree of anchoring estimated structurally. The authors model this via an &lt;em&gt;experience-based Kalman filter&lt;/em&gt; in which each agent&amp;rsquo;s forecast combines a common Kalman-filtered signal (derived from food prices) with a cohort-specific reference term built from the cohort&amp;rsquo;s entire prior sequence of expected inflation. The estimated history-weight parameter θ is negative, confirming that agents positively weight their inflation history rather than overreacting to current news — a pattern that holds not only in US SCE and Michigan Survey of Consumers data but also across six European countries in the ECB Consumer Expectations Survey. Embedded in a Blanchard–Yaari perpetual-youth OLG New Keynesian model — where households hold experience-based expectations but firms set prices under rational Calvo frictions — the mechanism produces qualitatively different aggregate dynamics from full-information rational expectations (FIRE): after inflationary shocks, expectations initially underreact (agents anchor to the low-inflation steady state) and then persist well beyond the shock horizon as high inflation is gradually incorporated into cohort memory, generating hump-shaped expectation dynamics. For monetary policy, the optimal Taylor rule must be &lt;em&gt;more aggressive&lt;/em&gt; after cost shocks than under FIRE: an energetic early response prevents the high-inflation episode from entering cohort memories, avoiding a self-reinforcing upward drift in inflation expectations. Applied to the 2021 high-inflation episode, the model predicts that the youngest cohorts — experiencing high inflation for the first time — will exhibit persistently elevated inflation expectations long after the supply shocks that caused the episode have dissipated.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Summary of a forthcoming paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="what-are-the-four-empirical-patterns-in-the-survey-data-and-do-they-hold-outside-the-us"&gt;What are the four empirical patterns in the survey data, and do they hold outside the US?&lt;/h3&gt;
&lt;p&gt;Using the New York Fed&amp;rsquo;s Survey of Consumer Expectations (a monthly panel from 2013), the paper documents four patterns: (i) inflation expectations differ substantially across birth cohorts; (ii) cohort-specific inflation experience is age-clustered; (iii) individual inflation history is positively correlated with individual inflation expectations; (iv) cohorts do not differ in how they update to current information once their own inflation history is controlled for. These patterns hold in the Michigan Survey of Consumers and, with cohort fixed effects, across the ECB Consumer Expectations Survey covering six European countries — suggesting the mechanism is not US-specific.&lt;/p&gt;
&lt;h3 id="how-does-the-experience-based-kalman-filter-work-and-what-does-estimation-yield"&gt;How does the experience-based Kalman filter work, and what does estimation yield?&lt;/h3&gt;
&lt;p&gt;Each consumer&amp;rsquo;s forecast has two components: a standard Kalman filter signal common to all agents (extracted from food price data) and a cohort-specific reference term that is a weighted average of all past expectations formed by that cohort, governed by the parameter θ. Structurally estimated from SCE data using time fixed effects, θ is negative — meaning consumers positively anchor to their inflation history rather than over-extrapolating from current news. In a goodness-of-fit regression, the experience-based Kalman filter predicts observed cohort-level heterogeneity with a slope coefficient of 1.069, dominating lifetime average inflation and lagged inflation as predictors.&lt;/p&gt;
&lt;h3 id="what-is-the-general-equilibrium-model-and-how-do-heterogeneous-expectations-enter-the-is-curve"&gt;What is the general equilibrium model, and how do heterogeneous expectations enter the IS curve?&lt;/h3&gt;
&lt;p&gt;The model is a Blanchard–Yaari perpetual-youth OLG New Keynesian economy. Each surviving cohort solves a standard Euler equation using the experience-based expectations operator rather than rational expectations, yielding a history-dependent IS curve in which the effective real rate depends on the weighted average of each cohort&amp;rsquo;s reference inflation. Intermediate goods producers set prices under Calvo frictions with &lt;em&gt;rational&lt;/em&gt; expectations, yielding a standard New Keynesian Phillips curve. The central bank follows a Taylor rule. The IS curve&amp;rsquo;s history-dependence means that past inflationary episodes — absorbed into cohort memory — affect present aggregate demand.&lt;/p&gt;
&lt;h3 id="what-do-the-impulse-responses-show-under-experience-based-versus-fire-expectations"&gt;What do the impulse responses show under experience-based versus FIRE expectations?&lt;/h3&gt;
&lt;p&gt;Under a taste (demand) shock, experience-based expectations generate lower inflation on impact — agents anchor to the low-inflation steady state — but inflation remains elevated for longer as the shock is incorporated into cohort memory. Under a cost (supply) shock, two forces compete: anchoring to the steady state damps initial price pressure, but rational firms can raise prices by more because the IS curve becomes more inelastic; the net effect requires a stronger interest rate response than under FIRE. In both cases, household expectation dynamics are hump-shaped — initial underreaction followed by gradual build-up — consistent with evidence in Angeletos et al. (2021) and Pfajfar and Roberts (2018).&lt;/p&gt;
&lt;h3 id="how-does-the-optimal-taylor-rule-change-under-experience-based-expectations"&gt;How does the optimal Taylor rule change under experience-based expectations?&lt;/h3&gt;
&lt;p&gt;After a cost shock the central bank should be more aggressive than under FIRE. The social cost of tolerating a transitory inflationary episode is much higher under experience-based expectations because it permanently shifts cohort memory upward, creating self-reinforcing dynamics in future periods. An aggressive early response prevents the episode from entering cohort references. After a taste shock the optimal response is similarly strong under both FIRE and experience-based expectations, so the memory channel adds little incremental urgency on the demand side.&lt;/p&gt;
&lt;h3 id="what-does-the-model-predict-about-the-2021-high-inflation-episode"&gt;What does the model predict about the 2021 high-inflation episode?&lt;/h3&gt;
&lt;p&gt;Feeding the model with actual monthly data through December 2021, average inflation expectations post-2021 are predicted to be both higher and more persistent under experience-based expectations than under FIRE or diagnostic expectations. Young cohorts, who experienced only low inflation in the 2010s, are updating their memory of inflation upward for the first time, creating a cohort-specific anchoring shift. The model implies that the 2021 episode could have long-lasting effects on consumer price expectations even if the supply shocks that caused it are fully transitory.&lt;/p&gt;</description></item><item><title>Income taxation across countries</title><link>https://macropaperwarehouse.com/papers/income-taxation-across-countries/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/income-taxation-across-countries/</guid><description>&lt;p&gt;The paper provides the most comprehensive cross-country empirical characterisation of effective income tax functions to date, estimating the two-parameter log-linear tax function — pioneered by Feldstein (1969) and applied in structural macroeconomics by Heathcote, Storesletten, and Violante (2017) — for over thirty countries across approximately four decades using harmonized household microdata from the Luxembourg Income Study (LIS). The log-linear function fits income tax systems worldwide with median R² of 0.984 (mean 0.976), extending a finding previously known mainly for the United States to essentially all LIS countries. Five main facts emerge. First, income tax progressivity (τ) and average tax level (λ) are positively correlated across countries: Northern European countries with the highest average tax rates — Belgium, Netherlands, Germany, Finland — also have the highest progressivity; countries such as Brazil, Colombia, Peru, and the Republic of Korea exhibit effectively flat income taxes (τ near zero or negative) despite progressive statutory codes, because actual enforcement and effective coverage are limited. Second, progressivity increases with economic development: richer countries systematically operate more progressive income tax systems, consistent with greater institutional capacity to enforce income taxation. Third, progressivity differs significantly by family structure: married couples with children face the highest progressivity across countries, single households without children the lowest, reflecting child tax credits, joint filing rules, and other family-based provisions. Fourth, the United States ranks toward the lower end of progressivity among high-income countries, with τ ≈ 0.046 in 2010; Belgium, Finland, Germany, Iceland, Ireland, the Netherlands, and Spain are more than twice as progressive as the US. Fifth, transfers account for most redistribution: the combined tax-and-transfer system&amp;rsquo;s progressivity substantially exceeds that of income taxes alone, indicating that analyses focusing solely on income tax progressivity understate total redistributive effort.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Summary of a forthcoming paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="what-is-the-log-linear-tax-function-and-why-does-the-paper-adopt-it-for-cross-country-comparison"&gt;What is the log-linear tax function and why does the paper adopt it for cross-country comparison?&lt;/h3&gt;
&lt;p&gt;The log-linear tax function expresses post-tax income as T(y) = λy^(1−τ) + (1−λ)y, equivalent to log(y − T(y)) = α + (1−τ)log(y), where τ measures progressivity (τ &amp;gt; 0: marginal rates rise with income; τ = 0: flat tax) and λ captures the average tax level. The function is attractive because it (a) is used widely in structural macro models, enabling direct calibration from these estimates; (b) can be estimated consistently from microdata with just two parameters; (c) permits clean cross-country and over-time comparisons. A richer functional form would sacrifice the comparability across 30+ countries and 40 years of data.&lt;/p&gt;
&lt;h3 id="how-well-does-the-log-linear-function-fit-income-tax-systems-across-all-countries-in-the-sample"&gt;How well does the log-linear function fit income tax systems across all countries in the sample?&lt;/h3&gt;
&lt;p&gt;Very well. Across all 200+ country-wave regressions, the median R² is 0.984 and the mean is 0.976. The fit is robust to different income definitions, imputation methods, and country-specific data sources. This extends the well-known finding for the United States (HSV 2017) to countries with very different income tax structures, suggesting the log-linear form is an adequate empirical approximation to real-world progressive tax schedules worldwide.&lt;/p&gt;
&lt;h3 id="what-is-the-cross-country-pattern-of-progressivity-in-2010"&gt;What is the cross-country pattern of progressivity in 2010?&lt;/h3&gt;
&lt;p&gt;Spain (τ ≈ 0.157), Belgium (τ ≈ 0.139), and the Netherlands (τ ≈ 0.127) have the most progressive income taxes in 2010. The Republic of Korea (τ ≈ −0.006) is slightly regressive in effective terms, along with Peru (τ ≈ 0.013) and other low-income countries where income tax coverage is limited. The United States has τ ≈ 0.046, placing it toward the lower end of progressivity among developed countries. In terms of the Progressivity Tax Wedge (PTW) — how much marginal tax rates rise between the average income earner and one at twice the average — Belgium, Finland, Germany, Iceland, Ireland, the Netherlands, and Spain are more than twice as progressive as the US.&lt;/p&gt;
&lt;h3 id="how-does-income-tax-progressivity-relate-to-economic-development"&gt;How does income tax progressivity relate to economic development?&lt;/h3&gt;
&lt;p&gt;The paper documents a systematic positive relationship: richer countries (measured by median income, mean income, or GDP per capita) have more progressive income tax systems. Low-income countries like Peru and Guatemala collect most revenue through goods and services taxes and exhibit low income tax progressivity; high-income Northern European countries have both high tax capacity (the institutional ability to enforce income taxation) and high progressivity. This complements the tax capacity literature and suggests that the development-progressivity link operates through institutional channels, not solely through political demand for redistribution.&lt;/p&gt;
&lt;h3 id="how-does-family-structure-affect-income-tax-progressivity"&gt;How does family structure affect income tax progressivity?&lt;/h3&gt;
&lt;p&gt;Estimated separately for four household types — single without children, single with children, married without children, married with children — progressivity is consistently highest for married couples with children and lowest for single households without children. This pattern holds across countries and over time, reflecting child tax credits, joint filing rules, and other family-based tax provisions that steepen the effective marginal tax schedule. The paper quantifies this heterogeneity by family type, filling a gap in cross-country comparisons that typically focus on single households without children.&lt;/p&gt;
&lt;h3 id="what-do-transfers-add-to-the-redistributive-picture-and-what-is-the-implication-for-welfare-analysis"&gt;What do transfers add to the redistributive picture, and what is the implication for welfare analysis?&lt;/h3&gt;
&lt;p&gt;When estimating a combined tax-and-transfer function (post-tax-and-transfer income regressed on pre-tax income), the progressivity of the combined system substantially exceeds that of income taxes alone. Countries with high income tax progressivity also tend to have high transfer system progressivity, but the transfer channel dominates. Analyses that focus solely on the income tax progressivity parameter τ therefore understate the total redistributive effort of high-income countries and overstate the tax-side role. This has direct implications for welfare analyses and cross-country comparisons using the log-linear framework.&lt;/p&gt;</description></item><item><title>On the Geographic Implications of Carbon Taxes</title><link>https://macropaperwarehouse.com/papers/on-the-geographic-implications-of-carbon-taxes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/on-the-geographic-implications-of-carbon-taxes/</guid><description>&lt;p&gt;Standard analyses of unilateral carbon taxes ignore the spatial reallocation of economic activity induced by the policy, leading them to overstate the costs and understate the effectiveness of such taxes. Using a multi-sector dynamic Spatial Integrated Assessment Model (S-IAM) calibrated to over 17,000 locations worldwide, the paper shows that a European Union carbon tax introduced unilaterally — if accompanied by &lt;em&gt;local rebating&lt;/em&gt; of tax revenues to the residents of the taxing region — expands the size of the EU economy and improves global welfare. The mechanism: the carbon tax falls disproportionately on non-agricultural, energy-intensive sectors and effectively shifts part of its incidence onto trading partners via higher goods prices, while the rebate accrues only to EU residents, raising EU income per capita and attracting migrants. Under a 40 USD/tCO₂ EU tax with local rebating, EU real income rises by 0.46% in 2021 and EU population rises by 1.1%; without rebating, EU real income falls by 4.96% in 2021. EU CO₂ emissions fall by 41% by 2100, but global emissions fall by only 3% due to carbon leakage — production shifts to US, Japanese, and other unregulated regions, raising US and Japanese emissions by 12% on impact. Global real income per capita declines by 0.63% by 2100 without rebating, while global welfare improves with local rebating as economic activity concentrates in high-productivity non-agricultural regions. Rebating revenues to developing countries instead of locally slows migration to the EU, reduces the spatial efficiency gain, and deteriorates global welfare relative to the local-rebating benchmark.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Summary of a forthcoming paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="why-does-the-standard-analysis-miss-the-spatial-channel-and-what-formal-result-does-the-paper-establish"&gt;Why does the standard analysis miss the spatial channel, and what formal result does the paper establish?&lt;/h3&gt;
&lt;p&gt;The paper proves formally that a unilateral carbon tax with local rebating can be expansionary for the implementing region: the tax shifts part of its incidence onto trading partners (by raising the price of goods in which the region has comparative advantage) while the rebate is returned only to locals, increasing local income per capita and attracting migrants. Standard models without trade, migration, and agglomeration externalities predict only contraction. The quantitative magnitude depends on whether the pre-existing spatial equilibrium is efficient; because it is generally not — due to agglomeration externalities and knowledge spillovers — the tax-induced reallocation can improve spatial efficiency.&lt;/p&gt;
&lt;h3 id="what-is-the-s-iam-and-what-makes-it-suited-to-quantifying-this-channel"&gt;What is the S-IAM, and what makes it suited to quantifying this channel?&lt;/h3&gt;
&lt;p&gt;The S-IAM (Conte et al. 2021) features over 17,000 locations with positive land mass, two sectors (agriculture and non-agriculture), multi-sector technology diffusion, trade subject to geography-specific iceberg costs, and migration subject to bilateral moving costs. The model is dynamic (2000–2100) and calibrated to observed sectoral specialization, trade flows, and income levels. Energy use generates CO₂ emissions that cause temperature increases reducing agricultural productivity differentially across latitudes, integrating the climate feedback with the economic geography. Without migration and agglomeration, the expansionary channel is absent.&lt;/p&gt;
&lt;h3 id="what-happens-to-eu-sectoral-specialization-under-the-two-rebating-regimes"&gt;What happens to EU sectoral specialization under the two rebating regimes?&lt;/h3&gt;
&lt;p&gt;Without rebating: the carbon tax erodes the EU&amp;rsquo;s comparative advantage in non-agriculture (which is more energy-intensive), shifting production toward agriculture; EU non-agricultural output falls 3.44% on impact, agricultural output rises 0.86%. With local rebating: the rebate disproportionately benefits non-agricultural regions (which pay more tax), raising their income per capita and drawing workers from the agricultural EU periphery; non-agricultural output grows while agriculture declines. The result is a spatial recentralization around the EU&amp;rsquo;s non-agricultural core, strengthening the high-productivity cluster.&lt;/p&gt;
&lt;h3 id="what-are-the-precise-global-welfare-effects-of-local-versus-alternative-rebating"&gt;What are the precise global welfare effects of local versus alternative rebating?&lt;/h3&gt;
&lt;p&gt;Under a 40 USD/tCO₂ EU tax with local rebating: EU real income rises 0.46% in 2021, EU population rises 1.1%, global welfare improves. Under no rebating: EU real income falls 4.96% in 2021, global real income per capita declines 0.63% by 2100, US and Japanese emissions rise 12% on impact due to carbon leakage, EU emissions fall 41% by 2100 while global emissions fall only 3%. Under rebating to developing countries: migration to the EU slows (developing countries become relatively more attractive), the spatial efficiency gain is smaller, and global welfare declines relative to local rebating.&lt;/p&gt;
&lt;h3 id="how-does-the-eu-carbon-tax-affect-sub-saharan-africa-and-the-developing-world"&gt;How does the EU carbon tax affect sub-Saharan Africa and the developing world?&lt;/h3&gt;
&lt;p&gt;Without rebating: sub-Saharan African real income per capita declines 2.36% by 2100, as the EU&amp;rsquo;s shift toward agriculture raises agricultural prices while simultaneously directing fewer imports toward agricultural exporters; South and East Asian real income per capita falls 1.35%. With local rebating: the EU&amp;rsquo;s shift toward non-agriculture reduces demand for agricultural imports, again hurting agricultural exporters. In both scenarios, equatorial and agricultural-exporting regions lose in the short-to-medium run; climate change mitigation benefits these regions in the very long run but the economic geography effect dominates over the 2100 horizon.&lt;/p&gt;
&lt;h3 id="how-do-results-for-a-us-unilateral-carbon-tax-compare-to-the-eu-case"&gt;How do results for a US unilateral carbon tax compare to the EU case?&lt;/h3&gt;
&lt;p&gt;A US unilateral carbon tax with local rebating generates qualitatively similar results: the US economy expands, population increases, and activity concentrates in the non-agricultural core. Without an EU tax, the EU incidentally benefits from a US carbon tax — US real income per capita rises 0.17% by 2100 under a US-only tax because the tax shifts activity toward non-agricultural regions including the EU — illustrating that unilateral action generates spatial spillovers beyond standard carbon-leakage accounting.&lt;/p&gt;</description></item></channel></rss>