<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Growth | Macro Paper Warehouse</title><link>https://macropaperwarehouse.com/topics/growth/</link><atom:link href="https://macropaperwarehouse.com/topics/growth/index.xml" rel="self" type="application/rss+xml"/><description>Growth</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Jan 2026 00:00:00 +0000</lastBuildDate><item><title>A Macro Study of the Unequal Effects of Climate Change</title><link>https://macropaperwarehouse.com/papers/a-macro-study-of-the-unequal-effects-of-climate-change/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/a-macro-study-of-the-unequal-effects-of-climate-change/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper develops a macro heterogeneous-agent model to quantify the distributional welfare impacts of higher temperatures from climate change across income groups in the United States. The motivation is that existing macro climate-economy models either abstract from heterogeneity entirely or focus on spatial heterogeneity across regions rather than income heterogeneity within regions. The paper fills this gap by modeling how the welfare consequences of temperature change depend on both the region a household lives in and its position in the income distribution.&lt;/p&gt;
&lt;p&gt;The model is calibrated to the US using five data sources: NIPA accounts from the BEA (averaged 1997–2020), the 2015 Residential Energy and Consumption Survey (RECS), PRISM climate data (1950–2022), a proprietary product-level data set of over 1,000 heaters, air conditioners, and heat pumps scraped from ecomfort.com in fall 2023, and county-level climate projections for year 2100 under RCP 8.5 from Rasmussen et al. (2016). The US is divided into five regions (cold, cool, mild, warm, and hot) of approximately equal population based on average county temperature. The quantitative exercise compares two stationary equilibria: a contemporary equilibrium using the current temperature distribution and a climate-change equilibrium using the projected 2100 distribution under RCP 8.5 (a no-large-scale-climate-policy scenario). Welfare is measured using the consumption-housing equivalent variation (CHEV), defined as the percent increase in consumption and housing a household would require in every period in the contemporary equilibrium to be indifferent between the two equilibria.&lt;/p&gt;
&lt;p&gt;Households adapt to temperature through two channels: an intensive margin (adjusting energy use for heating and cooling given existing equipment) and an extensive margin (deciding whether to purchase a heater, air conditioner, or heat pump, each carrying a fixed cost). The production functions for heating and cooling are estimated by OLS on the product-level data set, yielding equipment exponents of 0.35 (air conditioners), 0.28 (heaters), and 0.27 (heat pumps), and energy exponents of 0.77, 0.86, and 0.85, respectively, with R-squared values of 0.97, 0.79, and 1.00. A key analytical insight from a stylized model is that the outdoor temperature acts as a &amp;ldquo;transfer from nature&amp;rdquo; to households — warmer days in cold weather and cooler days in hot weather reduce the energy households must purchase, augmenting real income. Because this transfer is a larger share of income for lower-income households, its changes are distributionally regressive when the transfer falls (hotter regions warming further) and progressive when it rises (colder regions warming).&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. Among middle- and high-income households, climate change generates progressive welfare gains in colder regions — ranging from +0.71 percent of consumption-and-housing for households in the third income decile in the cool region to near-zero for the highest income households — and regressive welfare losses in hotter regions, ranging from −1.85 percent for third-decile households in the warm region to near-zero for high-income households. These patterns are driven by the intensive margin (changes in transfers from nature). For low-income households, the pattern reverses: low-income households in colder regions suffer welfare losses (the dominant effect is that climate change forces them to purchase their first air conditioner), while some low-income households in hotter regions experience welfare gains (they can forgo purchasing a heater). Climate change raises the Gini coefficient on lifetime welfare by 1.02, 1.01, and 0.50 percent in the cold, cool, and mild regions, and reduces it by 0.09 and 0.21 percent in the warm and hot regions. Aggregate welfare effects from the heterogeneous-agent model substantially exceed what a representative-agent model would imply: for example, in the mild region, climate change reduces aggregate welfare by 0.65 percent in the baseline but only 0.17 percent in the representative-agent version.&lt;/p&gt;
&lt;p&gt;Policy experiments reveal: (1) Fully offsetting the welfare costs of climate change for the lowest-income households would require government spending on energy assistance to more than double (a factor of 2.2 increase), with the largest increases concentrated in colder regions. (2) A universal heat-pump mandate eliminates the extensive-margin channel, producing monotonically progressive welfare gains in colder regions and monotonically regressive welfare losses in hotter regions across all income deciles. (3) Heat-pump cost parity with heaters largely increases adoption and moderates welfare costs, but low-income households in the hot region see limited improvement because they still prefer air conditioners. (4) Accounting for temperature effects on the labor productivity of outdoor workers (roughly 8 percent of the workforce, concentrated at lower incomes) amplifies welfare costs in hotter regions and moderates them in colder regions, with magnitudes tied to the share of workers affected.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper is a calibrated structural model rather than an empirical identification exercise. Identification in the sense of parameter estimation comes from two sources: (1) OLS estimation of heating and cooling production functions on cross-sectional product-level data, where manufacturers measure capacity and efficiency under standardized conditions, limiting TFP endogeneity concerns that plague aggregate production function estimation; and (2) internal calibration of remaining parameters to match a set of moments from RECS 2015 and NIPA. Threats to the structural analysis include the assumption that households treat housing and equipment as flow (rental) choices rather than durable stocks, abstracting from switching costs and adjustment costs over the transition — the paper explicitly notes this limits the analysis to long-run stationary equilibria. The small-open-economy assumption for capital removes domestic capital-market clearing as a constraint. The calibration uses 2015 RECS (not 2020) to avoid COVID-19 distortions to cooling budget shares. The paper abstracts from amenity values of outdoor temperature, mortality from temperature exposure (approximately 0.04 percent of US deaths from 1999–2020), and spatial migration responses.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-core-mechanisms-and-how-are-they-distinguished"&gt;Q2. What are the two core mechanisms and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;The two mechanisms are the intensive margin (how much energy to use given existing equipment) and the extensive margin (whether to purchase heating or cooling equipment at all). The paper distinguishes them analytically using the simple model, which isolates the intensive margin by assuming all households have equipment. The intuition from the simple model — outdoor temperature as a transfer from nature — explains why welfare effects are progressive in regions where climate change makes temperatures more moderate (transfers rise) and regressive where temperatures become more extreme (transfers fall). The extensive margin is then added in the quantitative model through fixed costs of heater, air conditioner, and heat pump equipment. The paper shows that climate change affects specialization favorability (the degree to which a temperature distribution favors concentrating on only heating or only cooling equipment), and that this extensive-margin channel is most important for lower-income households who are near a corner solution of specializing in only one type of equipment. The heat-pump-mandate counterfactual is used to isolate the intensive-margin channel: when all households use heat pumps in both equilibria, the extensive-margin decision is unchanged by climate change, and all welfare effects are driven purely by transfers from nature.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented-across-income-groups-and-regions"&gt;Q3. What heterogeneity is documented across income groups and regions?&lt;/h3&gt;
&lt;p&gt;Welfare effects vary dramatically in both sign and magnitude. Among middle- and high-income households, climate change generates progressive welfare gains in colder regions (e.g., +0.71 percent CHEV for third-decile households in the cool region, falling toward zero at the top) and regressive welfare losses in hotter regions (e.g., −1.85 percent CHEV for third-decile households in the warm region, again near-zero at the top). For low-income households, the pattern reverses: they experience welfare losses in colder regions (forced to buy first air conditioner) and welfare gains or smaller losses in hotter regions (can forgo purchasing a heater). Figure 2 in the paper shows these crossing patterns by income decile for all five regions simultaneously. The Gini coefficient changes by +1.02% (cold), +1.01% (cool), +0.50% (mild), −0.09% (warm), and −0.21% (hot). Migration incentives also differ: high-income households gain incentives to move to cooler regions (driven by transfers from nature), while low-income households gain incentives to move to warmer regions (driven by specialization changes).&lt;/p&gt;
&lt;h3 id="q4-what-is-the-transfers-from-nature-concept-and-why-does-it-produce-differential-welfare-effects"&gt;Q4. What is the &amp;rsquo;transfers from nature&amp;rsquo; concept and why does it produce differential welfare effects?&lt;/h3&gt;
&lt;p&gt;The paper formalizes the idea that outdoor temperature provides free heating or cooling that substitutes for costly purchased energy. On a cold day with outdoor temperature ζ, nature provides ζ degrees of heating for free, effectively augmenting household income by p_eh * ζ (the value of that heating at market prices). This transfer is identical in absolute terms for all households regardless of income, but it is a larger fraction of income for low-income households, so its loss or gain has greater proportional welfare impact on them. This parallels the progressivity of lump-sum transfers in public finance: losing a dollar matters more when income is lower. Consequently, when climate change moves a region to more moderate temperatures (colder regions), the resulting increase in transfers from nature is progressive — lower-income households gain proportionally more. When climate change moves a region to more extreme temperatures (hotter regions), the decrease in transfers is regressive — lower-income households lose proportionally more. The amenity value of outdoor temperature (distinct from the heating/cooling transfer) is abstracted from in the quantitative model on the grounds that, per the simple model, it does not affect the cross-income distribution of welfare changes if preferences over amenities are uncorrelated with income.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-extensive-margin-generate-the-reversal-of-welfare-effects-for-low-income-households"&gt;Q5. How does the extensive margin generate the reversal of welfare effects for low-income households?&lt;/h3&gt;
&lt;p&gt;The extensive margin works through what the paper calls &amp;lsquo;specialization favorability.&amp;rsquo; When a temperature distribution is dominated by cold days, households can optimally purchase only heater equipment, avoiding the additional fixed cost of an air conditioner; the reverse holds in hot climates. Climate change reduces the specialization favorability index in colder regions by adding more hot days, and increases it in hotter regions by reducing cold days. The welfare impact of moving between a corner solution (one type of equipment) and an interior solution (two types of equipment, or a heat pump) tends to be larger than moving between two interior solutions. In the cold region, climate change causes the majority of households in the bottom three income deciles to transition from not having air conditioning to having it (Figure 5, left panel). The fixed cost of buying an air conditioner for the first time exceeds the intensive-margin gains from more moderate temperatures, producing net welfare losses. In the hot region, many second-through-fourth decile households move from having heat in the contemporary equilibrium to not having heat in the climate-change equilibrium (Figure 5, right panel), saving the fixed cost and producing net welfare gains despite more extreme temperatures.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-model-calibrated-and-what-is-the-quality-of-fit"&gt;Q6. How is the model calibrated and what is the quality of fit?&lt;/h3&gt;
&lt;p&gt;Externally calibrated parameters include: capital income share α = 0.26 (Kiyotaki et al., 2011), depreciation rate δ = 0.066, interest rate r* = 0.04, CRRA coefficient σ = 2, bliss point temperature ζ* = 18°C, labor productivity process (ρ = 0.97, σ²_ε = 0.02, σ²_ξ = 0.66 from Kaplan, 2012), and production function exponents estimated from the ecomfort.com data. Internally calibrated parameters are jointly chosen to match: wealth-to-output ratio (3.0), housing-to-non-housing capital ratio (0.88), average heating budget share for non-heat-pump households (0.014), average cooling budget share (0.0055), energy budget share for heat-pump households (0.014), fractions of households with heating (0.95), cooling (0.86), and heat pumps (0.09), the ratio of energy budget shares between the fifth and first income quintile (0.12), the ratio of energy expenditures between high and low income (1.72), and energy assistance as a fraction of energy expenditures (0.83). Table 3 shows the model matches all targeted moments closely. External validation (untargeted moments) shows the model also replicates the associations between heating/cooling degree days and budget shares, equipment ownership, and indoor temperature choices, with similar signs and magnitudes to RECS 2015 data. One limitation is that the model overstates heat pump adoption (17% in model vs. 9% in 2015 RECS, though 14% in 2020 RECS), because it treats modern cold-weather-capable heat pumps as the default.&lt;/p&gt;
&lt;h3 id="q7-what-do-the-policy-counterfactuals-show"&gt;Q7. What do the policy counterfactuals show?&lt;/h3&gt;
&lt;p&gt;Four policy experiments are analyzed. First, scaling energy assistance proportionally to energy needs under climate change reduces assistance by 24% in cold and 20% in cool regions (where transfers from nature increase) and raises it by 9%, 36%, and 79% in mild, warm, and hot regions. Government spending increases by 25%, but the program remains smaller than 0.02% of output. This scaling partially offsets but does not eliminate the distributional distortions. Fully eliminating welfare costs for the lowest-income households would require multiplying energy assistance spending by a factor of 2.2. Second, a universal heat-pump mandate (analogous to natural gas bans like New York, Washington DC, or California&amp;rsquo;s post-2030 ban on natural gas furnaces) eliminates all extensive-margin effects because all households hold heat pumps in both equilibria. Under this mandate, climate change produces monotonically progressive welfare gains across all income groups in colder regions and monotonically regressive welfare costs in hotter regions. Third, heat-pump cost parity with heaters drives near-universal heat pump adoption and broadly moderates welfare costs relative to baseline, but the lowest-income households in the hot region see limited improvement because they still prefer air conditioners over heat pumps even at cost parity (air conditioners are cheaper and heat pumps&amp;rsquo; heating advantage is less valuable in an already-hot, increasingly-hotter climate). Fourth, the labor productivity extension (using the Richardson construction cost database adjustment factor of 1% per degree outside 40°F–85°F) implies that climate change raises low-income productivity by 2% in cold and 0.9% in cool regions and reduces it by 0.1%, 1.1%, and 2.2% in mild, warm, and hot regions. These labor-productivity changes modestly moderate welfare costs in colder regions and amplify them in hotter regions for low-income households.&lt;/p&gt;
&lt;h3 id="q8-why-does-income-heterogeneity-matter-for-aggregate-welfare-calculations"&gt;Q8. Why does income heterogeneity matter for aggregate welfare calculations?&lt;/h3&gt;
&lt;p&gt;The paper demonstrates that a representative-agent model substantially underestimates the aggregate welfare cost of climate change in all regions except the hot region. In the cold region, the aggregate CHEV is −1.03% in the baseline but the average (seventh-decile) household experiences small positive welfare effects (+0.19%), and the representative-agent model yields −0.00%. In the mild region, the aggregate is −0.65% but the representative-agent model gives −0.17%. The discrepancy arises because the welfare distribution is skewed: large losses for low-income households in colder regions are not offset by small or negative gains for high-income households, so the average is dominated by the tails. In the hot region the direction reverses: the baseline aggregate benefit (+0.24%) is driven by large gains at the bottom that the representative-agent model (−0.43%) misses entirely. This finding parallels the broader macroeconomics literature showing that income heterogeneity affects the aggregate welfare cost of business cycles, inflation, and asset pricing.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q9. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;The paper sits at the intersection of two literatures. The macro climate-economy literature (Acemoglu et al., 2012; Golosov et al., 2014; Barrage, 2020) typically uses representative-agent models that abstract from heterogeneity. The spatial heterogeneity literature (Cruz and Rossi-Hansberg, 2024; Bilal and Rossi-Hansberg, 2023; Rudik et al., 2022) studies how welfare consequences vary across regions based on their income levels and exposures but not within-region income differences. The within-region inequality literature (Dennig et al., 2015; Kornek et al., 2021; Belfori and Macera, 2022; Douenne et al., 2023) adds heterogeneous fixed income types to integrated assessment models, but does not model endogenous income and wealth distributions. Blanz (2023) is the closest precursor: it uses a standard incomplete-markets model to study food-price effects of climate change in developing countries, but does not model the temperature-equipment-energy production technology. The empirical literature (Hsiang et al., 2017; Park et al., 2018; Doremus et al., 2022) estimates reduced-form relationships between temperature and energy spending by income group, but cannot decompose intensive vs. extensive margin mechanisms or conduct structural policy counterfactuals. The key novel contributions are: (1) endogenous income and wealth heterogeneity within the Bewley-Huggett-Aiyagari tradition, (2) explicit modeling of both margins of temperature adaptation with estimated production functions, and (3) the ability to separately identify the roles of transfers from nature and specialization favorability.&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-are-conducted"&gt;Q10. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The paper reports several robustness checks. First, the main calibration uses the housing exponent γ = 0.1, but Appendix Figure D.1 shows results with γ = 0.4 (the upper bound implied by the RECS regression of energy on square footage, before controlling for quality), finding broadly similar qualitative results. Second, the 2015 RECS is used instead of the 2020 RECS due to COVID-19 distortions to cooling budget shares; the paper notes heating budget shares are similar between the two surveys while cooling shares are materially higher in 2020. Third, external validation of the model on untargeted moments (associations between HDD/CDD and heating/cooling budget shares, equipment ownership, and indoor temperatures) confirms the model&amp;rsquo;s predictive validity. Fourth, the welfare results are computed for both the main five-region model and a representative-agent version, documenting the magnitude of the aggregation bias. Fifth, the labor productivity extension bounds the relevant population (bottom 3% vs. bottom 16% of workers) to bracket the Occupational Requirements Survey estimate of 8% of workers constantly or frequently exposed outdoors.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-scope-conditions-and-limitations-of-the-main-results"&gt;Q11. What are the scope conditions and limitations of the main results?&lt;/h3&gt;
&lt;p&gt;Several important scope conditions apply. The analysis focuses exclusively on the direct effects of higher temperatures in the US; it does not cover other forms of climate damage (sea level rise, storm frequency, drought, wildfire) or effects in other countries. The model is solved for stationary equilibria, so it cannot speak to transition dynamics or the welfare costs of adjustment during the period when households are switching equipment. Housing and equipment are modeled as flow (rental) choices, abstracting from switching costs, adjustment frictions, and the interaction between homeownership and equipment decisions. The model abstracts from the amenity value of outdoor temperature (e.g., preference for pleasant weather), temperature-related mortality (about 0.04% of US deaths, 1999–2020, heavily concentrated among the unhoused population outside the model), and behavioral adaptation beyond energy and equipment choices (migration is analyzed only as a partial equilibrium incentive calculation, not as an equilibrium outcome). The capital market operates as a small open economy, so general equilibrium effects on interest rates are absent. Labor productivity effects of temperature are only explored for low-income workers in the outdoor sector, not for higher-income or indoor workers.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-migration-findings-and-their-caveats"&gt;Q12. What are the migration findings and their caveats?&lt;/h3&gt;
&lt;p&gt;The paper shows that climate change increases incentives for high-income households to migrate to cooler regions (driven by the transfers-from-nature channel — cooler regions offer larger increases in transfers) and increases incentives for low-income households to migrate to warmer regions (driven by the specialization channel — warmer regions allow forgoing heater equipment). The magnitude of the change in migratory pressure for high-income households is much smaller (order of magnitude roughly 0.15 on the paper&amp;rsquo;s scale) than for low-income households (order of magnitude roughly 3 on the same scale). The authors explicitly caveat that this is a partial equilibrium exercise: the model abstracts from the amenity value of temperature (which would reduce pressure to move to warmer regions by reducing the attractiveness of hot destinations) and from other dimensions of climate change (storm risk, fire risk) that would affect migration incentives independently.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Transfers from nature&lt;/strong&gt;: In this paper&amp;rsquo;s framework, outdoor temperature acts as a subsidy equivalent to income: on a cold day, nature provides degrees of heating for free, augmenting household real income by the value of that heating energy; on a hot day, it provides degrees of cooling. The transfer is the same in absolute terms for all households but represents a larger fraction of income for lower-income households, making changes in temperature distributionally progressive (when transfers rise) or regressive (when transfers fall).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extensive margin of temperature adaptation&lt;/strong&gt;: The binary decision of whether to purchase temperature-control equipment — a heater, air conditioner, or heat pump — each carrying a fixed cost. Households at the extensive margin may optimally forego one type of equipment entirely (complete specialization), and climate change can force them to acquire equipment they previously lacked or allow them to drop equipment they previously held.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive margin of temperature adaptation&lt;/strong&gt;: The continuous decision of how much energy to purchase to operate existing heating and cooling equipment in order to achieve a desired indoor temperature, conditional on having that equipment. Changes in the outdoor temperature distribution affect energy expenditures along this margin for all households that already own equipment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Specialization favorability index&lt;/strong&gt;: A region-level index S_n ∈ [0,1] defined as the absolute difference between total degrees of heating need and total degrees of cooling need, divided by their sum. Higher values indicate that the temperature distribution is more dominated by either heating or cooling demand, making it more efficient for households to specialize in a single type of temperature-control equipment rather than purchasing both. Climate change reduces specialization favorability in colder regions and increases it in hotter regions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption-housing equivalent variation (CHEV)&lt;/strong&gt;: The paper&amp;rsquo;s welfare metric: the percentage by which a household&amp;rsquo;s consumption and housing would need to increase in every period of the contemporary equilibrium for the household to be indifferent between remaining in the contemporary equilibrium and living in the climate-change equilibrium. Negative CHEV values indicate welfare losses from climate change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Temperature damage function D(T)&lt;/strong&gt;: A function mapping the deviation of indoor temperature from the bliss point to the fraction of full utility the household receives from housing services. D equals 1 when indoor temperature equals the bliss point (18°C in calibration) and falls below 1 as indoor temperature deviates in either direction, with the rate of decline governed by parameter χ. This function creates the motive to use energy for heating and cooling.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;RCP 8.5&lt;/strong&gt;: As used in this paper, a climate scenario from the CMIP archive representing emissions in the absence of large-scale climate policy, used to construct the 2100 temperature distribution in the climate-change equilibrium. County-level projections come from Rasmussen et al. (2016), probability-weighted across climate models.&lt;/p&gt;</description></item><item><title>Armed conflict exposure and trust: evidence from a natural experiment</title><link>https://macropaperwarehouse.com/papers/armed-conflict-exposure-and-trust-evidence-from-a-natural-experiment/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/armed-conflict-exposure-and-trust-evidence-from-a-natural-experiment/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks how individual-level exposure to internal armed conflict shapes social capital, specifically trust in institutions and trust in people. The question matters because trust is a core component of social capital that underpins cooperation, economic growth, financial development, political participation, and post-conflict recovery; yet the empirical literature is split between studies finding conflict erodes trust and studies finding &amp;ldquo;post-traumatic growth&amp;rdquo; that enhances pro-sociality. The authors argue prior work cannot cleanly identify causal effects because of non-random selection into exposure, attrition from migration/death, and confounding conflict-induced changes in the socio-economic environment.&lt;/p&gt;
&lt;p&gt;The empirical strategy exploits a natural experiment in Turkey: mandatory conscription assigns every male citizen, via a lottery, to a military base, and a significant share are randomly sent to bases in the eastern/south-eastern conflict zone where the state has fought the PKK since 1984. By sampling ex-recruits who live in peaceful western districts, exposure during military service is the respondents&amp;rsquo; only personal contact with the conflict, isolating individual-level effects from environmental confounds. Data come from a field survey of 5,024 randomly selected adult males in 29 western districts in summer/fall 2019 (response rate 83%); eligible men had completed service between 1984 and 2014. Only 5 respondents did not answer the military-service questions.&lt;/p&gt;
&lt;p&gt;Two exposure measures are built. ACE (Exposure to Armed Conflict Environment) is the standardized number of combatant casualties in the county and during the period of a respondent&amp;rsquo;s service, drawn from the Turkish State-PKK Conflict Event Database; its variation comes from four exogenous components (birthdate-driven timing, regulation-driven duration, clash intensity, and lottery-assigned location). TDE (Traumatic Direct Experiences) is a binary indicator equal to 1 if the respondent was wounded in armed clashes or had someone around them killed/hurt; 2% reported being wounded and 15% reported others around them killed or hurt. ACE and TDE correlate only 0.25. Two trust outcomes: Institutional Trust (average of 14 five-point items: army, judiciary, parliament, TV, newspapers, parties, clergy, universities, environmental orgs, charities, police, banks, private companies, EU) and Social Trust (trust in unfamiliar people / strangers). The army was the most trusted institution (~75% high trust vs. 43% for courts, 35% for parliament). Estimation is OLS with age, education, and minority controls, standard errors clustered at the living-block level.&lt;/p&gt;
&lt;p&gt;Main findings: the two exposure types have opposing effects. In the preferred specification including both measures, ACE raises Institutional Trust (about 0.02, significant at 5%) and Social Trust (about 0.03, significant at 5%), while TDE lowers Institutional Trust (about -0.15, 5%) and Social Trust (about -0.11, 1%). ACE is insignificant when TDE is omitted because it then pools traumatized and non-traumatized recruits, biasing it toward zero. There is no significant ACE-by-TDE interaction, so the negative trauma effect is independent of conflict intensity. Effects are similar in sign and magnitude across both trust dimensions, indicating an encompassing change rather than institution-specific distrust. Interactions with time-since-service are insignificant, implying the effects are permanent.&lt;/p&gt;
&lt;p&gt;Mechanism: the authors invoke Janoff-Bulman&amp;rsquo;s (1992) &amp;ldquo;shattered assumptions&amp;rdquo; theory. TDE is positively associated with depression and insecurity indexes, which in turn correlate negatively with both trust measures; ACE is not significantly related to depression/insecurity. There is no significant relationship between exposure and trust in the army, ruling out an accountability mechanism. Heterogeneity by in-group: TDE raises trust in family (coping mechanism) but, like strangers, friends show positive ACE and (insignificant) negative TDE effects, arguing against parochialism as the main driver. Implications: distinguish contextual from direct exposure; design psychological recovery programs for veterans; estimates are likely conservative given the limited 6-18 month exposure window.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Identification relies on Turkey&amp;rsquo;s conscription lottery, which randomly assigns drafted men to military bases, a significant share of which lie in the eastern/south-eastern conflict zone. Because the sample is drawn only from peaceful western districts, service is the respondents&amp;rsquo; sole exposure to the conflict, isolating individual-level effects from conflict-induced changes in the socio-economic environment. ACE&amp;rsquo;s variation comes from four exogenous components: birthdate-driven timing of service, regulation-driven service duration (18 months in the 80s, 15 in 1992, 18 in 1995, 15 in 2003, 12 in 2014), clash intensity around the base, and lottery-assigned location. Threats: (1) non-random base assignment - addressed by balance tests (Table 2) showing no systematic differences in age, ethnicity, or height by conflict-zone assignment; education differs because college graduates are slightly skewed toward western bases (40% of non-college-grads served in the east vs. 30% of college grads), but the difference vanishes when college graduates (9.3% of sample) are excluded, education is controlled in all specs, and a no-college-grad sample (Table A2) is robust; (2) self-selection into dangerous tasks/violence for TDE - addressed by the fact that task assignments are made by command at the start of service before behavior is observed, and Table 3 balance tests show wounded vs. non-wounded respondents do not differ on pre-military characteristics; an alternative TDE (observing a fellow soldier hurt/killed, immune to own risk-taking) yields similar results (Table A1).&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The proposed mechanism is a transformation of fundamental world assumptions (benevolence, meaning, safety of the world) per Janoff-Bulman (1992). Distinguishing tests: (1) TDE affects a broad range of trust dimensions but is NOT significantly related to trust in the army, ruling out an accountability interpretation (which would predict distrust concentrated on state security institutions) and a comradeship interpretation (which would predict effects only on social trust). (2) TDE is positively and significantly associated with depression and insecurity indexes (Tables 7-8), and these indexes are themselves negatively and significantly related to both trust measures, consistent with shattered world assumptions. (3) ACE is not significantly associated with depression/insecurity; the authors note these scales are worded to detect negative states and may miss the positive feelings ACE could elicit, and that indirect environmental exposure plausibly has weaker effects on fundamental beliefs than direct trauma.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;The central heterogeneity is by exposure type: contextual exposure (ACE) raises trust, direct trauma (TDE) lowers it. No significant ACE-by-TDE interaction, so trauma&amp;rsquo;s effect does not depend on conflict intensity. No significant moderation by time since service (Table 6), implying permanent effects. In-group heterogeneity (Table 9, ordered logit): TDE significantly raises trust in family (coefficient 0.26, 5%), interpreted as a coping mechanism of retreating to closest networks; trust in friends shows positive ACE (0.07, 5%) and negative but insignificant TDE, mirroring the stranger result. The similar pattern for strangers and friends argues against parochialism as the primary driver.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;(1) Alternative TDE defined as observing a fellow soldier hurt/killed, more immune to own risk-taking (Table A1) - results unchanged. (2) Excluding college graduates (Table A2) - results unchanged. (3) Tobit specification accounting for the censored nature of trust measures (Table A3) - similar results. (4) Including a conflict-zone dummy and base-district fixed effects (Tables A4-A5) to absorb unobserved location heterogeneity (though the authors note these likely absorb part of the ACE variation, so they are not in the baseline). (5) Separate results for each of the 14 institutional-trust dimensions (Table A6) and excluding one dimension at a time from the composite index - results stable. (6) Alternative standard-error clustering at home-district or region levels - unchanged.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It builds on the draft-lottery natural-experiment tradition (Angrist 1990 on Vietnam; Angrist-Chen 2011; Galiani et al. 2011; Grossman et al. 2015) and the conflict-and-social-capital literature (Rohner et al. 2013; Cassar et al. 2013; Bauer et al. 2016; Kijewski-Freitag 2018). It differs by: (1) cleanly identifying causal effects free of environmental confounds, since trust is measured in untouched western locations rather than in transformed post-conflict settings; (2) carefully separating contextual from direct exposure, which many studies cannot; (3) proposing a novel individual-level psychological mechanism (shattered world assumptions) rather than the economic/institutional-legacy channels (Besley-Reynal-Querol 2014; Nunn-Wantchekon 2011; Grosjean 2014) or the inter-group-competition/parochialism explanation (Bauer et al. 2016). The authors argue the heterogeneity they document can help reconcile the conflicting positive and negative findings in prior literature - prior &amp;lsquo;pro-social&amp;rsquo; effects may reflect coping-driven re-creation of safe social space (consistent with Grosjean&amp;rsquo;s (2014) &amp;lsquo;dark nature&amp;rsquo; of conflict-induced pro-sociality), not genuine restoration of trust.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Two main implications: (1) researchers and policy advisers should carefully distinguish contextual from direct conflict exposure when studying behavioral outcomes; (2) the findings inform the design of psychological and social recovery programs for combat veterans and victimized post-conflict populations. Scope conditions: the study is specific to the Turkish conflict setting and limited to male ex-combatants; it remains open whether effects generalize to women, civilians, or other countries. Because exposure lasted only a pre-determined 6-18 months after which recruits returned to peaceful lives, the authors argue estimates are conservative relative to populations living in protracted conflict environments.&lt;/p&gt;
&lt;h3 id="q7-what-additional-findings-or-caveats-are-noted"&gt;Q7. What additional findings or caveats are noted?&lt;/h3&gt;
&lt;p&gt;The authors report (results not shown) that individuals with traumatic experiences are more likely to participate in political organizations, and cite Kibris-Nelson (2021) that such individuals are more likely to start their own businesses (while being less successful at it), consistent with coping strategies of creating a controllable environment. They concede the mechanism evidence for the positive ACE effect is &amp;lsquo;somewhat less clear&amp;rsquo; than for TDE, and offer an alternative possibility that whether intense-environment survival raises trust may be moderated by how heroically the veteran&amp;rsquo;s social network views his service. The depression subscale is the 6-item Brief Symptoms Inventory; insecurity is an 8-item scale. Roughly 6.5 million of the 15 million men drafted since 1984 are estimated to have served in the conflict zone.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Exposure to Armed Conflict Environment (ACE)&lt;/strong&gt;: A standardized, individual-specific measure of contextual conflict exposure equal to the number of combatant casualties in the county and during the time period of a respondent&amp;rsquo;s military service. It captures immersion in the conflict environment with high geo-temporal precision and is treated as exogenous because its components (birthdate-driven timing, regulation-driven duration, clash intensity, lottery-assigned location) are outside the individual&amp;rsquo;s control.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Traumatic Direct Experiences (TDE)&lt;/strong&gt;: A binary indicator equal to 1 if a respondent was personally wounded in armed clashes or had someone around them killed or hurt during military service. It captures direct, personal experience of violence as distinct from mere presence in a conflict environment; in the sample 2% were wounded and 15% had others around them hurt/killed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Institutional Trust&lt;/strong&gt;: In the paper&amp;rsquo;s sense, the simple average of a respondent&amp;rsquo;s 5-point Likert trust ratings across 14 public and private organizations (army, judiciary, parliament, media, parties, clergy, universities, environmental orgs, charities, police, banks, private companies, EU) - deliberately broad so as not to over-weight state institutions directly tied to the conflict.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Trust&lt;/strong&gt;: A generalized form of trust measured by how much a respondent trusts people they are not familiar with (strangers), rather than the vaguer &amp;lsquo;most people&amp;rsquo; wording, chosen to minimize in-group/out-group and ethnic associations and isolate generalized trust in others.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Shattered assumptions&lt;/strong&gt;: The paper&amp;rsquo;s operative mechanism, drawn from Janoff-Bulman (1992): people hold core assumptions that the world is benevolent, meaningful, and safe; traumatizing experiences shatter these positive assumptions, eroding deeply rooted trust - whereas surviving a dangerous environment without mishap can instead reinforce them. Trust, depression, and insecurity are treated as observable implications of these otherwise-unobservable world assumptions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parochialism / parochial altruism&lt;/strong&gt;: The rival hypothesis (associated with Bauer et al. 2016) that conflict exposure increases in-group favoritism while eroding out-group trust. The paper tests and largely rejects it as the primary driver because ACE raises trust in both strangers and friends and the in-group (family) pattern does not match parochial predictions.&lt;/p&gt;</description></item><item><title>Business Cycle during Structural Change: Arthur Lewis' Theory from a Neoclassical Perspective</title><link>https://macropaperwarehouse.com/papers/business-cycle-during-structural-change-arthur-lewis-theory-from-a-neoclassical-perspective/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/business-cycle-during-structural-change-arthur-lewis-theory-from-a-neoclassical-perspective/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks why the nature of business cycles changes systematically as economies develop and shed their large agricultural sectors. The motivation is both empirical and theoretical. Empirically, countries with large declining agricultural sectors—most prominently China—exhibit business cycle patterns that depart sharply from the textbook procyclical-employment pattern seen in mature economies: aggregate employment is acyclical with respect to GDP, nonagricultural employment is strongly procyclical, agricultural employment is countercyclical, and the labor productivity gap between nonagriculture and agriculture narrows during booms. These cross-country regularities hold in a sample of 63–66 countries using ILO sectoral employment data over 1970–2015, with the correlation between aggregate employment and GDP declining monotonically as the agricultural employment share rises. The cross-country correlation between the agricultural employment share and log GDP per capita is −0.84. For China specifically over 1978–2012, the correlation between HP-filtered agricultural employment and GDP is −0.69, while the correlation for nonagricultural employment with GDP is 0.73. Agricultural employment fell from about 62.4% of total Chinese employment in 1985 to 33.6% in 2012.&lt;/p&gt;
&lt;p&gt;The authors construct a unified neoclassical model of growth, structural change, and business cycles. The economy produces a CES aggregate of agricultural and nonagricultural output (elasticity of substitution epsilon), with agriculture itself being a CES aggregate of modern and traditional sub-sectors (elasticity omega). Modern agriculture uses capital and labor (Cobb-Douglas), whereas traditional agriculture uses only labor. This nested structure means the effective elasticity of substitution between capital and labor in agriculture is variable and declines as the traditional sector shrinks—formalizing the Lewisian surplus-labor mechanism within a neoclassical framework. A time-invariant tax wedge tau on nonagricultural wages captures rural-urban earnings gaps and keeps agriculture inefficiently large.&lt;/p&gt;
&lt;p&gt;The deterministic model is estimated using Simulated Method of Moments on Chinese data from 1985 to 2012, targeting seven moment sequences: employment share in agriculture, capital share in agriculture, agricultural output-to-GDP ratio, agricultural expenditure share, aggregate GDP growth, the aggregate capital-output ratio path, and the change in the productivity gap. Key findings from estimation: the elasticity of substitution between agricultural and nonagricultural goods epsilon is estimated at 3.6 (significantly greater than 1 at 1% level), and the elasticity between modern and traditional agriculture omega is also very large. The estimated subsistence level in a Stone-Geary extension is small (11% of agricultural production in 1985), so nonhomothetic preferences play only a minor quantitative role. Nonagricultural TFP growth gM is estimated at 6.5% per year; modern-agricultural TFP growth gAM at 6.1% per year; traditional-sector TFP growth gS at 0.9% per year. The estimated labor wedge tau implies persistent misallocation.&lt;/p&gt;
&lt;p&gt;Stochastic TFP shocks (VAR(1) for each of the three sectors) are then estimated from observed data by exploiting the model&amp;rsquo;s equilibrium conditions. The persistence parameters are 0.63 (nonagriculture), 0.90 (modern agriculture), and 0.42 (traditional agriculture). The model, simulated 1,000 times starting in 1980, reproduces the salient Chinese business cycle features: the standard deviation of GDP is 1.7% (matching the data), agricultural employment is countercyclical (model correlation with GDP: −0.25; data: −0.23), nonagricultural employment is strongly procyclical (model: 0.99; data: 0.73), and aggregate employment has a low correlation with GDP (model: 0.42; data: 0.10). A variance decomposition shows nonagricultural TFP shocks account for approximately 95% of GDP fluctuations.&lt;/p&gt;
&lt;p&gt;The key mechanism is that a large traditional sector provides an elastic labor supply to nonagriculture at low marginal cost (a neoclassical Lewisian buffer). Positive TFP shocks to nonagriculture draw labor out of traditional agriculture, raising average capital intensity and labor productivity in agriculture—hence the countercyclical productivity gap. As structural change progresses and the traditional sector shrinks, this labor buffer disappears, the effective labor supply elasticity declines, and business cycle properties converge toward those of a standard neoclassical (Hansen-Prescott) economy. Out-of-sample simulations confirm this convergence: the correlation between total employment and GDP rises from around 40% to near 100% as the agricultural employment share falls below 10%. The paper also shows that positive TFP shocks in agriculture slow structural change, consistent with empirical evidence from the Green Revolution (Foster and Rosenzweig 2004; Bustos et al. 2016; Moscona 2018; Jayachandran 2006).&lt;/p&gt;
&lt;p&gt;Elasticity estimates using CES production functions for the US, Japan, and China from consumption value-added data yield epsilon of 2.49, 1.58, and 1.70 respectively, all significantly above unity at the 1% level—supporting the labor-pull interpretation of structural change. The authors find that imposing the symmetry restriction (epsilon = epsilon_ms) used by Herrendorf et al. (2013) replicates their near-zero estimate for the US, but relaxing that restriction reveals the agriculture-nonagriculture elasticity to be large while the manufacturing-services elasticity is near zero.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-four-key-business-cycle-stylized-facts-documented-for-countries-with-large-agricultural-sectors"&gt;Q1. What are the four key business cycle stylized facts documented for countries with large agricultural sectors?&lt;/h3&gt;
&lt;p&gt;The paper documents four regularities that hold across 63–66 countries (ILO data, 1970–2015): (1) aggregate employment is less correlated with GDP and less volatile; (2) agricultural employment is countercyclical; (3) the labor productivity gap (nonagriculture/agriculture) is negatively correlated with nonagricultural employment; (4) consumption is highly volatile relative to GDP. All four are quantitatively documented for China and compared with the US.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-core-theoretical-mechanism-distinguishing-this-paper-from-earlier-structural-change-models"&gt;Q2. What is the core theoretical mechanism distinguishing this paper from earlier structural-change models?&lt;/h3&gt;
&lt;p&gt;The paper adds an internal split of the agricultural sector into modern (capital-using Cobb-Douglas) and traditional (labor-only) sub-sectors that are imperfect substitutes. This nested structure generates a variable effective elasticity of labor supply to nonagriculture: when the traditional sector is large, labor can be released to industry at near-constant marginal cost (a continuous Lewisian surplus), dampening wage and price fluctuations and decoupling aggregate employment from GDP. As the traditional sector shrinks through capital accumulation and differential TFP growth, the effective labor-supply elasticity falls, progressively transforming the economy into a standard neoclassical one.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-handle-the-lack-of-a-steady-state-for-the-business-cycle-analysis"&gt;Q3. How does the paper handle the lack of a steady state for the business cycle analysis?&lt;/h3&gt;
&lt;p&gt;Because structural change is ongoing in China, approximating the model around a balanced growth path is infeasible. The authors instead solve the model recursively over 250 periods back from an assumed one-sector asymptotic balanced growth path (ABGP), using a 27-state Tauchen Markov chain for the three TFP shocks and piecewise linear decision rules on a 75-point grid for each of the two continuous state variables (kappa and kappa-tilde). They simulate 1,000 economies and compute rolling 28-year window statistics, which are then compared to the data.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-identification-strategy-for-the-elasticity-of-substitution-epsilon-and-what-are-the-main-threats"&gt;Q4. What is the identification strategy for the elasticity of substitution epsilon, and what are the main threats?&lt;/h3&gt;
&lt;p&gt;The primary strategy is Simulated Method of Moments on 143 moment conditions from Chinese data 1985–2012 (28 annual observations each for five moment series plus two level/change moments). A second strategy uses IFGNLS estimation of a Stone-Geary demand system for three countries (US, Japan, China) using both consumption value-added (Herrendorf et al. method) and production value-added (GGDC data). The main threats acknowledged: (a) endogeneity—both sides of the demand equations are driven by unobserved productivity and preference shocks with opposite sign implications (addressed by turning to exogenous Green Revolution shocks); (b) measurement error; (c) the symmetry restriction in prior work; (d) the model is closed-economy and abstracts from demand shocks.&lt;/p&gt;
&lt;h3 id="q5-what-role-do-agricultural-tfp-shocks-versus-nonagricultural-tfp-shocks-play-in-gdp-fluctuations"&gt;Q5. What role do agricultural TFP shocks versus nonagricultural TFP shocks play in GDP fluctuations?&lt;/h3&gt;
&lt;p&gt;A variance decomposition shows nonagricultural TFP shocks (ZM) account for approximately 95% of GDP fluctuations in the benchmark economy over 1985–2012. The logic is that positive TFP shocks to ZM reduce misallocation by drawing labor from the (inefficiently large) agricultural sector to nonagriculture, amplifying the GDP response. In contrast, positive TFP shocks to agriculture partially offset the direct productivity gain by worsening misallocation (labor stays in agriculture), so GDP barely responds. In the low-elasticity (epsilon = 0.5) alternative model, agricultural TFP shocks account for about half of GDP fluctuations—one reason the authors reject this alternative.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-models-prediction-for-business-cycle-evolution-as-structural-change-progresses-compare-to-cross-country-evidence"&gt;Q6. How does the model&amp;rsquo;s prediction for business cycle evolution as structural change progresses compare to cross-country evidence?&lt;/h3&gt;
&lt;p&gt;Using rolling 28-year windows of simulated data from 1985 to 2185, the paper documents four monotone transitions as the agricultural employment share falls: (a) the correlation between agricultural employment and the productivity gap falls toward zero; (b) the correlation between agricultural and nonagricultural employment rises from large and negative (around −0.75 for China&amp;rsquo;s current employment share of 40–50%) toward zero; (c) the correlation between total employment and GDP rises from about 40% to nearly 100%; (d) the volatility of employment relative to GDP rises toward the level of mature economies. All four patterns match the cross-country empirical patterns documented in Figure 5.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-labor-push-versus-labor-pull-debate-imply-for-the-estimated-elasticity-and-how-is-it-resolved"&gt;Q7. What does the labor-push versus labor-pull debate imply for the estimated elasticity, and how is it resolved?&lt;/h3&gt;
&lt;p&gt;With epsilon &amp;gt; 1 (gross substitutes), nonagricultural TFP growth attracts labor from agriculture (labor pull), whereas agricultural TFP growth keeps workers on farms and slows structural change. With epsilon &amp;lt; 1 (complements), agricultural TFP growth would instead push workers into industry. The structural estimate epsilon = 3.6 &amp;gt; 1 strongly favors the labor-pull interpretation. This is confirmed by the Green Revolution evidence: Foster and Rosenzweig (2004), Moscona (2018), Bustos et al. (2016), and Jayachandran (2006) all find that positive agricultural TFP shocks slow industrialization and expand agricultural employment—consistent with epsilon &amp;gt; 1 and inconsistent with epsilon &amp;lt; 1.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-are-run-on-the-business-cycle-model"&gt;Q8. What robustness checks are run on the business cycle model?&lt;/h3&gt;
&lt;p&gt;Four robustness exercises: (1) Low elasticity epsilon = 0.5 with a large food subsistence level—this version fails to generate the observed countercyclicality of the productivity gap and implies an empirically incorrect response to agricultural TFP shocks. (2) Sectoral capital adjustment costs (quadratic, kappa = 2.5)—improves the cyclical behavior of aggregate employment and consumption but makes investment too smooth. (3) Raising the persistence of traditional-sector TFP shocks to match that of modern agriculture (phi_S = phi_AM = 0.90)—reduces aggregate labor volatility and makes the relative volatility of employment monotonically increasing with development. (4) Orthogonal shocks (zero cross-sector correlation)—results are negligibly different from the benchmark. These exercises indicate that the qualitative conclusions are robust across specifications.&lt;/p&gt;
&lt;h3 id="q9-how-is-the-productivity-gap-between-nonagriculture-and-agriculture-generated-by-the-model-and-does-it-match-the-data"&gt;Q9. How is the productivity gap between nonagriculture and agriculture generated by the model, and does it match the data?&lt;/h3&gt;
&lt;p&gt;In the model, the productivity gap (nonagricultural output per worker divided by agricultural output per worker) declines with development because the traditional, labor-intensive sector shrinks, raising average labor productivity in agriculture. This is both a long-run trend prediction and a business-cycle prediction: positive TFP shocks to nonagriculture draw workers from the traditional sector, raising agricultural capital intensity and productivity, thereby reducing the gap. The model successfully captures the falling trend in the productivity gap for China. The correlation between the HP-filtered productivity gap and nonagricultural employment in the model is −0.74, close to the empirical value of −0.54 for China. The model predicts lower volatility of the productivity gap than observed in the data.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-estimated-role-of-nonhomothetic-preferences"&gt;Q10. What is the estimated role of nonhomothetic preferences?&lt;/h3&gt;
&lt;p&gt;The authors extend the baseline homothetic CES model to allow Stone-Geary preferences (agricultural good as a necessity). The estimated subsistence level c-bar corresponds to only 11% of agricultural production in 1985, making the income effect through nonhomotheticity quantitatively small. The estimated epsilon falls only marginally when Stone-Geary preferences are introduced. The remaining structural parameters are virtually unchanged. The authors interpret this as evidence that, at the macroeconomic level, technological factors (TFP growth differences and capital accumulation) rather than nonhomothetic preferences are the primary drivers of structural change in China—a finding consistent with Alvarez-Cuadrado and Poschke (2011).&lt;/p&gt;
&lt;h3 id="q11-how-does-this-paper-relate-to-acemoglu-and-guerrieri-2008-and-herrendorf-et-al-2013"&gt;Q11. How does this paper relate to Acemoglu and Guerrieri (2008) and Herrendorf et al. (2013)?&lt;/h3&gt;
&lt;p&gt;The model builds on Acemoglu and Guerrieri (2008) in having capital deepening and differential TFP growth drive reallocation from agriculture to nonagriculture, but adds the traditional sector (absent in Acemoglu-Guerrieri), which generates the Lewisian surplus-labor mechanism and the declining productivity gap. With respect to Herrendorf et al. (2013): their three-sector CES model imposes a common elasticity across agriculture, manufacturing, and services, yielding a near-Leontief (epsilon near zero) estimate for the US. The authors show this estimate is an artifact of the symmetry restriction: when that restriction is relaxed, the agriculture-nonagriculture elasticity is large (2.32–2.49 for the US) while the manufacturing-services elasticity is near zero. The asymmetric three-sector estimates for the US (2.49), Japan (1.58), and China (1.70) are all above unity at the 1% significance level.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-main-limitations-and-open-questions"&gt;Q12. What are the main limitations and open questions?&lt;/h3&gt;
&lt;p&gt;The paper explicitly identifies several limitations: (1) the business cycle analysis is restricted to productivity (TFP) shocks only and does not include demand shocks; (2) the model is closed-economy and ignores trade; (3) the distinction between traditional and modern agriculture is not directly observed in the data—the traditional sector&amp;rsquo;s TFP process is estimated indirectly, introducing potential measurement error that may exaggerate the volatility and understate the persistence of traditional-sector shocks; (4) the prediction that agricultural value added is positively correlated with nonagricultural labor (and negatively with agricultural labor) is inconsistent with Chinese data, a failure the paper acknowledges. Future work is flagged on demand shocks and open-economy extensions.&lt;/p&gt;
&lt;h3 id="q13-what-cross-country-empirical-evidence-beyond-china-is-presented"&gt;Q13. What cross-country empirical evidence beyond China is presented?&lt;/h3&gt;
&lt;p&gt;Using ILO sectoral employment data for 63–66 countries over 1970–2015 (requiring at least 15 consecutive years of observations), the authors document: the correlation between agricultural and nonagricultural HP-filtered employment shifts from positive for countries with small agricultural sectors to strongly negative for countries with large sectors; the correlation between total employment and GDP declines monotonically with the agricultural employment share; the productivity gap is negatively correlated with nonagricultural employment in countries with large agricultural sectors (correlation of −0.54 for China) but near zero in mature economies; consumption volatility relative to GDP declines with development. The US historical time series (1929–2015) shows that before 1960 NBER recessions were associated with reversals in structural change—mirroring today&amp;rsquo;s China—while this pattern ceased after 1960.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Traditional agriculture (subsistence sector)&lt;/strong&gt;: A sub-sector of the agricultural sector that uses only labor (no capital) and produces an imperfect substitute for modern agricultural output. Its presence generates a reserve pool of labor that can move to nonagriculture at low marginal cost, creating the Lewisian surplus-labor property within a neoclassical framework. As the economy develops, this sector is crowded out by capital-intensive modern agriculture.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Modern agriculture&lt;/strong&gt;: A Cobb-Douglas sub-sector within agriculture that uses both capital and labor. Its expansion—crowding out the traditional sector—constitutes the modernization of agriculture. As workers leave the traditional sector, average capital intensity and labor productivity in agriculture rise, generating the procyclical productivity-gap pattern observed in developing economies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Asymptotic Balanced Growth Path (ABGP)&lt;/strong&gt;: The long-run equilibrium toward which the model economy converges, characterized by a fully modernized (traditional sector vanished), small agricultural sector, constant growth rates of sectoral capitals, and standard neoclassical business cycle properties. The paper establishes conditions under which the ABGP is asymptotically stable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor wedge (tau)&lt;/strong&gt;: An exogenous, time-invariant tax on nonagricultural wages that prevents equalization of marginal products of labor across sectors, standing in for a variety of frictions (migration barriers, rural overpopulation, institutional barriers) that keep agriculture inefficiently large. Its presence means that positive TFP shocks to nonagriculture both raise productivity directly and reduce misallocation by drawing workers out of the oversized agricultural sector.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Elasticity of substitution between agriculture and nonagriculture (epsilon)&lt;/strong&gt;: The elasticity governing substitution between agricultural and nonagricultural goods in aggregate CES production. When epsilon &amp;gt; 1 (gross substitutes, as estimated: epsilon = 3.6 for China), positive TFP shocks to nonagriculture pull labor from agriculture (labor-pull structural change), while positive shocks to agriculture slow structural change—consistent with Green Revolution evidence. When epsilon &amp;lt; 1 (complements), the opposite holds, implying counterfactual predictions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Productivity gap&lt;/strong&gt;: The ratio of average labor productivity in nonagriculture to average labor productivity in agriculture. In the model and the data this gap declines over the course of development (because agriculture modernizes and raises its average productivity) and also narrows during booms in countries undergoing structural change (because booms draw workers from low-productivity traditional agriculture). The model relates the gap formally to the ratio of labor income shares: APLM/APLG = (1−tau) × (LISM/LISA)^(−1).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sullying effect of recessions on agriculture&lt;/strong&gt;: The paper&amp;rsquo;s terminology for the pattern—documented empirically for China by Zhang et al. (2001)—whereby recessions induce workers to return to or remain in the agricultural sector, reversing structural change and lowering average agricultural productivity. This is the cyclical analog of the Lewisian adjustment: in downturns, the labor buffer of traditional agriculture absorbs displaced workers, cushioning aggregate employment but impairing agricultural productivity.&lt;/p&gt;</description></item><item><title>Codification, Technology Absorption, and the Globalization of the Industrial Revolution</title><link>https://macropaperwarehouse.com/papers/codification-technology-absorption-and-the-globalization-of-the-industrial-revolution/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/codification-technology-absorption-and-the-globalization-of-the-industrial-revolution/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Why did the First Industrial Revolution (IR) spread to Meiji Japan—and to essentially no other non-Western country—during the first wave of globalization? The paper tests Mokyr&amp;rsquo;s hypothesis that &amp;ldquo;technical literacy,&amp;rdquo; i.e., the codification of engineering, commercial, and industrial knowledge in the local vernacular, was a necessary condition for absorbing IR technologies. The motivating puzzle: after opening to trade (1858) and the Meiji Restoration (1868), 80% of Japanese exports were still primary products as late as ~1883 and real per capita GDP growth was only 0.6%/yr (1870-1883/85); then in a brief 13-year window (1883-1896) the manufacturing export share tripled and stabilized at around 60% of exports until WWII.&lt;/p&gt;
&lt;p&gt;Data and setup: The authors build several novel datasets. (1) A cross-language measure of codification: scraping national/major libraries and WorldCat for technical books (agriculture, applied sciences, commerce, industry, technology) in 33 languages, 1500-1930. (2) &amp;ldquo;British Patent Relevance&amp;rdquo; (BPR): the cosine similarity (TF-IDF, unigrams+bigrams) between the digitized synopses of all British patents 1780-1852 (from Woodcroft 1857) and a hand-curated corpus of 460 English-language 19th-century technical manuals matched to SITC industries. BPR measures the world supply of codifiable IR knowledge by industry and is deliberately not based on what Japan translated (to avoid endogeneity). (3) The first harmonized, bilateral, industry-level trade dataset for the 19th century: 37 regions, 93 industries, quinquennial 1880-1910, built from reporting countries Japan, US, Belgium, Italy. Outcomes are annualized industry export growth ({1880,1885} to {1905,1910}) and, in robustness, productivity/comparative-advantage growth following Costinot et al. (2012) and Amiti-Weinstein (2018).&lt;/p&gt;
&lt;p&gt;Main findings (with magnitudes): A Japanese industry with a one-standard-deviation higher BPR experienced annual export growth ~12 percentage points faster and annual productivity (comparative-advantage) growth ~1.2 percentage points faster (coefficients 0.121*** and 0.012***). Cross-sectionally, the BPR-growth relationship is positive and significant only for Japan and other codifying countries: for non-Japan regions the BPR coefficient is negative (-0.030***), while English-, French-, and the &amp;ldquo;top-4 codified&amp;rdquo; (English/French/German/Italian) regions show positive coefficients (0.042**, 0.032**, 0.078***), smaller than Japan&amp;rsquo;s. Low-income and Asian regions tend negative (divergence), not always significant. Time-series: regressing Japanese export growth from 1875 to varying end-years, the BPR coefficient is negative/significant in the 1875-1880 placebo window (Japan resembled the periphery), flips around 1890, and is positive and significant at 1% by 1895—coinciding with Japan&amp;rsquo;s catch-up in codification.&lt;/p&gt;
&lt;p&gt;Mechanism and the Meiji &amp;ldquo;natural experiment&amp;rdquo;: In 1870, 84% of all technical books were in four languages (English, French, German, Italian); an Arabic-only reader had access to just 71 technical books. Japan started ordinary but codified explosively: technical-book growth jumped from 1.6%/yr (1600-1860) to 8.8%/yr (1870-1900); translated technical books rose from 8 (1500-1860) to 608 by 1900; Japanese technical books in the NDL grew from 706 (1880) to 2,823 (1890). State provision solved a public-goods/coordination problem: the government built English-Japanese dictionaries (ETSJ 1862/1866, FSEJ 1871) creating standardized Japanese jargon from Chinese glyphs, and 74% of identified technical-book translators (1870-1885) were government employees. Implication: low-cost vernacular access to technical knowledge was a necessary (not sufficient) condition for IR diffusion; where regions were linguistically/geographically distant from Western Europe, codification required state provision (a Gerschenkronian role for the state).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-the-main-threats-to-it"&gt;Q1. What is the identification strategy and the main threats to it?&lt;/h3&gt;
&lt;p&gt;Two-pronged. (1) Cross-sectional: regress region-industry export growth on BPR interacted with region-group dummies, with exporter fixed effects, exploiting that BPR is global (not Japan-specific) and that Japan was uniquely a codifier in the periphery. If codification is the mechanism, only codifying regions should show a positive BPR-growth link. (2) Time-series: exploit the sharp timing of Japanese codification (two well-demarcated periods—pre vs. post technical literacy in the 1880s) by estimating the BPR coefficient on Japanese export growth from 1875 to rolling end-years. The 1875-1880 window serves as a placebo (Japan not yet literate). Main threat is omitted-variable bias: that BPR is correlated with distance to the technology frontier, fundamental comparative advantage, Meiji institutional reforms, or industry steam-intensity. The cross-section addresses the &amp;lsquo;BPR matters everywhere&amp;rsquo; and income/geography confounds; the timing addresses slow-moving confounds (literacy, Tokugawa culture, gradual reforms) since reforms like tax/banking/railroads were mostly in place by 1875, 15-37 years before the BPR effect appears.&lt;/p&gt;
&lt;h3 id="q2-how-are-the-cross-section-and-time-series-results-distinguished-from-confounders-empirically"&gt;Q2. How are the cross-section and time-series results distinguished from confounders empirically?&lt;/h3&gt;
&lt;p&gt;In the cross-section, income terciles (High/Medium/Low) and an Asia dummy are added: no region group replicates Japan&amp;rsquo;s positive pattern; the poorest and Asian regions show negative (divergence) coefficients. The placebo (1875-1880) yields a negative significant BPR coefficient for Japan itself—identical in sign to non-codifiers—then flips positive/significant by 1895, which conventional &amp;lsquo;opening to trade&amp;rsquo; (1858) or &amp;lsquo;Meiji Restoration&amp;rsquo; (1868) stories cannot explain because the effect appears 37 and 27 years later, respectively.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Japan&amp;rsquo;s BPR coefficient is larger (though not always significantly) than that of European codifiers, consistent with Japan having more to learn from British patents as a late industrializer. Among non-codifiers, low-income and Asian regions show negative BPR-growth relationships (divergence). Within codifiers, English- and French-speaking regions individually have positive but smaller and less precisely estimated coefficients; pooling the top-4 codified languages sharpens significance (0.078***). The time-series point estimates for Japan slowly decline after 1900 (not significantly), consistent with Japan shifting to Second Industrial Revolution technologies and becoming less reliant on older IR ones.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;(1) Alternative patent corpora: results are nearly identical using British patents 1853-1879 (full text and AI-summarized) and US patents 1836-1860 and 1861-1879 (coefficients 0.121, 0.116, 0.111, 0.115), though later/US patents lower the R-squared, suggesting the 1780-1852 IR patents best explain Japanese export growth. (2) Productivity instead of exports (Costinot et al. 2012 comparative-advantage growth): qualitatively the same, 1.2 pp/yr for a 1-SD BPR increase, with deterioration in non-codifiers. (3) Confounders: controlling for British-colony status (insignificant) and industry steam-power intensity (French 1860s data) does not affect results. (4) Sample selection: dropping non-manufacturing sectors, excluding Asian destination markets, and dropping major export products (textiles, iron/metal) all leave the results intact, indicating broad-based change.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q5. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;It builds on Mokyr (2011) on &amp;rsquo;technical knowledge&amp;rsquo;/&amp;lsquo;access costs&amp;rsquo; for European industrialization, extending it outside Europe with a Gerschenkronian twist (state as provider of the codification public good). It contributes to the technology-adoption-lags literature (Comin and Hobijn 2010; ~45-year average lags) by offering a friction explanation. It departs from prior Meiji studies (Sussman-Yafeh 2000; Tang; Morck-Nakamura; Bernhofen-Brown) that found banking, railroads, constitutional/monetary reforms had little measurable growth impact—offering codification as the resolution to &amp;lsquo;what drove the Meiji Miracle,&amp;rsquo; consistent with Broadberry et al. (2025) dating Japan&amp;rsquo;s convergence to ~1890 driven by manufacturing productivity. It also extends the knowledge-codification literature (Dittmar 2011; Brown 2024; Abramitzky-Sin 2014) by linking codified vernacular knowledge directly to industry growth rather than indirect outcomes like city growth.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Public provision of technical knowledge in the vernacular can relax a critical bottleneck to industrialization, especially for regions linguistically/geographically distant from the technology frontier where the market undersupplies this public good. Scope conditions: codification is necessary but NOT sufficient. The Meiji model required complementary investments—language/jargon standardization, mass education for absorptive capacity (literacy &amp;gt;90% for army conscripts by 1909; ~40% of elementary class time on science), tacit-knowledge acquisition (2,400 hired foreigners providing 9,506 person-years of training; study-abroad missions), and tax capacity (1873 Land Tax Reform). China&amp;rsquo;s post-1949 codification under Zhou did not yield sustained growth until Maoist policies (Great Leap, Cultural Revolution) ended—&amp;rsquo;the exception that proves the rule.&amp;rsquo;&lt;/p&gt;
&lt;h3 id="q7-what-external-validity-evidence-is-offered-beyond-japan"&gt;Q7. What external-validity evidence is offered beyond Japan?&lt;/h3&gt;
&lt;p&gt;The Meiji codification model was studied and transplanted by Park Chung Hee in South Korea (took power 1961; KIST; researcher counts rose sharply) and Zhou Enlai in China (premier 1949; Russian-language translation drive with USSR as the &amp;lsquo;Britain&amp;rsquo;). In 1950, Japan had ~70,000 technical books, China ~1,000, Korea &amp;lt;100; China surpassed 30,000 by the early 1960s. Korea&amp;rsquo;s per capita income clearly rises after Park; China&amp;rsquo;s codification did not translate into growth until after 1976. These are explicitly presented as suggestive/non-causal, plus appendix discussions of British India and Late Imperial Russia.&lt;/p&gt;
&lt;h3 id="q8-what-are-notable-caveats-and-measurement-choices"&gt;Q8. What are notable caveats and measurement choices?&lt;/h3&gt;
&lt;p&gt;BPR uses British 1780-1852 patent synopses and English manuals deliberately (Britain as IR leader; Japan hired British instructors and used British textbooks; avoids endogeneity from Japanese translation choices). It excludes tacit knowledge and secrecy-protected innovation by design. English codification is likely underestimated (British Library was un-scrapable after a 2023 cyberattack; Library of Congress used instead). German patents/trade data were excluded for coverage/reliability reasons. Linguistic-distance evidence on 1870/1913 GDP is explicitly not interpreted causally. The aggregate growth correlations for Japan, Korea, and China are described as suggestive, not causal.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Codification (of technical knowledge)&lt;/strong&gt;: The creation of a means of transmitting engineering, commercial, and industrial knowledge—via language creation and written messages (manuals, textbooks, dictionaries)—that does not require direct contact between the knowledge originator and the recipient (Cowan and Foray 1997). In the paper&amp;rsquo;s sense it is a non-rival public good that the market undersupplies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Technical literacy / technical knowledge&lt;/strong&gt;: Following Stevens (1995) and Mokyr, the codified engineering, commercial, and industrial practices a practitioner needs to set up and run modern factory-based manufacturing; the paper measures it as the stock of vernacular technical books (agriculture, applied sciences, commerce, industry, technology), excluding theoretical/hard-science and non-firm subjects like medicine.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;British Patent Relevance (BPR)&lt;/strong&gt;: An industry-level measure equal to the cosine similarity (TF-IDF weighted) between the vectorized text of British patent synopses (1780-1852) and the vectorized text of English technical manuals for that industry; it proxies how much codifiable IR knowledge a given industry stood to gain, and is independent of what was actually translated into Japanese.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Access costs&lt;/strong&gt;: Mokyr&amp;rsquo;s (2011) term for the cost of obtaining usable technical knowledge; the paper argues vernacular codification (dictionaries, translations) lowered these costs, and that linguistic distance from English/Latin-Greek roots and physical distance from Europe raised them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Technology absorption / absorptive capacity&lt;/strong&gt;: The complementary conditions needed to use codified knowledge—prior language/jargon development, literacy and scientific training, and tacit knowledge—all of which the Meiji state invested in (dictionaries, compulsory education, &amp;rsquo;live machines&amp;rsquo;/foreign instructors, study-abroad missions).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Defensive modernization (Gerschenkronian state role)&lt;/strong&gt;: The paper&amp;rsquo;s reading that an existential external threat aligned the Japanese elite behind aggressive state-led adoption of Western science, casting the state as the critical agent supplying the codification public good in late industrialization—a Gerschenkronian extension of Mokyr applied outside Europe.&lt;/p&gt;</description></item><item><title>Diet, Economic Development and Climate Change</title><link>https://macropaperwarehouse.com/papers/diet-economic-development-and-climate-change/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/diet-economic-development-and-climate-change/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Food production accounts for roughly one-third of global greenhouse gas (GHG) emissions, and richer nations contribute disproportionately through meat-intensive diets and input-intensive farming. This paper asks how much of that disparity will be exported to the developing world as it grows, and which policies can most cost-effectively reduce agricultural emissions during that transition. The answer requires separately identifying two distinct channels—demand-side dietary change and supply-side technological change—and tracing their general equilibrium consequences through global food markets.&lt;/p&gt;
&lt;p&gt;The authors build a quantitative multi-country general equilibrium model calibrated to 90 countries (plus a rest-of-world aggregate) and 47 food products for 2010. The demand side features nested non-homothetic CES preferences, which allow income elasticities to differ across food products—the core mechanism of the nutrition transition. The supply side, built on Farrokhi and Pellegrina (2023), operates at a granular grid-cell level covering the Earth&amp;rsquo;s surface, with producers on each plot choosing both which crop to grow and whether to use a modern, input-intensive (higher-GHG) technology or a traditional, labor-intensive one—the core mechanism of agricultural modernization. GHG emissions are tracked from both production and transportation. Data on calorie intake come from FAO Food Balance Sheets; emissions from Poore and Nemecek (2018) and EDGAR-FOOD; yields from FAO-GAEZ (approximately 1.1 million fields).&lt;/p&gt;
&lt;p&gt;A key methodological contribution is an identification result for income elasticities that requires no price data. In open-economy models, trade shares provide a sufficient statistic for consumer prices, so the model&amp;rsquo;s implicit Marshallian demand equations can be estimated using only expenditure shares and bilateral trade flows—a cleaner identification than prior closed-economy approaches. Structural elasticity estimates are validated against reduced-form regressions that regress product-level log absorption on log GDP per capita interacted with the product&amp;rsquo;s GHG intensity; the cross-method correlation has a slope of 0.64–0.77 and R² of 0.93–0.95.&lt;/p&gt;
&lt;p&gt;Four empirical patterns motivate the model. First, diet composition alone drives large variation in emissions: if the whole world adopted the US diet (holding total calories fixed), the food share of global GHG emissions would rise from 30% to 42%; adopting the Argentinian diet would raise it to 74%; adopting the Ethiopian diet would lower it to 12%. Second, GHG emissions per capita from food rise strongly with GDP per capita (elasticity 0.39 in the cross-section); about one-third of this is a pure scale effect (more calories) and two-thirds is a compositional shift toward higher-emission foods (elasticity of emissions per calorie with respect to GDP per capita is 0.23–0.28). Third, products with higher GHG emissions per calorie have higher income elasticities; a 1% rise in a product&amp;rsquo;s GHG intensity is associated with a 0.17–0.21% higher income elasticity, robust to excluding all meat products. Fourth, emissions from fertilizers and energy use as a share of total agricultural emissions rise with GDP per capita (slope 0.82), indicating that agricultural modernization independently amplifies GHG emissions within each crop.&lt;/p&gt;
&lt;p&gt;Model decompositions reveal that about two-thirds of the cross-sectional correlation between food emissions per capita and GDP per capita is attributable to intrinsic dietary preferences (culture, religion, demographics) rather than to income itself, and about one-half of the correlation for emissions per calorie. This implies that the causal effect of economic growth on emissions is substantially smaller than raw correlations suggest.&lt;/p&gt;
&lt;p&gt;Policy counterfactuals (Table 4) are the paper&amp;rsquo;s centerpiece. A uniform 10% TFP shock across all modern agricultural, non-agricultural, and input producers raises global welfare by 14.9% and increases global agricultural GHG emissions by 5.0% (approximately 0.6 Gt CO₂ from production, 0.004 Gt from transport). Shutting down the nutrition transition channel reduces this emission increase by 28%; shutting down agricultural modernization reduces it by a further 16%; shutting both down reduces it by 42%—so the two mechanisms together account for more than one-third of the growth-induced emission increase. Crucially, ignoring general equilibrium supply responses would overstate the emission impact of economic growth by 100%: higher food demand raises production prices, which dampens both consumption growth and further technology adoption.&lt;/p&gt;
&lt;p&gt;For dietary restrictions: a global no-beef mandate would reduce agricultural GHG emissions by 20%, at a global welfare cost of 0.6%, with large concentrated losses in major beef-producing and consuming countries (Argentina −3–5%; Uruguay −4%). A global vegetarian mandate would reduce emissions by 30% (approximately the same 20% figure is given in the abstract with apparent inconsistency but Table 4 column 3 shows −20% for no-beef and −30% for vegetarian), at a welfare cost of 2.8% globally and with greater inequality impacts for developing countries. Back-of-the-envelope calculations that ignore general equilibrium overstate the emission reductions from dietary restrictions by roughly one-third.&lt;/p&gt;
&lt;p&gt;For food trade policy: raising trade costs enough to cut transportation emissions by 75% reduces total agricultural GHG emissions by 11.9%, but at a global welfare cost of 17.8%—a ratio far worse than dietary policies. The welfare loss is highly unequal: countries in the bottom quartile of the GDP per capita distribution face welfare losses of up to 41% (the abstract states this figure; Table 4 col. 2 shows the Q4/Q1 inequality worsening by 4.9 percentage points in the eat-local scenario). The conclusion is that dietary policies dominate food trade policies on both effectiveness and equity grounds.&lt;/p&gt;
&lt;p&gt;Transportation emissions account for only about 5% of agricultural GHG (0.7 Gt CO₂ vs. 16.5 Gt from production), so policies targeting transport emissions alone have limited aggregate impact.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-for-income-elasticities-and-why-is-it-novel"&gt;Q1. What is the core identification strategy for income elasticities, and why is it novel?&lt;/h3&gt;
&lt;p&gt;Standard non-homothetic CES estimation requires price data because the demand equation depends on price indices. In a closed economy this problem is severe. The authors show that in an open economy, bilateral trade shares provide a sufficient statistic for variety price indices: averaging trade shares across a country&amp;rsquo;s import partners yields a geometric mean of production prices that can be differenced out using fixed effects. The key estimating equation (40) regresses an adjusted expenditure share on log income per capita, with fixed effects absorbing production-price variation through the set of import partners. No price data is needed. This is exact—not an approximation—unlike the approximate methods in Comin et al. (2021) or Caron and Fally (2022), which either impose additional assumptions about price variation across consumer groups or require proxies for crop-specific trade costs such as gravity variables.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-threats-to-identification-and-how-are-they-addressed"&gt;Q2. What are the main threats to identification and how are they addressed?&lt;/h3&gt;
&lt;p&gt;The key concern is that income is correlated with prices and preference shifters that also affect food expenditure shares. In the reduced-form regressions (equation 1), country-year and product-year fixed effects control for country-specific factors (including regional technology change) and global product-specific factors (including product-specific technological progress). In the structural estimation (equation 40), the model&amp;rsquo;s functional form is used to control fully for endogeneity arising through prices, since trade shares substitute out unobservable price indices exactly. The close agreement between reduced-form and structural income elasticity estimates (slope 0.64–0.77, R² 0.93–0.95 in cross-validation) is reassuring that the two quite different identifying assumptions yield similar results. One remaining concern is unobservable preference shifters (ai,k and ã_i,s), which appear as residuals; identification requires income variation orthogonal to these shifters, and the authors follow the precedent of assuming fixed effects are sufficient. Household-level data from Brazil&amp;rsquo;s Consumer Expenditure Survey (POF) bolster the reduced-form patterns using within-country income variation.&lt;/p&gt;
&lt;h3 id="q3-how-are-the-nutrition-transition-and-agricultural-modernization-distinguished-empirically-and-in-the-model"&gt;Q3. How are the nutrition transition and agricultural modernization distinguished empirically and in the model?&lt;/h3&gt;
&lt;p&gt;These are fundamentally different economic mechanisms. The nutrition transition operates through demand: as incomes rise, consumers shift toward food products that, for reasons of taste or nutrition, happen to have higher GHG emissions per calorie. It is a between-product phenomenon captured by non-homothetic income elasticities. Agricultural modernization operates through supply: as wages rise, producers substitute away from labor-intensive traditional technologies toward input-intensive modern technologies (fertilizers, machinery) that emit more GHG per calorie of output, for any given crop. It is a within-product phenomenon captured by the endogenous technology-choice margin in the agricultural production model. In the counterfactual decompositions, the authors shut down each channel independently: the nutrition transition is shut down by setting all within-sector income elasticity parameters (ε_k) equal; agricultural modernization is shut down by fixing the land share in each technology exogenously. Doing so reveals that the nutrition transition accounts for 28% and modernization for 16% of the emission increase from a 10% TFP shock (jointly 42%), with the remainder attributable to scale effects and general equilibrium price responses.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-role-of-general-equilibrium-supply-responses-and-why-do-they-matter-so-much"&gt;Q4. What is the role of general equilibrium supply responses and why do they matter so much?&lt;/h3&gt;
&lt;p&gt;A central finding is that ignoring supply-side equilibrium price responses would overstate the emission impact of economic growth by 100%. The mechanism is straightforward: economic growth raises income and thus food demand, which pushes up production prices (because agricultural supply is upward-sloping due to limited land and heterogeneous productivity across grid cells). Higher prices dampen consumption, which partially offsets the demand-driven emission increase. For dietary restriction policies, back-of-the-envelope calculations that simply remove the GHG attributable to banned food products overstate the emission reduction by roughly one-third, because consumers substitute toward other food products and global agricultural production reorganizes. The model&amp;rsquo;s general equilibrium structure is therefore essential for obtaining credible policy counterfactuals, and a main conclusion of the paper is that the literature&amp;rsquo;s existing back-of-the-envelope calculations in environmental science substantially overstate both the emission risks from growth and the emission benefits from dietary policies.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented-across-countries-and-products"&gt;Q5. What heterogeneity is documented across countries and products?&lt;/h3&gt;
&lt;p&gt;Across countries: diet composition varies enormously. Counterfactual calculations show that if all countries adopted the Argentinian diet (holding total calories fixed), the global food share of total emissions would rise to 74%; adopting the Ethiopian diet would lower it to 12%, compared to the factual 30%. The income elasticity of the agricultural sector as a whole is 0.39, close to Comin et al. (2021)&amp;rsquo;s 0.37. Rich countries have a higher share of modern technology in production, higher fertilizer and energy use per unit of land, higher food GHG per capita, and higher food GHG per calorie. About two-thirds of the cross-sectional gradient in food GHG per capita is attributable to intrinsic preferences rather than income per se. Religion is documented as one driver: Islamic-majority countries show lower preference for pork; Hindu-majority countries show higher preference for lamb, mutton, and poultry relative to other meats. Across products: GHG emissions per 1,000 kcal range from above 35 kg CO₂ for beef and coffee to below 5 kg CO₂ for wheat and rye. Income elasticity parameters (ε_k) range from lowest for staples (yams, sweet potatoes, millet, sorghum, rice) to highest for luxury fruits and vegetables (berries, asparagus, cucumbers, watermelon). Notably, the income-GHG gradient persists after excluding all meat products: vegetables and fruits have higher GHG per calorie than staples, so the nutrition transition is broader than a simple meat-consumption story.&lt;/p&gt;
&lt;h3 id="q6-how-do-the-diet-restriction-and-food-trade-policy-counterfactuals-compare-on-welfare-and-effectiveness"&gt;Q6. How do the diet restriction and food trade policy counterfactuals compare on welfare and effectiveness?&lt;/h3&gt;
&lt;p&gt;Diet restriction (no-beef): global GHG emissions fall 20%, global welfare falls 0.6%. The welfare effect is highly concentrated—Argentina experiences −3–5% welfare loss, Uruguay approximately −4% in the no-beef scenario, because they are large meat producers and exporters. Inequality between rich (Q4) and poor (Q1) countries worsens by 1.0 percentage point. Diet restriction (vegetarian): global GHG emissions fall 30%, global welfare falls 2.8%. Inequality worsens by 6.0 percentage points, indicating developing countries bear more of the cost because a larger share of their income goes to food, and their income sources (agriculture) are more directly affected. Food trade policy (&amp;rsquo;eat local&amp;rsquo;, raising trade costs to cut transportation emissions by 75%): global GHG emissions fall 11.9%, but global welfare falls 17.8%—roughly 25–30 times the welfare cost per percentage point of emission reduction compared to dietary policies. Inequality worsens substantially more: Q4/Q1 ratio worsens by 4.9 percentage points. Countries in the bottom GDP quartile face welfare losses up to 41%. The paper concludes that dietary restrictions are both substantially more effective in reducing GHG emissions and far more equitable in their welfare consequences than food trade policies.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-share-of-agricultural-ghg-from-transportation-versus-production-and-what-are-the-implications"&gt;Q7. What is the share of agricultural GHG from transportation versus production, and what are the implications?&lt;/h3&gt;
&lt;p&gt;In the 2010 data, GHG emissions from food transportation account for approximately 5% of total agricultural GHG (0.7 Gt CO₂ out of approximately 17.2 Gt total). Production accounts for 95% (16.5 Gt CO₂). This has two implications. First, in the economic growth counterfactual, transportation emissions increase by 2.2%, but because transportation is only 5% of total, its contribution to total emission growth (0.004 Gt) is negligible. Second, it implies that policies targeting food &amp;lsquo;food miles&amp;rsquo; or local eating are poorly targeted: even a dramatic 75% reduction in transportation emissions only mechanically eliminates 4.6% of total agricultural GHG, and the actual general equilibrium reduction (11.9%) comes mostly from production effects (agricultural trade restrictions reduce global production and consumption), accompanied by very large welfare costs.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-and-validation-exercises-are-conducted"&gt;Q8. What robustness checks and validation exercises are conducted?&lt;/h3&gt;
&lt;p&gt;The paper provides several validation exercises. (1) The reduced-form income elasticity regressions are run both with all crops and excluding all meat products (beef, lamb and mutton, pig meat, poultry), yielding nearly identical coefficients of 0.176 and 0.175 (columns 1 and 2 of Table 1), and with country-year and product-year fixed effects (columns 3–4), showing similar results across specifications. (2) The structural income elasticities are compared to the reduced-form estimates, with a cross-method slope of 0.64–0.77 and R² of 0.93–0.95, reassuring given the two methods make different identifying assumptions. (3) Model fit is checked against six untargeted empirical regularities (Figure 6): declining agricultural employment share, rising input cost share, rising modern technology land share, rising food GHG per capita, rising calories per capita, and rising food GHG per calorie—all with GDP per capita. The model matches the sign and approximate magnitude of each relationship. (4) Household-level estimates using Brazil&amp;rsquo;s POF survey replicate the cross-country finding that higher-GHG products have higher income elasticities, controlling for fixed effects, food price proxies, and excluding meat. (5) The decomposition of the cross-sectional income-emissions gradient shows that equalizing comparative advantage (column 3) or trade costs (column 4) across countries leaves the gradient approximately unchanged, supporting the focus on preferences and technology.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-prior-work-and-where-does-it-depart-from-it"&gt;Q9. How does this paper relate to prior work and where does it depart from it?&lt;/h3&gt;
&lt;p&gt;The paper sits at the intersection of several literatures. It builds on Farrokhi and Pellegrina (2023) for the granular grid-cell production model with technology choice; on Costinot, Donaldson, and Smith (2016) for the agricultural field structure; and on Comin, Lashkari, and Mestieri (2021) for non-homothetic CES preferences and the identification of income elasticities. Key departures: (a) Relative to Comin et al. (2021), the authors extend identification to nested CES preferences and to an open-economy without requiring price data—their method is exact rather than approximate. (b) Relative to the environmental science literature (e.g., Hoolohan et al., 2013; Perignon et al., 2017; Tilman et al., 2011), the paper endogenizes general equilibrium supply responses, which the authors show dramatically attenuate the effect of both income growth and dietary policies on emissions. (c) Relative to prior quantitative spatial models of climate change (e.g., Shapiro 2016 on trade costs and CO₂), this paper focuses on agricultural emissions specifically and introduces nutrition transition and technology choice. (d) The authors claim to be the first to analyze both dietary restrictions and food trade policies on agricultural emissions within quantitative trade models. (e) Relative to Chen et al. (2022), who use a computable general equilibrium model with general equilibrium supply adjustments, this paper includes far more food products (47 vs. their smaller set) and endogenizes technology choice, both of which are quantitatively important for capturing the nutrition transition.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-papers-mechanism-for-why-vegetable-and-fruit-consumption-also-raises-ghg-emissions-as-income-rises-even-without-meat"&gt;Q10. What is the paper&amp;rsquo;s mechanism for why vegetable and fruit consumption also raises GHG emissions as income rises, even without meat?&lt;/h3&gt;
&lt;p&gt;The paper notes in footnote 1 that the positive correlation between income elasticities and GHG emissions per calorie persists even when meat products are excluded from the sample (Table 1, columns 3–4). The reason is that vegetables and fruits—which become more preferred as countries grow richer—emit more GHG per calorie than staple foods such as yams and potatoes. Staples require little processing or refrigeration and are typically produced with traditional, low-input technologies. By contrast, fresh fruits and vegetables (especially high-value items such as berries, asparagus, grapes, and coffee) require more energy-intensive transportation, storage, and sometimes greenhouse production. This means that the nutrition transition generates rising emissions not merely through the beef channel emphasized in much of the public debate, but through a broader shift away from calorie-dense staples toward diverse, lower-calorie-density products that happen to have higher GHG footprints per calorie.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-model-imply-about-the-environmental-kuznets-curve-for-food-emissions"&gt;Q11. What does the model imply about the Environmental Kuznets Curve for food emissions?&lt;/h3&gt;
&lt;p&gt;The paper explicitly tests for and finds no evidence of an Environmental Kuznets Curve (EKC) in food emissions—that is, no inverse-U shape in which emissions per capita eventually decline as countries become very rich, as might be expected if wealthy nations adopt more sustainable diets or stricter environmental regulations. The income-emission relationship is found to be approximately log-linear across all levels of development (footnote 8). This is consistent with the broader empirical literature on the EKC (cited survey by Dinda, 2004). The implication is that there is no automatic &amp;lsquo;greening&amp;rsquo; of diets as countries develop; active policy intervention would be needed.&lt;/p&gt;
&lt;h3 id="q12-how-is-economic-development-modeled-in-the-policy-counterfactuals-and-what-are-the-scope-conditions"&gt;Q12. How is economic development modeled in the policy counterfactuals, and what are the scope conditions?&lt;/h3&gt;
&lt;p&gt;Economic development is modeled as a uniform 10% increase in TFP for three types of agents: (i) modern agricultural producers, (ii) non-agricultural producers, and (iii) agricultural input producers (fertilizers, machinery, pesticides). Traditional agricultural technology is not subject to productivity growth, following Gollin, Parente, and Rogerson (2007). This creates both income effects (via higher wages) and substitution effects (via changes in relative input prices that favor modern, input-intensive technology). The scope conditions are important: the results apply specifically to a uniform global TFP shock, not to individual-country development. For individual-country TFP shocks, the analytical decomposition (equation 34) shows that general equilibrium income spillovers to foreign countries can attenuate the nutrition transition if foreign incomes fall (e.g., due to terms-of-trade effects). The model does not incorporate dynamics (it is a static model calibrated to 2010), so it cannot directly speak to transition paths or time horizons for emission convergence.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-welfare-implications-for-developing-countries-under-different-policies-and-why-do-dietary-policies-dominate"&gt;Q13. What are the welfare implications for developing countries under different policies, and why do dietary policies dominate?&lt;/h3&gt;
&lt;p&gt;Under economic growth (10% TFP shock), global welfare rises 14.9% with a modest increase in Q4/Q1 inequality of 0.4 percentage points, indicating relatively even welfare gains. Under no-beef, global welfare falls 0.6% but inequality worsens by 1.0 pp; under vegetarian, welfare falls 2.8% and inequality worsens by 6.0 pp—developing countries lose more because more of their income is spent on food and the agricultural sector is a larger share of their economy. Under eat-local (food trade restrictions), welfare falls 17.8% and the Q4/Q1 ratio worsens by 4.9 pp, with countries in the bottom GDP quartile facing losses up to 41%. The stark dominance of dietary policies over trade policies reflects two structural features: (a) food trade restrictions reduce the gains from comparative advantage in food production, which are particularly large for food-exporting developing countries; and (b) the welfare cost per unit of GHG reduction is far higher for trade policies because they distort production allocation without addressing the underlying demand-side emissions driver.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Nutrition Transition&lt;/strong&gt;: As defined and used in this paper: the demand-side process by which rising income causes consumers to shift their caloric intake away from staple foods (yams, potatoes, rice, millet) toward food products with higher GHG emissions per calorie (meats, fruits, vegetables, coffee). The transition is captured in the model by non-homothetic income elasticity parameters ε_k that are higher for more emissions-intensive products and is operative even after excluding all meat products.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agricultural Modernization&lt;/strong&gt;: As defined and used in this paper: the supply-side process by which rising wages induce producers to substitute from traditional, labor-intensive agricultural technology (τ=0, no purchased intermediate inputs) toward modern, input-intensive technology (τ=1, fertilizers, machinery, pesticides), which emits more GHG per calorie of output. This operates within each crop and is captured in the model by endogenous technology choice at the plot level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-Homothetic CES Preferences (Nested)&lt;/strong&gt;: A three-tier preference structure in which the expenditure share of a food product k depends on income through a product-specific parameter ε_k that governs how fast the product&amp;rsquo;s preference weight grows with utility. Products with higher ε_k have higher income elasticities; the overall income elasticity of the agricultural sector (0.39 in this paper&amp;rsquo;s calibration) is an expenditure-weighted average of the ε_k values. The nested structure allows the agricultural sector&amp;rsquo;s income elasticity relative to non-agriculture to be determined separately from the income elasticities of individual food products within agriculture.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implicit Marshallian Demand&lt;/strong&gt;: The demand equation derived from non-homothetic CES preferences by substituting out unobservable price indices using a base good, yielding a demand specification that depends on observable expenditure shares and income rather than on prices directly. In this paper&amp;rsquo;s open-economy extension, trade shares further substitute out unobservable variety price indices, making the estimation equation fully price-data-free.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GHG Emission Intensity (per calorie)&lt;/strong&gt;: In this paper: the parameter φ_k (crop-specific) and φ_τ (technology-specific), where φ_kτ = φ_k × φ_τ is the kg CO₂-equivalent emitted per 1,000 kcal of crop k produced under technology τ. This is the key cross-product heterogeneity that, combined with income elasticity heterogeneity, drives the environmental consequences of the nutrition transition. In the data: ranges from below 5 kg CO₂ per 1,000 kcal for wheat and rye to above 35 kg for beef and coffee.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Grid-Cell Production Model&lt;/strong&gt;: A representation of the agricultural supply side in which the Earth&amp;rsquo;s land surface is divided into approximately 1.1 million fields (FAO-GAEZ), each with agro-climatically determined potential yields by crop and technology that are independent of market conditions. Within each field, a continuum of plots is allocated to crops and technologies via Fréchet productivity draws, yielding smooth aggregate supply functions and allowing for realistic specialization patterns and technology gradients across geography.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Back-of-the-Envelope (Demand Mechanism) Benchmark&lt;/strong&gt;: In this paper: a partial-equilibrium counterfactual calculation that takes observed or baseline food demand quantities and simply attributes changes to them from a policy without allowing supply prices, production, or trade flows to adjust. The paper systematically compares model general equilibrium results against this benchmark (column 9 of Table 4) to quantify how much supply-side adjustments matter, finding that the back-of-the-envelope approach overstates the emission impact of economic growth by approximately three times, and overstates the emission reduction from dietary policies by roughly one-third.&lt;/p&gt;</description></item><item><title>Distortions, Producer Dynamics, and Aggregate Productivity: A General Equilibrium Analysis</title><link>https://macropaperwarehouse.com/papers/distortions-producer-dynamics-and-aggregate-productivity-a-general-equilibrium-analysis/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/distortions-producer-dynamics-and-aggregate-productivity-a-general-equilibrium-analysis/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks how institutional distortions to factor markets affect not only the static allocation of inputs across farms but also the dynamic choices — crop selection and productivity-enhancing investment — that determine the long-run distribution of farm productivities and hence aggregate agricultural TFP. The question matters because prior work on misallocation has largely treated the productivity distribution as exogenous; this paper endogenizes it, showing that the dynamic channels can be quantitatively larger than the static factor-misallocation channel.&lt;/p&gt;
&lt;p&gt;The empirical foundation is the Vietnam Access to Resources Household Survey (VARHS), a balanced panel of 2,118 farm households surveyed biennially from 2006 to 2016 across twelve provinces in north and south Vietnam. Vietnam provides a natural laboratory: post-1986 reforms decollectivized agriculture nationally, but deeply divergent pre-reform institutions (collective agriculture in the north for more than three decades; private household farming in the south throughout) produced durable differences in land-market functioning, crop-choice restrictions, and property-rights security. Measured TFP is more than 2.5 times higher in the south than the north (the observed log TFP ratio implies roughly a 2.5-fold level difference). The elasticity of land use with respect to farm TFP is 0.554 in the south versus 0.152 in the north, and the elasticity of labor use is 0.382 versus 0.122 — three to four times larger in the south — indicating far more efficient resource allocation in the south. The share of perennial-crop farmers (high-value cash crops, especially coffee) is 33% in the south and roughly 5% in the north. Average biennial TFP growth is 6.2% in the south versus 2.6% in the north.&lt;/p&gt;
&lt;p&gt;The authors build a dynamic general equilibrium model of heterogeneous farm managers (following Lucas 1978) in which farm productivity has four components: a permanent farmer-specific component, a random transitory component, an endogenous managerial ability component accumulated through investment, and a crop-specific component tied to endogenous crop choice. Institutional distortions are modeled as idiosyncratic revenue taxes correlated with farm productivity (following Restuccia and Rogerson 2008), with the key parameter being the elasticity of distortions with respect to farm productivity (rho). A higher rho means more-productive farms face proportionately larger distortions, which (i) compresses the gap between large and small farms in equilibrium factor use, and (ii) reduces the private return to investing in ability. The model also incorporates government-imposed crop restrictions that force a fraction of farms to grow rice regardless of profitability. The model is calibrated to south Vietnam moments: average TFP growth, dispersion in TFP and growth, the land-size distribution, the measured elasticity of distortions, and crop shares. Measurement error in output and inputs is explicitly modeled following Bils, Klenow, and Ruane (2021); the estimated BKR statistic is 0.906 for the south and 0.987 for the north, indicating relatively limited measurement error by manufacturing-sector standards.&lt;/p&gt;
&lt;p&gt;The main counterfactual imposes north Vietnam distortion parameters on the south-calibrated benchmark economy. Three distortion parameters differ: (1) the distortion elasticity rho rises from 0.79 (south) to 0.91 (north); (2) crop-specific distortions flip sign — in the south perennials face lower effective taxes than rice (phi_perennial = 1.61 &amp;gt; 1), while in the north perennials face higher effective taxes than rice (phi_perennial = 0.68 &amp;lt; 1); (3) the share of farms subject to government-imposed crop restrictions rises from 23% to 43%.&lt;/p&gt;
&lt;p&gt;The counterfactual experiment produces four main quantitative results. First, aggregate TFP falls by 41% relative to the benchmark, accounting for 61% of the observed productivity gap between north and south Vietnam (the observed ratio is 0.42; the counterfactual ratio is 0.59). Second, the average biennial farm TFP growth rate falls by 1.6 percentage points (from 6.23% to 4.60%), accounting for just under half of the observed 3.6 percentage-point north-south gap. Third, TFP dispersion (standard deviation of log TFP) falls by 8 percentage points, more than half of the 14-percentage-point lower dispersion observed in the north. Fourth, the share of perennial farmers collapses from 33% to 9%, closely matching the observed 5% in the north.&lt;/p&gt;
&lt;p&gt;Channel decomposition reveals that static factor misallocation alone reduces output by 19.4% (one-third of the total 40.8% gap, proportionately allocated), while the crop-choice channel reduces output by 8.0% and the farm-ability channel (endogenous investment) reduces output by 31.6%. Together, the dynamic channels (crop choice plus farm ability) account for approximately two-thirds of the total productivity loss, more than doubling the contribution of static misallocation. Among individual distortions, the distortion elasticity rho alone accounts for a 38.3% output reduction, crop-specific distortions account for 7.5%, and government crop restrictions account for only 1.4%. The key mechanism is that a small increase in rho (from 0.79 to 0.91) has large productivity consequences because the productivity cost is convex in rho and accelerates as rho approaches one — at rho = 1, distortions fully absorb all incremental profits from higher ability, eliminating investment incentives entirely.&lt;/p&gt;
&lt;p&gt;The paper shows that measurement error has limited impact on the north-south comparison (since the main experiment is a within-survey, within-country comparison), but substantially inflates the level gains from removing all distortions: removing measurement error from the model more than doubles the estimated gains from moving to a first-best economy, underscoring that measurement error matters most in cross-economy level comparisons.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-the-main-threats-to-validity"&gt;Q1. What is the identification strategy and the main threats to validity?&lt;/h3&gt;
&lt;p&gt;The identification exploits within-country, within-survey variation between north and south Vietnam, which share a common currency, survey instrument, and price measurement methodology. The main threat is that technology and geography differ across regions beyond institutions. The paper addresses this in two ways. First, it restricts comparisons to the two rice-growing delta regions — the Red River Delta (north) and Mekong Delta (south) — where technology and geographic differences are minimal, and shows the same patterns hold: measured distortion elasticity in the Mekong Delta is 0.79 versus 0.94 in the Red River Delta, and growth is higher and productivity more dispersed in the south. Second, the paper uses FAO Global Agro-Ecological Zones data to show land quality differences are negligible between north and south and, if anything, slightly favor the north; when scaled through the production function (land share times span-of-control = 0.35), land quality cannot account for the observed TFP gap. A second threat is measurement error inflating wedge dispersion and the estimated distortion elasticity. The paper addresses this by embedding explicit measurement error in the calibration and by using the Bils-Klenow-Ruane (2021) methodology, finding BKR statistics of 0.91 (south) and 0.99 (north), suggesting measurement error is modest in agriculture relative to manufacturing. The calibrated true distortion elasticity for the south is rho = 0.79, versus a measured elasticity of 0.86, a bias of around 0.06 — consistent with BKR estimates.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-productivity-channels-and-how-is-each-measured"&gt;Q2. What are the three productivity channels and how is each measured?&lt;/h3&gt;
&lt;p&gt;The three channels are (1) static factor misallocation, (2) crop distribution, and (3) farm ability. Each is isolated by a sequential decomposition: for factor misallocation, counterfactual distortions rho and phi are imposed while holding the crop and ability distributions fixed at benchmark-economy values, yielding an output loss of 19.4%. For crop distribution, the crop shares are adjusted to the counterfactual economy while holding within-crop ability distributions fixed at benchmark values; output falls by 8.0%. For farm ability, the ability distribution conditional on crop type is adjusted to the counterfactual while holding crop shares fixed; output falls by 31.6%. The sum (59.0%) exceeds the total gap (40.8%) because of negative interactions among channels — factor misallocation has a smaller bite when the productivity distribution is more compressed, as in the counterfactual.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-role-of-the-distortion-elasticity-parameter-rho-and-why-does-it-generate-outsized-productivity-losses-from-a-small-change"&gt;Q3. What is the role of the distortion elasticity parameter rho and why does it generate outsized productivity losses from a small change?&lt;/h3&gt;
&lt;p&gt;The parameter rho governs the extent to which more productive farms face proportionately larger distortions. At rho = 0, distortions are orthogonal to productivity; at rho = 1, distortions grow one-for-one with productivity, fully taxing away any incremental profit from increasing ability. The investment return to moving up the ability ladder is proportional to the incremental profit gained, which equals (1 - rho) times the increment in revenue. As rho rises toward 1, this return collapses toward zero. Because the South&amp;rsquo;s calibrated rho is already 0.79 — close to 1 on the relevant scale — a further increase to 0.91 is disproportionately large in terms of investment disincentives. The paper demonstrates this asymmetry explicitly in Appendix C.6: a symmetric increase and decrease of rho by 0.1 (set to the observed North-South difference in measured elasticity) reduces output by 42% when rho rises but only 39% when rho falls, driven primarily by the farm-ability channel (27 log points versus 21 log points difference in log output).&lt;/p&gt;
&lt;h3 id="q4-how-do-crop-specific-distortions-and-government-crop-restrictions-work-and-what-is-their-quantitative-contribution"&gt;Q4. How do crop-specific distortions and government crop restrictions work and what is their quantitative contribution?&lt;/h3&gt;
&lt;p&gt;Crop-specific distortions phi_i create wedges that differ across crop types. In the south, phi_perennial = 1.61 (perennial growers face lower effective taxes than rice farmers), while in the north phi_perennial = 0.68 (perennial growers face higher effective taxes). This reversal in relative distortions discourages perennial farming in the north both directly (lower profits) and dynamically (perennial farmers, who tend to be higher-ability, invest less). Unilaterally imposing north crop-specific distortions on the south benchmark reduces output by 7.5%. Government-imposed crop restrictions force a fraction omega of farms to grow rice regardless of profitability, with omega rising from 23% to 43% north-south. This channel has the smallest impact (1.4% output loss) because: (a) a large fraction of restricted farmers would have chosen rice anyway, and (b) back-of-envelope calculation shows the loss amounts to reducing productivity of only about 7% of farmers (the 20 percentage-point change in omega times the 33% perennial share) by about 20% (measured perennial-rice TFP gap).&lt;/p&gt;
&lt;h3 id="q5-what-empirical-evidence-motivates-the-endogenous-investment-mechanism"&gt;Q5. What empirical evidence motivates the endogenous investment mechanism?&lt;/h3&gt;
&lt;p&gt;Table 3 shows that in both north and south Vietnam, farm investment (cash or labor investment in irrigation or soil/water conservation) and extension-service participation are positively correlated with farm TFP and negatively correlated with farm-level distortion wedges, indicating that more distorted farms invest less. In the south, both investment and extension services are significantly positively associated with subsequent TFP growth. In the north, only extension-service participation is positively associated with future growth, while physical investment is not — suggesting the return to investment is suppressed in the north. The data also document a life-cycle profile (Figure 3) in which farm TFP rises steeply for young farms and then levels off, much more sharply in the south than in the north, consistent with faster ability accumulation in the less-distorted south.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-is-documented-across-crop-types-within-each-region"&gt;Q6. What heterogeneity is documented across crop types within each region?&lt;/h3&gt;
&lt;p&gt;In the south, perennial farmers have higher average output (log output 10.6 vs. 9.9 for rice), more land (3.9 acres vs. 2.4), more labor, higher TFP (above mean relative to rice), and far higher biennial TFP growth (10.9% vs. 4.9%). In the north, the pattern reverses: perennial farmers underperform relative to rice farmers in output (-0.583 log points, significant), land, labor, and TFP (-0.413 log points). This reversal occurs because crop-specific distortions disproportionately penalize perennial farming in the north. Despite the average gaps, there is substantial productivity overlap across crop types within both regions (Figure A.1), with many unproductive perennial farmers and productive rice farmers coexisting. This overlap motivates the paper&amp;rsquo;s modeling of crop selection as a utility-cost decision with idiosyncratic taste heterogeneity (Frechet distribution), rather than a pure productivity-cutoff rule.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-conducted-and-what-do-they-show"&gt;Q7. What robustness checks are conducted and what do they show?&lt;/h3&gt;
&lt;p&gt;Four robustness exercises are conducted. First, re-calibrating with fixed quadratic investment-cost curvature (zeta = 2) instead of the estimated 1.74 yields a counterfactual output ratio of 58.8%, similar to the baseline 59.2%. Second, lowering the targeted average growth rate by 2 percentage points (addressing the concern that aggregate TFP growth partly reflects economy-wide technology rather than ability investment) produces a counterfactual output ratio of 58.6% — essentially unchanged. Third, lowering the targeted growth rate by 4 percentage points produces 62.4%, still economically large. Fourth, two model extensions are explored: (a) incorporating a hump-shaped life-cycle profile with a young-to-old transition produces a 43% productivity loss, similar to the 41% baseline; (b) allowing entrants to draw ability from a distribution dependent on the exiting predecessor&amp;rsquo;s ability produces a 57% productivity loss — larger than baseline because investment creates positive spillovers to future entrants.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-the-prior-misallocation-literature"&gt;Q8. How does this paper relate to and differ from the prior misallocation literature?&lt;/h3&gt;
&lt;p&gt;The paper builds on Restuccia and Rogerson (2008) and Hsieh and Klenow (2009), who model static misallocation via idiosyncratic wedges. It contributes three extensions. First, it endogenizes the farm productivity distribution by adding investment and crop choice, so that the same wedges that generate static misallocation also distort dynamics — this doubles the productivity cost. Second, the experiment is a within-country comparison between two regions rather than a comparison against a hypothetical undistorted economy, avoiding the criticism that the undistorted benchmark is unrealistic. The re-calibrated north model accounts for 100% of the observed north-south TFP ratio (40.7% model vs. 42% data). Third, the dynamic model generates falsifiable predictions about farm TFP growth rates, TFP dispersion, and crop distributions — all of which move in the right directions — providing a richer validation test than static models allow. The paper also relates to Hsieh and Klenow (2014), who document faster life-cycle productivity growth in less distorted economies (India and Mexico vs. US), and to Adamopoulos and Restuccia (2020), who study land reform in Vietnam but with exogenous productivity distributions; the current paper finds that endogenizing productivity distributions significantly amplifies the costs of distortions. The measurement-error treatment follows Bils, Klenow, and Ruane (2021) and Adamopoulos et al. (2022).&lt;/p&gt;
&lt;h3 id="q9-what-does-the-model-imply-about-a-hypothetical-undistorted-economy"&gt;Q9. What does the model imply about a hypothetical undistorted economy?&lt;/h3&gt;
&lt;p&gt;Removing all distortions (rho = 0, phi_i = 1 for all crops, omega = 0, sigma_epsilon = 0) increases TFP by a factor of 3.37 relative to the south benchmark (Appendix C.5, Table C.11), meaning the first-best economy is more than three times as productive. Static reallocation gains alone (holding the productivity distribution fixed) account for roughly 70% of this gap. The remaining gains come from the endogenous shift in the ability distribution — in the undistorted economy, lower ability farmers invest less (because higher general equilibrium wages lower profits) but higher ability farmers invest more (because distortions no longer claw back incremental profits). The net result is a more polarized ability distribution with a heavier right tail, consistent with the concentrated structure of agriculture in advanced economies. Importantly, the paper cautions that abstracting from measurement error inflates the estimated undistorted-economy gains by more than a factor of two: a model without measurement error yields gains more than twice as large as the calibrated model that accounts for it.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central policy implication is that institutions distorting factor markets — particularly those that generate a positive correlation between farm productivity and the effective tax rate (captured by rho) — reduce agricultural TFP through three compounding channels, with two-thirds of the loss arising from dynamic distortions (investment suppression and crop selection) rather than static factor reallocation. This means that standard static calculations of misallocation costs substantially understate the true costs. Land accumulation restrictions that prevent productive farms from expanding (the historical legacy in north Vietnam, where 82.8% of Red River Delta agricultural land was state-allocated) are particularly costly because they are the empirical analog of high rho. The scope conditions are: (1) the analysis applies to the Vietnamese agricultural context in 2006-2016, a period well after initial reform but still characterized by persistent institutional differences; (2) the model abstracts from occupational choice and structural transformation, which other work has shown amplify distortion costs further; (3) the main results are robust to the north-south within-country design but level estimates (gains from the first-best) are sensitive to measurement error treatment. The paper suggests that reducing the productivity-distortion correlation — e.g., through secure land titles and functioning land rental markets — would unlock gains exceeding what static misallocation calculations imply.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Distortion elasticity (rho)&lt;/strong&gt;: The parameter governing how strongly institutional distortions — modeled as idiosyncratic revenue taxes — are correlated with farm-level productivity. A higher rho means more productive farms face proportionately larger distortions, compressing both static factor allocation and the dynamic return to investing in ability. In the paper&amp;rsquo;s calibration, rho = 0.79 for south Vietnam and 0.91 for north Vietnam; the difference accounts for the majority of the measured North-South productivity gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Managerial ability ladder&lt;/strong&gt;: The endogenous component of farm productivity that farmers accumulate through investment. A farmer at ability node h has productivity phi^h; investing e units of output raises ability to the next node with probability x = (e/a)^(1/zeta). The investment return depends on the incremental profit gain from higher ability, which is suppressed when the distortion elasticity rho is large, creating a tight link between static institutional distortions and dynamic farm growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crop-specific distortion (phi_i)&lt;/strong&gt;: A factor in the distortion specification that captures institutional barriers differentially affecting specific crops. In south Vietnam, phi_perennial = 1.61, meaning perennial-crop growers face lower effective taxes than rice farmers; in north Vietnam, phi_perennial = 0.68, reversing the ranking. This parameter embeds market-access barriers, infrastructure gaps, and regulatory disadvantages specific to particular crops.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Government-imposed crop restriction (omega)&lt;/strong&gt;: The share of farms legally required to grow rice regardless of relative profitability or household preferences, reflecting Vietnamese national food-security policies. The restriction is more prevalent in the north (43% of farms) than the south (23%). Unlike idiosyncratic distortions, crop restrictions enter the model as a direct constraint on the discrete crop-choice decision rather than as a tax on revenue, and the paper finds their productivity cost is relatively small (1.4% output loss) because many restricted farmers would have chosen rice anyway.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic misallocation&lt;/strong&gt;: The productivity losses arising from distortions&amp;rsquo; effects on farms&amp;rsquo; forward-looking decisions — specifically the choice of crop (crop selection) and investment in managerial ability — as opposed to the static misallocation of given factor inputs across farms with fixed productivities. In the paper, dynamic misallocation accounts for two-thirds of the total productivity gap, more than doubling the contribution of static factor misallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;BKR measurement-error statistic&lt;/strong&gt;: A diagnostic from Bils, Klenow, and Ruane (2021) that estimates the ratio of true wedge dispersion to observed wedge dispersion using the cross-term in a regression of log output changes on log wedge, log input, and their interaction. Values near one indicate little measurement error; values near zero indicate the observed wedge is mostly noise. The paper finds BKR = 0.906 for south Vietnam and 0.987 for north Vietnam, indicating measurement error is modest and is unlikely to confound the north-south comparison.&lt;/p&gt;</description></item><item><title>Distributional Consequences of Becoming Climate-Neutral</title><link>https://macropaperwarehouse.com/papers/distributional-consequences-of-becoming-climate-neutral/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/distributional-consequences-of-becoming-climate-neutral/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates how the EU&amp;rsquo;s Fit-for-55 climate package will affect aggregate output and distribute its costs across the income distribution. The question matters because energy is a necessity good — poorer households devote a larger share of spending to energy — so policies that raise energy prices are regressive in their first-order incidence. Despite a large literature on the aggregate macroeconomics of the green transition, distributional consequences have received limited attention.&lt;/p&gt;
&lt;p&gt;The authors build a parsimonious dynamic general-equilibrium model with two infinitely-lived households (rich and poor), a standard output-producing firm that treats energy as a complementary CES input alongside the capital-labor aggregate, and an energy-producing sector that combines a carbon-intensive brown technology with a carbon-free green technology as imperfect substitutes (CES with elasticity of substitution calibrated to 3 following Papageorgiou et al. 2017). The novel feature is Price Independent Generalized Linearity (PIGL) non-homothetic preferences following Boppart (2014), which generate nonlinear Engel curves: the poor agent&amp;rsquo;s energy expenditure share exceeds the rich agent&amp;rsquo;s, matching Eurostat Household Finance and Consumption Survey data (2015) showing the bottom income quintile has more than twice the energy expenditure share of the top quintile. The model targets an 18% energy expenditure share for the poor agent and 7.5% for the rich agent. The rich agent holds all financial wealth; the poor agent lives on labor income alone. The government taxes the brown technology and recycles revenue as a green-technology subsidy under a balanced budget, representing the ETS. Agents have perfect foresight. The paper simulates perfect-foresight transitions from an initial steady state to a new climate-neutral steady state, with the transition path endogenously determining the new steady state — a nonstandard feature arising from non-homothetic preferences.&lt;/p&gt;
&lt;p&gt;In the baseline scenario (linear tax ramp over 25 years), achieving an 85% reduction in brown energy use requires a 168% tax on the brown technology. This drives the price of energy services up by 49%, GDP down by 9.3% in the new steady state, energy as a production input down by 10.9%, and capital input down by 9.3%, while the real wage falls by roughly 7% and the real interest rate is nearly unchanged (dropping by only 0.02 percentage points transiently). The welfare cost measured in expenditure-equivalent terms is a 10.8% loss for the rich agent and a 16.2% loss for the poor agent — the poor agent suffers approximately 50% more. To finance consumption during the transition the poor agent accumulates debt equal to 38.8% of annual income.&lt;/p&gt;
&lt;p&gt;Results are highly sensitive to the brown-green substitution elasticity: raising it from 3 to 5 roughly halves the required tax (to 78.6%) and halves GDP losses (to 4.7%); lowering it to 2 roughly doubles the tax (to 354%) and GDP losses (to 17.7%). Non-homothetic preferences matter quantitatively: switching to homothetic preferences (while preserving different expenditure shares) shrinks aggregate GDP losses by 26% and eliminates nearly all distributional disparity, confirming that the non-homotheticity — not merely different expenditure levels — is the operative distributional mechanism. If the Fit-for-55 energy efficiency improvement target of 1.49% per year is simultaneously achieved, the required tax falls to 136%, the price of energy actually declines by 5.5%, and GDP rises by 1.1% in the new steady state, with the poor agent benefiting slightly more and accumulating assets (4% of annual income) rather than debt.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-modeling-and-calibration-strategy-and-what-are-the-main-threats"&gt;Q1. What is the core modeling and calibration strategy, and what are the main threats?&lt;/h3&gt;
&lt;p&gt;The paper is a quantitative theory exercise with no econometric identification. Calibration targets HFCS Eurostat data (2015) for energy expenditure shares by income quintile, the Papageorgiou et al. (2017) estimate of the brown-green substitution elasticity (ρE = 3), and stylized facts on wealth and income distribution from Krueger, Mitman, and Perri (2016). The main threat is parameter uncertainty around ρE, which the paper acknowledges is poorly identified empirically and which drives the results almost one-for-one. The sensitivity analysis explores ρE ∈ {2, 3, 5}, a range the paper concedes is narrow relative to the literature&amp;rsquo;s full dispersion.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-generating-the-distributional-gap-between-rich-and-poor"&gt;Q2. What are the main mechanisms generating the distributional gap between rich and poor?&lt;/h3&gt;
&lt;p&gt;Three reinforcing channels: (1) Non-homothetic preferences give the poor agent a higher energy expenditure share (18% vs. 7.5%), so the 49% energy price increase hits the poor&amp;rsquo;s budget much harder as a share of income. (2) The poor agent cannot buffer the shock through wealth drawdowns (holding zero net assets initially), forcing it to accumulate debt of 38.8% of annual income. (3) Non-homothetic preferences alter the labor supply response: as expenditures fall, the poor agent&amp;rsquo;s labor supply declines less than the rich agent&amp;rsquo;s (the rich agent decreases labor supply by 0.2 percentage points more), reflecting that leisure is a luxury good in this preference system. In the new steady state the rich agent&amp;rsquo;s consumption of the consumption good drops sharply while the rich agent front-loads consumption at the announcement, immediately jumping 2% higher.&lt;/p&gt;
&lt;h3 id="q3-how-are-non-homothetic-preferences-distinguished-empirically-and-in-the-model-from-simply-having-different-expenditure-shares"&gt;Q3. How are non-homothetic preferences distinguished empirically and in the model from simply having different expenditure shares?&lt;/h3&gt;
&lt;p&gt;Section 4.4 runs a counterfactual with homothetic preferences (ε = 0) but preserves identical initial expenditure shares for each agent (7.5% and 18%) by making ν agent-specific. Under homotheticity the expenditure shares do not vary with income as the transition unfolds. The comparison shows that GDP losses shrink by 26% (from 9.3% to 6.9%) and the distributional gap nearly vanishes — both agents experience almost identical welfare losses. This decomposition isolates the effect of non-homotheticity itself: it is the income-dependent adjustment of expenditure shares during the transition, not merely the different initial levels, that drives both larger aggregate losses and the distributional disparity.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented-and-along-what-dimensions"&gt;Q4. What heterogeneity is documented and along what dimensions?&lt;/h3&gt;
&lt;p&gt;Heterogeneity is modeled along two dimensions: initial wealth (rich holds all assets; poor holds zero) and energy expenditure shares (18% for poor, 7.5% for rich) arising from non-homothetic preferences. The model produces no within-group heterogeneity by construction (two-agent framework). The paper documents the time paths of consumption, expenditures, expenditure equivalents, energy expenditure shares, and wealth shares for each agent separately along the transition, showing that both agents cut energy consumption by roughly 15% while the poor agent cuts consumption-good spending by substantially more than the rich agent.&lt;/p&gt;
&lt;h3 id="q5-what-alternative-transition-timing-paths-are-explored-and-what-do-they-imply"&gt;Q5. What alternative transition timing paths are explored and what do they imply?&lt;/h3&gt;
&lt;p&gt;Three alternatives supplement the linear baseline: tax introduction after 1 year, after 12.5 years, and after 25 years of the announcement. Key findings: (a) the required final tax rate is nearly insensitive to timing — the 25-year-delayed scenario requires 172% vs. 168% in the baseline; (b) conditional on excluding climate damages, it is always welfare-superior to delay implementation, with the poor agent gaining close to 3.5 percentage points in expenditure equivalent welfare by delaying to 25 years vs. implementing after 1 year; (c) gradual vs. immediate introduction yields similar welfare outcomes in the benchmark without adjustment costs, but with investment adjustment costs (χ = 10) a sudden implementation causes a brief sharp drop in the real interest rate without large quantity effects.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-gdp-measure-differ-from-aggregate-output-in-the-model"&gt;Q6. How does the GDP measure differ from aggregate output in the model?&lt;/h3&gt;
&lt;p&gt;GDP is defined to exclude the share of final output used as input into energy production. Aggregate output Y falls 7.3% in the new steady state, but GDP falls 9.3%. The gap (approximately 2 percentage points) reflects the increased resource cost of energy production under the green transition: because the brown and green technologies are imperfect substitutes, satisfying the emission reduction target requires devoting a larger share of final output to producing energy services, a real resource drain captured in the GDP definition but excluded from raw output Y.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-energy-efficiency-scenario-imply-and-what-is-its-key-caveat"&gt;Q7. What does the energy efficiency scenario imply, and what is its key caveat?&lt;/h3&gt;
&lt;p&gt;If energy efficiency improves at 1.49% per year over 25 years (a 45% cumulative gain in energy-producing-firm total factor productivity), the required tax falls to 136.3%, the price of energy declines by 5.5% (rather than rising 49%), and GDP rises 1.1% rather than falling 9.3%. The poor agent benefits more from the efficiency gains and accumulates assets worth 4% of annual income rather than debt. The critical caveat is that the efficiency improvement is modeled as purely exogenous and costless. The paper explicitly acknowledges that achieving these efficiency gains may require investment that is not modeled, so the results should be interpreted as an upper bound on the offsetting potential.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-relate-to-and-differ-from-the-most-closely-related-prior-work"&gt;Q8. How does the paper relate to and differ from the most closely related prior work?&lt;/h3&gt;
&lt;p&gt;Ascari et al. (2025) is the closest related paper (developed independently). Differences: (i) Ascari et al. use a Bewley-type incomplete-markets model generating heterogeneity through random discount factors, whereas this paper uses a two-agent complete-markets construct with exogenously fixed initial wealth; (ii) this paper allows endogenous labor supply, which increases short-run flexibility; (iii) this paper does not consider transfer schemes to redistribute away from distributional consequences. Results are described as broadly consistent. Fried, Novan, and Peterman (2018) and Boehl and Budianto (2024) use OLG models and find inequality implications but focus on inter-generational rather than intra-generational distributional effects.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The core implications are: (1) the Fit-for-55 emission tax alone is regressive — the poor bear a welfare loss 50% larger than the rich and end up with 38.8% of annual income in additional debt; (2) delaying tax implementation (with early announcement) is welfare-improving in the absence of climate damage modeling — the welfare difference is nearly 3.5 percentage points for the poor between fastest and latest implementation; (3) if energy efficiency targets are met exogenously, the transition is nearly costless and distributional concerns vanish; (4) the regressive result is conditional on the government recycling tax revenues to green-technology subsidies rather than to household transfers. All these implications are conditional on European economies where climate damages are plausibly small and the model abstracts from open-economy dynamics, endogenous technology, and within-income-group heterogeneity.&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-are-reported"&gt;Q10. What robustness checks are reported?&lt;/h3&gt;
&lt;p&gt;Five robustness exercises are reported: (1) investment adjustment costs raised from χ = 0 to χ = 10 — minimal effect on welfare or quantities in the smooth baseline, though sudden tax introduction produces a brief interest-rate plunge; (2) homothetic preferences counterfactual while maintaining initial expenditure shares (Section 4.4); (3) elasticity of substitution between brown and green technology at ρE = 2 and ρE = 5 (Section 4.3, Table 2); (4) alternative transition timing (1 year, 12.5 years, 25 years post-announcement; Section 4.2); (5) simultaneous energy efficiency improvement of 1.49% per year (Section 4.5). A New Keynesian extension with Rotemberg price adjustment costs and a Taylor rule (Appendix B) is also provided for robustness on inflation dynamics.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-main-caveats-or-limitations-acknowledged-by-the-authors"&gt;Q11. What are the main caveats or limitations acknowledged by the authors?&lt;/h3&gt;
&lt;p&gt;Climate damages are excluded, so the paper understates the case for early action and cannot provide a full welfare comparison between acting early and acting late. Energy efficiency improvement is modeled as exogenous and costless, overstating the net gain from that channel. The two-agent framework abstracts from within-group heterogeneity and overlapping generations. Open-economy dynamics are not modeled; the brown-technology structure serves as a reduced-form for energy imports but does not capture international price feedback. The elasticity of substitution between brown and green technology is uncertain, and results are nearly proportional to this parameter. The model has no endogenous innovation or directed technical change, limiting applicability to long-run transition analysis.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Non-homothetic PIGL preferences&lt;/strong&gt;: Preferences of the Price Independent Generalized Linearity class (Boppart 2014) where energy expenditure shares depend on income level, making energy a necessity good (share declining in income) and consumption goods a luxury. Parameter ε ∈ (0,1) controls non-homotheticity; ε = 0 recovers homothetic preferences. The paper calibrates γ = 0.639 from CEX data, implying an elasticity of substitution between consumption and energy goods of approximately 0.4.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Brown vs. green technology&lt;/strong&gt;: Two imperfectly substitutable technologies for producing energy services within the model&amp;rsquo;s energy sector. The brown technology converts units of final output into energy services using a carbon-intensive (emission-producing) process; the green technology is emission-free. They enter a CES aggregator for energy production with elasticity ρE calibrated to 3. Imperfect substitutability means the green transition raises the cost of energy services even with subsidies to green technology.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Expenditure equivalent loss&lt;/strong&gt;: The welfare metric used in the paper: the percentage change in expenditures in the initial steady state (without any tax) that would make an agent indifferent between remaining in the initial steady state and living through the actual transition path. Defined implicitly by equating flow utility at scaled initial expenditures to flow utility along the transition. Baseline results: -10.8% for the rich agent and -16.2% for the poor agent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tax on the brown technology&lt;/strong&gt;: The policy instrument modeled as capturing the essence of EU ETS and national carbon schemes. It raises the unit cost of the emission-intensive energy input; revenue is recycled as a subsidy to the green technology within a balanced government budget rather than distributed to households. A 168% tax achieves the 85% emission reduction target in the baseline, implying fossil fuel prices nearly triple.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous final steady state&lt;/strong&gt;: The model&amp;rsquo;s new steady state after the green transition is not predetermined; it depends on the wealth distribution that emerges endogenously during the transition. Because markets are complete and preferences are non-homothetic, different transition paths generate different terminal wealth distributions and therefore different aggregate outcomes in the new steady state. This prevents backward solution and requires a fully nonlinear transition path solver.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Energy expenditure share by income quintile&lt;/strong&gt;: The empirical regularity, documented from Eurostat HFCS data (2015), that the bottom income quintile devotes more than twice the fraction of disposable income to energy (electricity, gas, fuels for personal transport) as the top quintile. This fact calibrates the non-homotheticity of preferences (targeting 18% for the poor agent and 7.5% for the rich agent) and motivates the paper&amp;rsquo;s focus on distributional consequences.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Elasticity of substitution between brown and green technology (ρE)&lt;/strong&gt;: The key production-side parameter governing how easily the energy sector can switch from fossil-fuel to clean inputs. Calibrated to ρE = 3 from Papageorgiou et al. (2017). Results are nearly proportional to this parameter: ρE = 5 halves and ρE = 2 roughly doubles the required tax, GDP losses, and welfare costs. The paper identifies this as the dominant source of quantitative uncertainty.&lt;/p&gt;</description></item><item><title>Entrepreneurial Investment Dynamics and the Wealth Distribution</title><link>https://macropaperwarehouse.com/papers/entrepreneurial-investment-dynamics-and-the-wealth-distribution/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/entrepreneurial-investment-dynamics-and-the-wealth-distribution/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates how the illiquidity of entrepreneurial capital shapes investment dynamics and wealth inequality. The central question is whether entrepreneurship drives wealth heterogeneity or merely attracts the already-wealthy — and, specifically, whether the investment behavior of nascent entrepreneurs can be rationalized by frictions on capital reallocation rather than financial constraints alone.&lt;/p&gt;
&lt;p&gt;The empirical foundation is the restricted Kauffman Firm Survey (KFS), a single-cohort panel of 3,140 U.S. firms founded in 2004 and tracked through 2011. The key measurement is the log average revenue product of capital (log ARPK), residualized on two-digit NAICS industry fixed effects and time dummies. Two striking facts emerge. First, the cross-sectional distribution of log ARPK is left-skewed (skewness approximately -0.33, mean -0.49, standard deviation 1.75, kurtosis 5.7). Second, the distribution shows asymmetric persistence: the autocorrelation of log ARPK in the bottom quintile (ρ₁ = 0.897) is statistically significantly larger than in the top quintile (ρ₅ = 0.443), and the diagonal entry of the estimated transition matrix for the first quintile (0.614) substantially exceeds that for the fifth (0.568). These facts are inconsistent with standard models: a frictionless dynamic investment model with time-to-build predicts i.i.d. ARPK; one with collateral constraints predicts right-skewness and right-tail persistence.&lt;/p&gt;
&lt;p&gt;The model extends Cagetti and De Nardi (2006) by distinguishing between liquid bonds and illiquid entrepreneurial capital. Capital adjustment generates four friction types: a proportional fixed cost (fs) on upward investment, a proportional transaction cost (λ) on downsizing, an additional proportional cost (ζ) on exit, and a minimum capital requirement on entry. The model is calibrated via indirect inference to identifying moments from the KFS (persistence and skewness of log ARPK, investment rate distribution, share of employer firms, entry and exit rates) plus economy-wide targets (entrepreneur fraction, interest rate of 3–4%).&lt;/p&gt;
&lt;p&gt;The FULL-sample calibration yields λ = 0.43 (43% loss on capital sold by continuing entrepreneurs) and ζ = 0.55 (additional 55% write-down upon exit), with a proportional fixed cost fs = 0.035 (3.5%). The effective net collateral constraint is approximately 44% of the real capital value. These frictions are quantitatively large: eliminating them under general equilibrium raises aggregate TFP in the entrepreneurial sector by 23.3% and average welfare by 23.1% in consumption equivalent variation terms. Decomposing the welfare losses relative to a complete-markets benchmark shows that approximately 89% of the total welfare loss (relative to full frictions) is attributable to market incompleteness and financial frictions, with the remaining 11% directly attributable to the illiquidity frictions — that is, frictions alone account for roughly 7.15 percentage points of a total 64.8% lifetime consumption welfare loss.&lt;/p&gt;
&lt;p&gt;A key finding on wealth inequality contradicts prior literature. When calibrated to KFS micro-data, the model generates a Gini coefficient of 0.65 (FULL sample) or 0.53 (NAICS54), well below the empirical U.S. Gini of approximately 0.8. The top 1% hold only 26% of wealth in the FULL calibration versus roughly 30% empirically. This contrasts with Quadrini (2000) and Cagetti and De Nardi (2006), who match the wealth distribution by calibrating to PSID or SCF household survey data. The reason for the gap is the left-skewed, illiquidity-depressed returns to entrepreneurship in the KFS: the calibrated returns to scale (ν = 0.79 FULL, 0.82 NAICS54) and the transaction costs together suppress the variance of capital income returns. Removing illiquidity frictions raises the Gini from 0.65 to 0.77 (fixed-r partial equilibrium) or 0.72 (general equilibrium), demonstrating that capital illiquidity compresses the wealth distribution by depressing average entrepreneurial returns.&lt;/p&gt;
&lt;p&gt;Three policy experiments — credit expansion (reducing borrowing spreads à la SBA 7(a) programs), a government buyer-of-last-resort for used capital (Resale I), and exit-cost reduction (Fire sale) — all raise welfare by 0.07–0.15% in consumption equivalent terms and TFP by 0.5–0.9% relative to benchmark. Resale policies are preferred by entrepreneurs; workers prefer the credit policy. All three policies benefit lower-wealth households more than wealthy ones (the richest decile suffers welfare losses due to the savings tax used to finance the programs). The paper concludes that policies addressing capital illiquidity can yield welfare gains comparable to or exceeding standard credit provision programs, and that the distinction between illiquidity risk and financial constraint risk has first-order importance for policy design.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-two-core-empirical-facts-from-the-kfs-that-motivate-the-paper-and-why-do-standard-models-fail-to-generate-them"&gt;Q1. What are the two core empirical facts from the KFS that motivate the paper, and why do standard models fail to generate them?&lt;/h3&gt;
&lt;p&gt;First, the cross-sectional distribution of log ARPK among KFS firms is left-skewed (skewness ≈ -0.33), not symmetric or right-skewed. Second, log ARPK shows higher persistence in the left tail (autocorrelation ρ₁ = 0.897 for bottom-quintile firms) than in the right tail (ρ₅ = 0.443). A frictionless dynamic model with time-to-build predicts i.i.d. log ARPK that inherits the distribution of TFP innovations, generating no skewness under Gaussian shocks and no persistence. Models with collateral constraints (as in Cagetti and De Nardi 2006) generate right-skewed ARPK with right-tail persistence, because constrained firms operate below optimal scale, pushing ARPK above the unconstrained optimum. Neither class of models can produce the left-skewed, left-tail-persistent pattern in the KFS.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-mechanism-by-which-partial-irreversibility-generates-left-skewness-and-left-tail-persistence"&gt;Q2. What is the mechanism by which partial irreversibility generates left-skewness and left-tail persistence?&lt;/h3&gt;
&lt;p&gt;Partial irreversibility creates an asymmetry between the purchase price and the resale price of capital (the resale price being 1 − λ per unit). When a bad productivity shock hits, the option value of waiting to recover is higher than the cost of holding excess capital, so entrepreneurs adopt a &amp;lsquo;wait-and-see&amp;rsquo; attitude and maintain oversized firms rather than downsizing immediately. This creates a left tail of low-ARPK, large-capital firms. Moreover, since the incentive to wait is itself persistent (the transitory bad shock must resolve before the entrepreneur will downsize), the left tail displays higher autocorrelation. The exit cost ζ amplifies this for the exit margin: entrepreneurs with poor draws stay in business longer than is efficient, further extending the left tail. The right tail is not symmetrically elongated because entrepreneurs seeking to expand face a different option value (the call option value of capital rises), leading them to invest to smaller sizes, slightly thickening the right tail — but not enough to overcome the left-tail extension.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-calibration-strategy-and-which-parameters-are-identified-by-which-moments"&gt;Q3. What is the calibration strategy, and which parameters are identified by which moments?&lt;/h3&gt;
&lt;p&gt;Eleven parameters are jointly calibrated to KFS moments via indirect inference. The key mappings are: the downsizing transaction cost λ is identified by the asymmetric left-tail persistence of log ARPK (the ratio ρ₁/ρ₅ increases monotonically in λ); the exit cost ζ is identified by the skewness of log ARPK (higher ζ monotonically increases left skewness); the collateral constraint ϕ also affects skewness but has no monotone effect on ρ₁/ρ₅, aiding separation; the returns to scale ν is identified by the coefficient from a log-revenue on log-capital regression for employer firms; the fixed investment cost fs is identified by the fraction reporting positive investment; TFP shock autocorrelation ρ_z is identified by investment rate autocorrelation; the shock standard deviation σ_z by the coefficient of variation of investment rates; and the worker signal distortion and entrepreneur signal distortion parameters control entry and exit rates respectively. The discount factor β pins down the interest rate. Two separate calibrations are run: one targeting full KFS sample moments (FULL) and one targeting the modal industry — Professional, Scientific and Technical Services (NAICS54, 24.7% of the sample) — as a robustness check.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-main-calibrated-parameter-values-and-how-do-they-compare-across-the-full-and-naics54-calibrations"&gt;Q4. What are the main calibrated parameter values and how do they compare across the FULL and NAICS54 calibrations?&lt;/h3&gt;
&lt;p&gt;For the FULL calibration: λ = 0.43, ζ = 0.55, ϕ = 0.92, fs = 0.035, ρ_z = 0.66, σ_z = 0.43, ν = 0.79, β = 0.9265, α_e = 0.63. For NAICS54: λ = 0.53, ζ = 0.75, ϕ = 0.035, fs = 0.23, ρ_z = 0.66, σ_z = 0.43, ν = 0.82, β = 0.94, α_e = 0.50. The illiquidity parameters (λ and ζ) are larger in NAICS54 than in FULL. The collateral constraint parameter ϕ differs substantially (0.92 FULL versus 0.035 NAICS54), though the net effective collateral constraint (accounting for λ and depreciation) converges to a similar range in both calibrations.&lt;/p&gt;
&lt;h3 id="q5-how-are-the-illiquidity-and-financial-friction-channels-distinguished-both-theoretically-and-empirically"&gt;Q5. How are the illiquidity and financial friction channels distinguished both theoretically and empirically?&lt;/h3&gt;
&lt;p&gt;Theoretically, collateral constraints (parameterized by ϕ) make the lower support of log ARPK truncated from the left (log ARPK ≥ log(r+δ) - log α), generating right-skewness and right-tail persistence. Illiquidity frictions (λ and ζ), by contrast, induce a wait-and-see option value that extends the left tail of ARPK while leaving the right tail relatively thinner, generating left-skewness and left-tail persistence. Empirically, the paper proposes using the sign and magnitude of the skewness of log ARPK (negative implies illiquidity dominates; positive implies financial frictions dominate) and the ratio of left-tail to right-tail persistence (ρ₁/ρ₅ &amp;gt; 1 indicates illiquidity frictions, &amp;lt; 1 indicates financial frictions) as discriminating statistics. Separately, the portfolio composition of entrepreneurs offers a further discriminating test: increasing illiquidity drives entrepreneurs to hold more liquid assets (flight to liquidity), while tightening collateral constraints pushes entrepreneurs toward more illiquid assets in their portfolios.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-aggregate-tfp-and-welfare-findings-from-the-counterfactual-analysis"&gt;Q6. What are the aggregate TFP and welfare findings from the counterfactual analysis?&lt;/h3&gt;
&lt;p&gt;Under general equilibrium, removing all illiquidity frictions (λ = ζ = fs = 0) raises entrepreneurial sector TFP by 23.3% and average economy-wide welfare by 23.1% in consumption equivalent variation. Under partial equilibrium (fixed interest rate), welfare gains are even larger: 24.8% (entrepreneur subgroup) and 58.3% (worker subgroup), for an economy-wide average of 16.6%. The GE result is somewhat lower because the interest rate adjusts when more capital flows into entrepreneurship. The average productivity of entrepreneurs (conditional on being an entrepreneur) is 8.8% higher in the no-friction world than in the benchmark. The TFP gains arise from both extensive-margin selection (higher-productivity entrepreneurs enter; lower-productivity ones exit) and intensive-margin reallocation (high-productivity firms operate closer to optimal scale; low-productivity firms downsize rather than persist).&lt;/p&gt;
&lt;h3 id="q7-how-does-the-paper-decompose-total-welfare-losses-between-market-incompleteness-and-the-illiquidity-distortions"&gt;Q7. How does the paper decompose total welfare losses between market incompleteness and the illiquidity distortions?&lt;/h3&gt;
&lt;p&gt;Following Buera and Shin (2011), the paper computes welfare as a fraction of lifetime consumption relative to a complete-markets benchmark (a social planner&amp;rsquo;s problem where the planner allocates occupational choice and capital optimally). Relative to complete markets, the economy with no illiquidity frictions but with market incompleteness loses approximately 57.7% of lifetime consumption. The benchmark economy (with all frictions) loses approximately 64.8% of lifetime consumption relative to complete markets. The difference — approximately 7.15 percentage points — is attributed to the illiquidity frictions. As a share of the total frictional loss, about 89% is attributable to market incompleteness and financial frictions, and 11% to the illiquidity frictions. While 11% may seem small as a fraction, in absolute terms it is economically non-trivial.&lt;/p&gt;
&lt;h3 id="q8-why-does-the-paper-find-that-entrepreneurship-cannot-match-the-empirical-wealth-distribution-when-calibrated-to-the-kfs"&gt;Q8. Why does the paper find that entrepreneurship cannot match the empirical wealth distribution when calibrated to the KFS?&lt;/h3&gt;
&lt;p&gt;The model generates a Gini of 0.65 (FULL) or 0.53 (NAICS54) against a U.S. empirical Gini of approximately 0.8. The top 1% holds roughly 26% of wealth in the FULL calibration versus around 30% empirically. Two factors suppress capital income risk in the KFS-calibrated model. First, the calibrated returns to scale (ν = 0.79 FULL, 0.82 NAICS54) are lower than those used by Cagetti and De Nardi (2006) (ν ≈ 0.88), which were calibrated to PSID/SCF data on large-ish successful firms. Lower ν translates exponentially into lower variance of capital income. Second, the illiquidity frictions directly depress average returns to entrepreneurship by raising the user cost of capital and forcing entrepreneurs into suboptimal firm sizes. These two forces together prevent the model from generating the thick right tail of wealth needed to match empirical distributions. The paper argues that the KFS captures &amp;lsquo;broad&amp;rsquo; small-scale entrepreneurship, not the high-growth, high-return entrepreneurs who likely account for the top of the wealth distribution.&lt;/p&gt;
&lt;h3 id="q9-how-does-capital-illiquidity-affect-the-wealth-distribution-conditional-on-holding-returns-to-scale-fixed"&gt;Q9. How does capital illiquidity affect the wealth distribution conditional on holding returns to scale fixed?&lt;/h3&gt;
&lt;p&gt;More illiquid capital (higher λ or ζ) compresses the wealth distribution and lowers the Gini coefficient. The Gini rises from 0.65 (benchmark FULL calibration) to 0.77 under partial equilibrium without illiquidity frictions, and to 0.72 under general equilibrium without illiquidity frictions (while holding the net collateral constraint constant). The NAICS54 benchmark Gini is 0.53, rising to 0.76 (PE) or 0.68 (GE) without illiquidity frictions. The mechanism is that illiquid capital depresses the average return to entrepreneurial wealth, which compresses the income process and reduces the variance of wealth accumulation. Additionally, illiquid capital forces entrepreneurs to hold more bonds as a liquidity buffer, reducing the overall scale of their business investment and thus their lifetime income.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-three-policy-experiments-and-their-comparative-findings"&gt;Q10. What are the three policy experiments and their comparative findings?&lt;/h3&gt;
&lt;p&gt;The three policies are all financed by a proportional tax on bond savings returns. (1) Credit expansion: the government subsidizes borrowing intermediation costs (analogous to SBA 7(a)/CDC 504 programs), reducing the spread between the saving and borrowing rate. Economy-wide welfare rises by about 0.147%; TFP rises by about 0.9% relative to benchmark. Workers benefit more (0.169%) than entrepreneurs (-0.006% average for all entrepreneurs, since most wealthy entrepreneurs do not borrow and pay the tax). (2) Resale policy I (Buyer of last resort for all used capital): government offers a higher resale price q ≥ 1 − λ. Economy-wide welfare rises about 0.076%; TFP rises 0.6%. Entrepreneurs gain (0.084%) while workers also gain (0.074%) indirectly through the option value of future entrepreneurship. (3) Fire-sale (exit cost reduction only, Resale II): government subsidizes exiting entrepreneurs&amp;rsquo; capital resale. Economy-wide welfare rises 0.073%; TFP rises 0.5%. Workers prefer credit; entrepreneurs prefer resale policies. Wealthiest decile suffers welfare losses under all three policies. All welfare numbers are in consumption equivalent variation.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-relate-to-cagetti-and-de-nardi-2006-and-where-does-it-diverge"&gt;Q11. How does the paper relate to Cagetti and De Nardi (2006) and where does it diverge?&lt;/h3&gt;
&lt;p&gt;The paper builds directly on the Cagetti and De Nardi (2006) framework of occupational choice and incomplete markets with collateral constraints, extending it by separating liquid bonds from illiquid physical capital. In Cagetti and De Nardi (2006), bonds and capital are perfect substitutes; the sole friction is a collateral constraint that limits investment. The paper shows that this one-asset framework generates right-skewed ARPK and right-tail persistence — inconsistent with KFS facts. The paper&amp;rsquo;s two-asset framework with partial irreversibility generates left-skewed ARPK and left-tail persistence. Furthermore, Cagetti and De Nardi (2006) calibrate to PSID/SCF income data and successfully match the wealth distribution; the paper shows this success partly reflects the higher returns to scale implied by those data. When calibrated directly to KFS firm-level data, the model substantially undershoots the empirical wealth inequality, because the KFS captures a representative sample of small-scale entrepreneurs with genuinely lower returns to scale and significant illiquidity frictions.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-role-of-the-options-value-effect-and-the-collateral-constraint-channel-in-the-model-and-how-do-they-differ"&gt;Q12. What is the role of the options value effect and the collateral constraint channel in the model, and how do they differ?&lt;/h3&gt;
&lt;p&gt;The options value effect is described as the primary distortion. When capital is illiquid (λ or ζ &amp;gt; 0), the put option value of capital falls (selling capital is costly), raising the threshold signal required for workers to enter entrepreneurship, and raising the threshold signal required for incumbents to exit. As a result, entry rates fall, exit rates fall, potential entrepreneurs delay entry, and poorly performing entrepreneurs overstay. Along the intensive margin, the asymmetric purchase/resale price leads entrepreneurs planning to downsize to wait (operating larger-than-optimal firms) and entrepreneurs planning to invest to be more cautious (operating smaller-than-optimal firms). The collateral constraint channel is a secondary effect: illiquid capital reduces the net resale value that can serve as collateral (effective constraint = (1-λ)(1-δ)(ϕ)k&amp;rsquo;), tightening the borrowing constraint even when the formal collateral parameter ϕ is moderate. Crucially, while tighter ϕ forces entrepreneurs to hold more illiquid capital (no flight to liquidity), higher λ forces entrepreneurs to hold more liquid assets (flight to liquidity) — a key empirical distinction.&lt;/p&gt;
&lt;h3 id="q13-what-robustness-exercises-does-the-paper-conduct"&gt;Q13. What robustness exercises does the paper conduct?&lt;/h3&gt;
&lt;p&gt;The paper runs two separate full calibrations: one to the entire KFS sample (FULL) and one to the modal industry NAICS54 (Professional, Scientific and Technical Services, 24.7% of the sample). Both calibrations are used to assess the wealth distribution findings. The paper also examines moments at the two-digit industry level (only one industry shows statistically significant results due to small sample size, though most show economically significant signs). An additional measurement error parameter is explored in the appendix, where capital is assumed to be observed with multiplicative log-normal error; this helps improve model fit to the data. All policy experiments are computed under both partial equilibrium (fixed interest rate) and general equilibrium. The paper also analytically proves (in the appendix) the ARPK distribution properties for the four benchmark frameworks (frictionless, time-to-build only, static collateral constraints, and dynamic collateral constraints), establishing the theoretical necessity of partial irreversibility for the facts.&lt;/p&gt;
&lt;h3 id="q14-what-heterogeneity-in-welfare-effects-is-documented-across-the-wealth-distribution"&gt;Q14. What heterogeneity in welfare effects is documented across the wealth distribution?&lt;/h3&gt;
&lt;p&gt;Under all three policy experiments, welfare gains decrease with wealth. The poorest households gain the most in consumption equivalent variation terms because they receive a disproportionate share of the program&amp;rsquo;s benefits (better borrowing conditions, higher resale prices, improved option value of entrepreneurship) while paying a smaller absolute share of the savings tax used to finance the programs. The top 10% richest households — who are the primary taxpayers — experience welfare losses under all three policies. This pattern holds across credit, resale, and fire-sale policies, though the magnitude varies. Separately, entrepreneurs (who are wealthier on average, with over 50% concentrated in the top wealth decile) mostly lose from the credit policy (they fund it but don&amp;rsquo;t directly borrow) while gaining from resale policies (they benefit from higher capital resale prices regardless of wealth position). Workers (who are generally poorer) overwhelmingly gain from credit policies since the option value of switching to entrepreneurship rises substantially.&lt;/p&gt;
&lt;h3 id="q15-what-does-the-paper-imply-for-interpreting-the-literature-on-financial-constraints-and-entrepreneurship"&gt;Q15. What does the paper imply for interpreting the literature on financial constraints and entrepreneurship?&lt;/h3&gt;
&lt;p&gt;The paper issues several cautionary findings. First, the implied formal collateral parameter is relatively loose (ϕ = 0.92), consistent with Hurst and Lusardi (2004), Nanda (2011), and Robb and Robinson (2014) — who find no evidence that average entrepreneurs face severe financial constraints. However, once illiquidity is accounted for, the effective (net) collateral constraint is only about 44% of real capital value, consistent with Evans and Jovanovic (1989) and Cagetti and De Nardi (2006). This suggests that what appears empirically as &amp;lsquo;financial constraint&amp;rsquo; is partly a manifestation of capital illiquidity: banks lend less against entrepreneurial capital because its resale value is low, not primarily because of limited commitment. Second, empirical studies using regional variation in financial conditions to identify financial constraint effects may suffer from omitted variable bias, since resale prices of capital are also highly correlated with local financial conditions. Third, aggregate statistics such as startup rates and investment levels cannot distinguish between illiquidity shocks and financial constraint shocks; portfolio composition (the ratio of liquid to illiquid assets) is a more informative diagnostic.&lt;/p&gt;
&lt;h3 id="q16-what-is-the-papers-contribution-to-the-misallocation-literature-relative-to-hsieh-and-klenow-2009-asker-et-al-2014-and-midrigan-and-xu-2014"&gt;Q16. What is the paper&amp;rsquo;s contribution to the misallocation literature relative to Hsieh and Klenow (2009), Asker et al. (2014), and Midrigan and Xu (2014)?&lt;/h3&gt;
&lt;p&gt;Hsieh and Klenow (2009) and Asker et al. (2014) focus on the dispersion of log MRPK as a measure of misallocation, where adjustment costs (similar to fs and λ here) can generate observed dispersion without implying inefficiency. Midrigan and Xu (2014) focus on financial constraints (similar to ϕ) as the source of misallocation. The paper argues that these frameworks produce observationally equivalent outcomes in terms of log MRPK dispersion alone, making it impossible to distinguish between the two. The paper&amp;rsquo;s contribution is to show that the skewness of log ARPK and the asymmetric tail persistence are additional moments that can discriminate between the two types of frictions: negative skewness and left-tail dominance point to illiquidity frictions, while positive skewness and right-tail dominance point to financial frictions. This provides a new empirical diagnostic tool for decomposing sources of capital misallocation.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Average Revenue Product of Capital (ARPK)&lt;/strong&gt;: In the paper&amp;rsquo;s usage, ARPK = Y_it / K_{i,t-1}, the ratio of a firm&amp;rsquo;s real revenue to its beginning-of-period real capital stock, used as the primary measure of capital productivity. Log ARPK is residualized on two-digit NAICS industry fixed effects and time dummies before analysis, removing industry-level heterogeneity in capital shares and aggregate shocks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Partial irreversibility&lt;/strong&gt;: The friction arising from an asymmetry between the purchase price of new capital (normalized to 1) and the resale price of used capital (1 − λ for downsizing incumbents, and (1 − ζ)(1 − λ) for exiting entrepreneurs). This is modeled as a proportional transaction cost on capital sales and is interpreted as the difficulty of recouping original investment, analogous to a low resale value of used entrepreneurial equipment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wait-and-see attitude&lt;/strong&gt;: The behavioral response of entrepreneurs facing downside productivity shocks when capital is illiquid: rather than immediately downsizing or exiting upon a bad shock, they maintain larger-than-optimal firm sizes while waiting for conditions to improve. This is optimal because the transaction cost of selling capital makes the option of waiting (and possibly recovering) more valuable than the cost of operating an oversized firm.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Net collateral constraint (effective collateral parameter)&lt;/strong&gt;: Denoted ϕ̃ = (1 − λ)(1 − δ)ϕ, this is the fraction of entrepreneurial capital&amp;rsquo;s real value that can actually be pledged as collateral, after accounting for the reduced resale value from illiquidity (1 − λ) and physical depreciation (1 − δ). The paper distinguishes this from the formal limited-commitment parameter ϕ to show that observed financial constraints partly reflect capital illiquidity rather than contracting failures.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Options value effect&lt;/strong&gt;: The mechanism through which capital illiquidity distorts both the entry/exit decision and the intensive margin of investment. For downsizing incumbents, the put option value of capital (the option to sell it) falls when the resale price is low, inducing them to delay disinvestment. For potential entrants, the call option value of capital (the upside of entering) falls because losses upon exit are larger, raising the productivity signal threshold for entry. This is described as the primary distortion channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Span-of-control parameter (returns to scale, ν)&lt;/strong&gt;: The parameter ν ∈ (0,1) in the entrepreneurial production function y = z(k^{α_e} l^{1-α_e})^ν, capturing the extent to which managerial talent becomes diluted as firm size increases. The paper identifies ν = 0.79 (FULL) from the coefficient of a log-revenue on log-capital regression for employer firms, and shows that ν is the dominant determinant of the variance of capital income returns and hence the model&amp;rsquo;s ability to generate wealth inequality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption equivalent variation (CEV)&lt;/strong&gt;: The welfare metric used throughout the paper. For each household i, CEV µ_i is defined as the percentage increase in reference-economy consumption (or lifetime consumption stream) that makes the household indifferent between the reference economy and the economy of interest. Positive CEV means the new economy is preferred. Aggregate welfare is the distribution-weighted average of individual CEVs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Asymmetric persistence&lt;/strong&gt;: The empirical fact, documented in the KFS, that log ARPK shows higher autocorrelation at the bottom quintile (ρ₁ = 0.897) than at the top quintile (ρ₅ = 0.443), confirmed by both a conditional autocorrelation regression and a quintile transition matrix. This asymmetry is a key moment used to identify and distinguish illiquidity frictions (which produce left-tail persistence) from collateral constraints (which produce right-tail persistence).&lt;/p&gt;</description></item><item><title>Entry decision, the option to delay entry, and business cycles</title><link>https://macropaperwarehouse.com/papers/entry-decision-the-option-to-delay-entry-and-business-cycles/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/entry-decision-the-option-to-delay-entry-and-business-cycles/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; US cohorts of establishments born in recessions persistently employ fewer workers at entry and over their life cycle, yet are on average more productive than expansionary cohorts; the number of entrants is procyclical and roughly four times as volatile as aggregate employment. Standard firm-dynamics models cannot reproduce this strong, persistent selection of entrants without generating excessive variation in aggregate variables, because the expected lifetime value of entry is relatively insensitive to aggregate shocks of reasonable magnitude. The paper asks what makes initial aggregate conditions matter so much for the selection of entrants, and answers: potential entrants&amp;rsquo; ability to delay entry, a margin missing from existing frameworks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model setup.&lt;/strong&gt; The author builds a discrete-time, infinite-horizon firm-dynamics model with endogenous entry and exit, building on Moreira (2015) in the style of Hopenhayn (1992). The only aggregate shock is an exogenous AR(1) aggregate demand shock z. Heterogeneous incumbents differ in idiosyncratic productivity s (AR(1)) and customer capital b (accumulated from past sales, depreciating at rate δ), operate under monopolistic competition, draw a random fixed operating cost each period, and may exit endogenously or via a random exit shock γ. A constant mass of potential entrants holds heterogeneous signals q about post-entry productivity, drawn from a time-invariant Pareto distribution W(q). The key deviation: entrants may keep their signal and delay, observing a new z next period (probability τ of retaining the signal; τ=0 nests the standard model, τ=1 is the baseline). This creates a non-negative option value of delay V^w(q,z) that rises with q and with z.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings (with magnitudes).&lt;/strong&gt; The option to delay generates a countercyclical opportunity cost of entry: for reasonable parameters, entrants postpone until the present value of entry is up to twice the fixed entry cost. The threshold signal is countercyclical, so recessionary cohorts are fewer but more productive. Expected delay duration ranges from zero to six periods (years), negatively correlated with q. Calibrated to BDS establishment data 1977-2015 (a period is a year), with ρz=0.57, σz=0.0022, and τ=1 (an alternative identification gives τ=0.965, with nearly identical dynamics). The mechanism raises the variance of the number of entrants, for a given shock process, by about seven times. Recessionary (expansionary) cohorts employ 5.7% fewer (5.0% more) workers than the average cohort, persisting beyond 15 years; shutting down delay (τ=0) collapses this to ~1%, so ~80% of cohort-employment variation comes from delayers. Average recessionary productivity is ~3% higher under τ=1 vs only 0.4% under τ=0. The full model explains more than three-fourths of the persistence and variance of aggregate employment (model autocorrelation 0.57 vs data 0.61; std 0.012 vs 0.015). Empirically, cohort-level employment differences are driven by the composition (high-productivity/high-growth share), not the number, of entrants; the persistent customer-capital process plays a minor role (&amp;lt;7%).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implications.&lt;/strong&gt; Validating against the Great Recession: cohorts entering 2008-2016 account for ~45% of the depth (of an 8.9% drop in 2012) and ~85% of the slow recovery by 2016 in the data; the model reproduces ~39% of the 2012 depth and ~75% by 2016, with most of it coming from the entry margin. A standard model without delay, calibrated to the same facts, requires σz ~7x larger, yields aggregate-employment variance 1.7x the data, and predicts a Great-Recession employment drop twice as large as observed. Matching aggregate employment instead requires aggregate-demand-shock autocorrelation 1.40x and variance 25x higher. Ignoring the option to delay therefore yields misleading predictions about entrants&amp;rsquo; responses to permanent, temporary, and anticipated (news) policy shocks.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-mechanism-that-amplifies-the-effect-of-initial-aggregate-conditions-on-entrant-selection"&gt;Q1. What is the core mechanism that amplifies the effect of initial aggregate conditions on entrant selection?&lt;/h3&gt;
&lt;p&gt;The option to delay entry. Because entering today and entering tomorrow are mutually exclusive, waiting carries a non-negative option value V^w(q,z) that rises with the signal q and with aggregate demand z. With this intertemporal choice, a firm enters only if its gross value of entry exceeds the &lt;em&gt;total&lt;/em&gt; opportunity cost = fixed entry cost ce + option value of delay. This total cost is countercyclical (up to twice ce in recessions), so the threshold signal q*(z) becomes much more elastic to z. Even a small change in the relative benefit of entering today vs tomorrow shifts selection substantially, whereas without delay (τ=0) entry follows a neoclassical rule — enter if net lifetime benefits are non-negative — and the threshold barely moves with z.&lt;/p&gt;
&lt;h3 id="q2-why-does-a-firm-ever-find-it-optimal-to-delay-given-it-forgoes-period-profits"&gt;Q2. Why does a firm ever find it optimal to delay, given it forgoes period profits?&lt;/h3&gt;
&lt;p&gt;The decision hinges on the net value of waiting, V^w(q,z) − (V^gross(q,z) − ce). The aggregate demand level at entry affects not only first-period profits but also the expected post-entry survival rate (1−γ)G(c*_f), which is procyclical: in recessions the expected long-run value is lower, raising the risk of premature post-entry failure. This procyclical &amp;lsquo;discount factor&amp;rsquo; makes entry during expansions more valuable. Medium-productivity firms wait until the expected survival rate is high enough to compensate for low early-life demand. The author stresses that without irreversible and endogenous exit, the benefits of waiting would always be negative — endogenous exit risk is essential to the mechanism.&lt;/p&gt;
&lt;h3 id="q3-who-delays-and-who-does-not"&gt;Q3. Who delays, and who does not?&lt;/h3&gt;
&lt;p&gt;Delay has no effect on high- and low-productivity potential entrants; only medium-range-signal firms (q in [q*&lt;em&gt;{τ=0}(z), q*&lt;/em&gt;{τ=1}(z)]) find it profitable to wait for better aggregate demand. The lower the aggregate demand, the wider this range. At the business-cycle peak, nobody delays, so selection coincides with and without the option.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-empirical-identification-strategy-and-its-main-threat"&gt;Q4. What is the empirical identification strategy and its main threat?&lt;/h3&gt;
&lt;p&gt;Using the Business Formation Statistics (BFS), based on IRS EIN/SS-4 applications matched to BDS new employer businesses, the author separates applications that form a business within the first four quarters (First 4Q) from the second four quarters (Second 4Q), 2004Q3-2016Q4. The &amp;lsquo;wait-and-see&amp;rsquo; channel is identified from the share of late start-ups = Second4Q/(First4Q+Second8Q), which is significantly countercyclical (Fact 2). The main confound (Fact 3&amp;rsquo;s threat): bad aggregate conditions could lengthen the &lt;em&gt;time required to build&lt;/em&gt; a business (e.g., harder credit access in recessions) rather than reflecting deliberate waiting. The author controls for this using the average duration of business formation within the first four quarters and the total number of formations within eight quarters; the countercyclical share of late start-ups survives (Table 2, coefficient -0.304*** on HP real-GDP cycle). A separate caveat: the author cannot evaluate the &lt;em&gt;economic&lt;/em&gt; magnitude of the channel from data, because entrants who delay AND delay applying for EINs, or who apply but never return, are unobserved — hence the quantitative role is assessed via the structural model.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-testable-implication-that-distinguishes-the-mechanism-and-is-it-borne-out-in-data"&gt;Q5. What is the testable implication that distinguishes the mechanism, and is it borne out in data?&lt;/h3&gt;
&lt;p&gt;The model predicts that recessionary cohorts have, on average, HIGHER long-run survival rates than expansionary cohorts (countercyclical survival), because firms wait until expected survival is high enough. Without the option (τ=0) the model produces acyclical survival rates. In BDS data 1979-2015, cohort survival rates at ages g=1..5 are persistently negatively correlated with aggregate conditions at entry (e.g., for S3, corr with HP real-GDP cycle = -0.38, p=0.02; corr with Ihp = -0.46, p=0.00), robust across HP, linear-trend, unemployment, and NBER indicators, and across firm- vs establishment-level units. Note two counteracting forces: low demand directly lowers survival (higher failure) but raises it via selection; the net countercyclicality supports the selection channel.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-model-calibrated"&gt;Q6. How is the model calibrated?&lt;/h3&gt;
&lt;p&gt;17 parameters; a period = a year, unit = establishment. β=0.96 (4% riskless rate). Demand/customer-capital/productivity parameters from Foster et al. (2008, 2016): ρs=0.814, price elasticity ρ=1.622, demand-to-customer-capital elasticity η=0.919, depreciation δ=0.188. Entrant-distribution, selection, survival, size, and growth parameters (q, ξ, ce, μf, σf, γ, b0, σ_s, σ_e, α) jointly matched to BDS cohort moments (average entry rate ~12.1%, entrant employment share, size and survival to 30 years, employment share to age 5). The aggregate demand process (ρz=0.57, σz=0.0022) is calibrated to the autocorrelation (0.25) and std (0.06) of the HP-filtered (smoothing 100) entry rate. τ set to 1; an alternative strategy using the aggregate-employment time series identifies τ=0.965.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-paper-decompose-the-source-of-persistent-cohort-employment-differences"&gt;Q7. How does the paper decompose the source of persistent cohort-employment differences?&lt;/h3&gt;
&lt;p&gt;Counterfactuals (Table 6) hold the variation in the &lt;em&gt;number&lt;/em&gt; of entrants fixed while varying composition. &amp;lsquo;Adjust lowest s&amp;rsquo; (number variation from low-productivity firms) yields small, transient cohort-employment effects; &amp;lsquo;adjust highest s&amp;rsquo; yields large, persistent effects. The baseline lies between them: medium-productivity firms that delay amplify the procyclical variation in &lt;em&gt;high-productivity&lt;/em&gt; entrants, raising persistence. This matches Decker et al. (2014) and Pugsley-Sedlacek-Sterk: a small share of high-growth firms drives cohort contributions, and ex-ante entrant types explain most post-entry performance. The &amp;lsquo;only selection&amp;rsquo; counterfactual (shutting demand effects on post-entry firms) shows the customer-capital process contributes less than 7% to cohort-employment persistence.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-impulse-response-analysis-illustrate-propagation"&gt;Q8. How does the impulse-response analysis illustrate propagation?&lt;/h3&gt;
&lt;p&gt;A one-time negative demand shock sized to cut entrants by 25% (the Great-Recession magnitude): the baseline economy takes 3 years to recover half the employment decline and another 12 years to recover an additional 25%. An economy where the shock does not affect the entry margin recovers three-fourths of the decline in only 2 years, even when the shock is enlarged to match the baseline&amp;rsquo;s initial employment drop. Persistent entry-margin shocks accumulate, substantially deepening and prolonging the downturn (Table 9).&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;With the option to delay, entrant responses depend on the &lt;em&gt;relative&lt;/em&gt; benefit of entering today vs tomorrow, so policy effects vary with type, magnitude, timing, and duration. (1) A temporary cut in fixed entry cost raises the number of entrants more than a permanent cut during recessions, with equal effect in expansions; marginal entrants are high-productivity firms in recessions, low-productivity in expansions. Without the option, the response is invariant to policy duration. (2) News of a future entry-cost cut (after T periods) weakly &lt;em&gt;raises&lt;/em&gt; the threshold signal in all states — i.e., reduces entry today — and for small T this indirect, entry-deterring effect can dominate the eventual entry boost; standard models would only transmit such news through general-equilibrium channels. Scope: results derive from a partial-equilibrium reduced form; the author argues (Appendix A.3) that in general equilibrium the option value stays non-negative, so the entry threshold is weakly higher than in models without persistent signals, though procyclical wages partly offset the procyclical-discount-factor force.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-paper-relate-to-and-differ-from-prior-work"&gt;Q10. How does the paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;It addresses the Samaniego (2008) result that entry/exit are insensitive to reasonable productivity shocks and the Lee-Mukoyama (2018) &amp;lsquo;puzzle&amp;rsquo; of generating strong entrant selection. Rather than imposing cyclical entry costs (Lee-Mukoyama 2018), an entry function (Sedlacek-Sterk 2019), or exogenous entry-specific shocks (Clementi-Palazzo 2016; Sedlacek-Sterk 2017), it derives amplified selection endogenously from the option to delay. It complements &amp;lsquo;missing generation&amp;rsquo; (Gourio-Messer-Siemer) and demand-side (Sedlacek-Sterk; Moreira) explanations of procyclical cohort employment, extends the real-options literature (Bernanke 1993; Dixit-Pindyck 1994; Pindyck 2009; Bloom 2009) to the entry margin, and reinforces Sedlacek-Sterk&amp;rsquo;s finding that entry-stage selection, not post-entry choices, drives cohort contributions to aggregate fluctuations.&lt;/p&gt;
&lt;h3 id="q11-what-extensions-and-robustness-checks-are-provided"&gt;Q11. What extensions and robustness checks are provided?&lt;/h3&gt;
&lt;p&gt;(1) A two-stage entry phase (Appendix A.1) micro-founds the constant mass of potential entrants by adding an &amp;lsquo;aspiring start-up&amp;rsquo; free-entry stage, calibrated so only ~13% of aspiring start-ups (cq=0.022) become actual entrants, reconciling the low BFS application-to-employer-business transition rate (~14% over two years). (2) Allowing accumulation of delayed potential entrants (Appendix A.2) &lt;em&gt;amplifies&lt;/em&gt; cyclical differences across cohorts and increases procyclical entry-rate variation. (3) A general-equilibrium version (Appendix A.3) shows the model performs at least as well as standard models. Empirical results are robust to alternative cycle definitions (HP, linear trend, unemployment deviations, NBER), to firm- vs establishment-level units, to annual vs quarterly BFS data, and to ten-year pre-crisis cohort averages in the Great-Recession exercise.&lt;/p&gt;
&lt;h3 id="q12-what-caveats-does-the-author-flag"&gt;Q12. What caveats does the author flag?&lt;/h3&gt;
&lt;p&gt;The model generates a countercyclical average entrant size (consistent with Lee-Mukoyama 2015 for manufacturing plants) but at odds with Sedlacek-Sterk&amp;rsquo;s finding of procyclical entrant size in BDS; the author conjectures that allowing procyclical initial customer capital would only widen cyclical cohort-employment differences. The economic magnitude of the wait-and-see channel cannot be measured directly because key delaying groups are unobserved in BFS. Other Great-Recession forces (credit crunch, structural change in entrants) are not modeled and could also explain the 2008-2016 cohort employment drop. Explaining whether delayed entrants actually return to the market is left for future research.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Option value of delay (V^w(q,z))&lt;/strong&gt;: The present value a potential entrant forgoes by entering today instead of retaining its productivity signal and entering in a future period. It is non-negative everywhere, weakly increases in the signal q and in aggregate demand z, and exists only because exit is irreversible and endogenous (otherwise waiting would never pay).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Countercyclical opportunity cost of entry&lt;/strong&gt;: The total cost of entering — fixed entry cost ce plus the option value of delay — which rises in recessions (up to twice ce). It endogenously raises the elasticity of entry to aggregate demand and creates a group of firms that stay out despite positive expected net profits.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;Threshold signal q&lt;/em&gt;_τ(z)&lt;/em&gt;*: The minimum productivity signal at which a potential entrant chooses to enter at aggregate state z. It is countercyclical; under τ=1 it equals the signal at which gross entry value equals the total opportunity cost, and it is far more elastic to z than the τ=0 (no-delay) threshold.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Signal q and probability of recalling the signal τ&lt;/strong&gt;: q is a potential entrant&amp;rsquo;s heterogeneous, time-invariant signal about its initial post-entry productivity (drawn from Pareto W(q)). τ is the probability a delaying entrant keeps that signal next period; τ=0 collapses the model to a standard framework, τ=1 is the baseline (calibrated; identified value τ=0.965).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Customer capital (b)&lt;/strong&gt;: A demand-side stock tied to a firm&amp;rsquo;s past sales, depreciating at rate δ, that shifts demand for its differentiated good. Because it accumulates from prior sales, it slows firms&amp;rsquo; demand adjustment and creates persistence in production and employment, distinct from productivity differences (per Foster et al. 2016).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wait-and-see channel&lt;/strong&gt;: The empirical counterpart of the option-to-delay mechanism: a bad aggregate state at entry induces some potential entrants to postpone forming a business, raising the (countercyclical) share of late start-ups in BFS data, distinct from recessions merely lengthening the time required to build a business.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Recessionary vs expansionary cohorts&lt;/strong&gt;: Cohorts of establishments that begin operating when aggregate demand is below (z&amp;lt;1) vs above (z&amp;gt;1) the stochastic steady state. Recessionary cohorts are fewer, more productive, higher-survival, and persistently smaller in employment.&lt;/p&gt;</description></item><item><title>Environmental Subsidies to Mitigate Net-Zero Transition Costs</title><link>https://macropaperwarehouse.com/papers/environmental-subsidies-to-mitigate-net-zero-transition-costs/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/environmental-subsidies-to-mitigate-net-zero-transition-costs/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether public subsidies to green-technology producers, financed by a carbon tax, can materially reduce the macroeconomic cost of reaching net-zero CO2 emissions by 2060. The motivation is a market-structure failure that standard environmental models ignore: the abatement goods sector is initially immature and highly concentrated, with 10 percent of firms capturing roughly 80 percent of operating revenue (Eurostat/Ecorys data). Under such conditions a carbon tax alone raises the cost of abatement inputs, depresses competition, and generates a deep and prolonged GDP recession — even if it achieves the emissions target. The paper shows that redirecting carbon tax revenues toward subsidizing this sector can substantially offset the recession.&lt;/p&gt;
&lt;p&gt;The analytical vehicle is an environmental dynamic stochastic general equilibrium (E-DSGE) model for the world economy, built by merging three bodies of work: the DICE climate block (Nordhaus 1992, 2018), a real-business-cycle production structure in the spirit of Smets and Wouters (2007), and an endogenous market-structure framework for the abatement goods sector following Bilbiie, Ghironi, and Melitz (2012). Firm entry into the abatement sector responds to expected future profits, which depend on sunk costs. Two margins of adjustment are distinguished: the intensive margin (existing firms expanding production) and the extensive margin (startups creating new varieties). Competition in the abatement sector is a central object of analysis: higher firm numbers reduce the abatement price, which in turn lowers the carbon tax burden on final-goods producers.&lt;/p&gt;
&lt;p&gt;The model is estimated using Bayesian methods on five annual world time series from 1961 to 2019: real GDP growth, real consumption growth, CO2 emissions growth, the change in surface temperature anomaly, and the growth rate of environment-related patents (OECD). Because the model has stochastic growth trends, the authors use the extended-path solution method (Fair and Taylor 1983) rather than standard linearization, and an inversion filter to form the likelihood function. Posterior draws from 320,000 MCMC iterations (8 parallel chains, ~30 percent acceptance) pin down five structural parameters and ten shock parameters. Estimated initial output growth is approximately 4.99 percent per year and the initial emissions-to-output decoupling rate is 1.13 percent per year, both consistent with Nordhaus (1992) benchmarks. The temperature elasticity to radiative forcing (ξ_T) is estimated at 0.084, the abatement-sector exit rate at 0.06, and the entry congestion cost at 5.63.&lt;/p&gt;
&lt;p&gt;The paper implements projections from 2019 to 2100 under three IPCC-aligned scenarios (SSP1–1.9, SSP2–4.5, SSP3–7.0), focusing on the Paris Agreement target of limiting warming to below 2 degrees Celsius. In the laissez-faire (no-policy) scenario, emissions peak near 57 Gt CO2 in 2060 and 70 Gt in 2100, producing roughly 4 degrees Celsius of warming by 2100, with damages reaching 4 percent of GDP per year. In the below-2-degree scenario with a carbon tax only, the carbon tax must rise to approximately $480 per ton by 2080, abatement cost reaches 3.4 percent of GDP in 2060, and cumulative GDP loss from 2019 to 2060 totals $258 trillion (averaging $6.3 trillion per year, or 4.9 percent of 2019 world GDP). This is the baseline against which subsidies are evaluated.&lt;/p&gt;
&lt;p&gt;Two subsidy experiments are run, both fully financed by carbon tax revenue (budget neutral by construction). First, a subsidy targeted only at incumbent abatement firms (intensive margin): this immediately compresses the abatement price from 2.5 times to 1.5 times the price of the final good, reduces aggregate abatement cost from 2 percent to 0.8 percent of GDP in 2040, and brings the carbon tax needed to hit the emissions target down from $300 to $160 per ton in 2040. However, by lowering incumbents&amp;rsquo; labor costs and raising the equilibrium wage, the intensive-margin subsidy raises the cost of startup entry and reduces the number of abatement firms over time, deteriorating long-run competition.&lt;/p&gt;
&lt;p&gt;Second, an optimal subsidy that allocates carbon revenues between incumbents and startups. The optimal split is determined by maximizing social welfare (the infinite discounted sum of household utility) over a grid of subsidy shares. The welfare function is concave in the startup share, with a maximum at 60 percent of revenues to startups and 40 percent to incumbents. Under this optimal policy, the number of firms in the abatement sector nearly doubles relative to the baseline by 2050, the abatement price falls sharply, and the carbon tax needed to achieve the same emissions path drops to $125 per ton in 2040 versus $300 in the no-subsidy baseline. Cumulative GDP loss from 2019 to 2060 falls to $141 trillion ($138 trillion in one presentation, $141 trillion in another), saving approximately $120 to $123 trillion relative to the carbon-tax-only scenario, equivalent to roughly $2.9 trillion per year. The abatement price is reduced by more than a factor of 2.5 under the optimal subsidy regime.&lt;/p&gt;
&lt;p&gt;Present-value GDP subsidy multipliers (the ratio of discounted GDP gain to discounted subsidy expenditure) exceed 2.0 through 2035 and remain above 1.78 through 2060, with consumption multipliers ranging from 1.42 to 1.90 over the same horizon. These large multipliers reflect the competition-enhancing effect of startup subsidies: by accelerating firm entry, the policy lowers abatement prices for all final-goods producers, amplifying the direct subsidy impact. The largest GDP gains are concentrated in the first decade (2019–2030), when subsidies rapidly reduce the abatement price and induce firm entry. The scope condition for these results is the below-2-degree (SSP1–1.9) scenario with a simultaneous carbon-tax-and-subsidy announcement in 2019, a world-representative aggregate model, and the assumption that carbon tax revenues are fully recycled into the abatement sector rather than used for general government expenditure.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-central-market-failure-the-paper-addresses-and-why-does-it-make-a-carbon-tax-alone-insufficient"&gt;Q1. What is the central market failure the paper addresses, and why does it make a carbon tax alone insufficient?&lt;/h3&gt;
&lt;p&gt;The abatement goods sector is initially immature and highly concentrated (10 percent of firms account for roughly 80 percent of operating revenue). In the decentralized equilibrium, each final-goods firm is atomistic with respect to climate damage and so does not voluntarily abate. The carbon tax corrects this free-rider problem, but because the abatement market is imperfectly competitive, abatement goods are priced at a monopolistic markup (the abatement price begins at 2.5 times the price of the final good). The high abatement price raises the cost of reducing emissions, depresses the optimal abatement effort, and magnifies the GDP recession. A carbon tax alone thus generates a $258 trillion cumulative GDP loss by 2060. The paper&amp;rsquo;s main point is that subsidizing entry into the abatement sector introduces competition that compresses the markup, lowering both the abatement price and the required carbon tax rate.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-models-identification-strategy-and-what-are-the-main-econometric-challenges"&gt;Q2. What is the model&amp;rsquo;s identification strategy and what are the main econometric challenges?&lt;/h3&gt;
&lt;p&gt;The model is identified through full-information Bayesian maximum likelihood on five world aggregate series, 1961–2019. Climate block parameters are largely taken from DICE (Nordhaus 1992, 2018), narrowing the estimation to five structural parameters: initial output growth rate, initial emissions-to-output decoupling rate, temperature elasticity to radiative forcing (ξ_T), abatement-sector exit rate (δ_A), and entry congestion cost (χ). The main econometric challenges are (i) stochastic growth trends, which make standard linearization around a fixed point invalid — addressed with the extended-path solution method — and (ii) forming the likelihood for a nonlinear model, addressed with an inversion filter (Fair and Taylor 1983; Guerrieri and Iacoviello 2017) rather than computationally expensive particle filters. A drawback acknowledged by the authors is that Jensen&amp;rsquo;s inequality collapses to equality in the extended-path approach, so nonlinear uncertainty from future shocks is not captured — the same limitation that applies to standard linearized DSGE models.&lt;/p&gt;
&lt;h3 id="q3-how-are-the-intensive-and-extensive-margins-of-adjustment-to-the-carbon-tax-distinguished-in-the-model-and-why-does-this-distinction-matter-for-policy"&gt;Q3. How are the intensive and extensive margins of adjustment to the carbon tax distinguished in the model, and why does this distinction matter for policy?&lt;/h3&gt;
&lt;p&gt;The intensive margin refers to incumbent abatement firms increasing the quantity produced of existing varieties. The extensive margin refers to households creating new startups that introduce additional varieties of abatement goods. The distinction matters because (i) more varieties increase competition and compress the abatement price (via a price-index formula: aggregate abatement price falls with firm numbers), and (ii) the two margins respond differently to subsidy design. A subsidy only to incumbents immediately lowers production costs and the abatement price but raises the equilibrium wage, which increases the sunk cost for prospective entrants and crowds out startup entry over time, ultimately harming competition. A subsidy to startups has a delayed effect — startups take one period to begin producing — but generates a sustained competitive effect that eventually exceeds the immediate gain from the incumbent-only policy. The welfare-maximizing policy therefore combines both, weighting startups at 60 percent.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-optimal-subsidy-split-and-how-is-it-determined"&gt;Q4. What is the optimal subsidy split and how is it determined?&lt;/h3&gt;
&lt;p&gt;The optimal split allocates 60 percent of carbon tax revenues to subsidizing startups&amp;rsquo; sunk entry costs and 40 percent to reducing incumbents&amp;rsquo; production costs (labor input subsidies). This is determined by computing the present value of household welfare (infinite discounted sum of utilities evaluated at 2019 when the policy is announced) for each value of the subsidy share on a fine grid. The welfare function is strictly concave in the startup share, rising until the startup share reaches 0.6 and declining thereafter. The intuition for concavity is that subsidizing startups has a long-horizon payoff (gradual entry and competition), while subsidizing incumbents has an immediate payoff (price reduction) but a long-run cost (reduced entry incentive). The optimum balances these dynamics.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-quantitative-effects-of-the-optimal-subsidy-on-the-carbon-tax-path-abatement-prices-and-firm-numbers"&gt;Q5. What are the quantitative effects of the optimal subsidy on the carbon tax path, abatement prices, and firm numbers?&lt;/h3&gt;
&lt;p&gt;Relative to the no-subsidy carbon-tax-only baseline: (1) The carbon tax needed to hit net-zero by 2060 falls from approximately $300 per ton in 2040 to $125 per ton under the optimal subsidy, and from approximately $390–$480 per ton in later years to correspondingly lower values. (2) The abatement price is reduced by more than a factor of 2.5 over the horizon. (3) The number of firms in the abatement goods sector nearly doubles by 2050 relative to the baseline. (4) Abatement cost as a share of output falls substantially, from the baseline peak of approximately 3.4 percent of GDP in 2060 to a lower trajectory. (5) Detrended output in 2040 improves from approximately -3 percent (baseline) to -1 percent under the optimal subsidy, and from -3.2 percent to -2 percent in 2050. These numbers are conditional on the below-2-degree warming scenario and the announced policy starting in 2019.&lt;/p&gt;
&lt;h3 id="q6-how-large-are-the-subsidy-fiscal-multipliers-and-what-drives-them"&gt;Q6. How large are the subsidy fiscal multipliers and what drives them?&lt;/h3&gt;
&lt;p&gt;GDP subsidy multipliers (present value of GDP gain per unit of present value of subsidy expenditure) are approximately 2.27 at the 2030 horizon, 2.03 at 2035, 1.89 at 2040, 1.81 at 2045, 1.78 at 2050, 1.80 at 2055, and 1.85 at 2060. Consumption multipliers are uniformly lower but remain above 1.4 throughout. The high multipliers are driven by the competition channel: each dollar of subsidy to startups reduces the abatement price for all final-goods producers economy-wide, amplifying the direct expenditure effect many times over. Multipliers exceed 2 in the early years when startup entry is most rapid and the abatement-price reduction is sharpest. The slight uptick in multipliers at the 2060 horizon reflects the long-run dynamics of the abatement sector reaching a more competitive equilibrium.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-role-of-the-dice-climate-block-and-what-simplifications-are-made-relative-to-state-of-the-art-climate-science"&gt;Q7. What is the role of the DICE climate block and what simplifications are made relative to state-of-the-art climate science?&lt;/h3&gt;
&lt;p&gt;The climate block is taken directly from DICE-1992 and DICE-2016R2 (Nordhaus 1992, 2018). It models atmospheric CO2 accumulation, radiative forcing from CO2 and non-CO2 sources, and two-box (surface and deep-ocean) temperature dynamics. Key DICE parameters (φ_11, φ_12, φ_21, φ_22, ξ_M, M_1750, damage cost a) are calibrated to match DICE values. The temperature sensitivity parameter ξ_T is estimated from the data rather than calibrated, yielding 0.084, slightly below DICE 2013 and 2016 values. The authors explicitly note that more advanced climate blocks are important for physical risk assessment but have &amp;rsquo;little added value&amp;rsquo; for transition risk analysis, which concerns the costs of policy, not the physical hazard. The non-CO2 radiative forcing follows a deterministic path that caps at F_max by 2100. The damage function is quadratic in surface temperature: Φ(T_t) = 1/(1+aT_t^2). In the laissez-faire scenario, this implies damages of 1.5 percent of GDP by 2050 and 4 percent by 2100.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-compare-to-the-standard-dice-model-and-what-does-the-comparison-reveal"&gt;Q8. How does the paper compare to the standard DICE model and what does the comparison reveal?&lt;/h3&gt;
&lt;p&gt;The authors estimate both the E-DSGE (with endogenous firm entry in the abatement sector) and a version equivalent to DICE (with perfect competition and no firm-entry dynamics) on the same data. Both models match the empirical second moments (standard deviations and autocorrelations of the five observables) comparably, so standard information criteria cannot discriminate between them. The key difference is that the E-DSGE model reproduces the standard deviation and autocorrelation of patent growth (the proxy for abatement-sector entry), which the DICE version cannot by construction (it has no entry shock). In DICE-like environments, the abatement sector is assumed competitive from the outset and the abatement price equals 1 (the final-goods price), so there are no dynamics in abatement pricing or firm numbers. This means DICE models understate transition costs when the abatement market is initially concentrated, and miss the welfare gain from competition-enhancing policies.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-role-of-the-endogenous-market-structure-mechanism-and-how-does-it-relate-to-solar-photovoltaic-markets"&gt;Q9. What is the role of the endogenous market structure mechanism and how does it relate to solar photovoltaic markets?&lt;/h3&gt;
&lt;p&gt;The paper argues the solar PV market provides historical validation of the model mechanism. From the late 1970s to 2019, the cumulative number of solar PV patents increased dramatically while module costs fell precipitously (the cost of solar PV modules in 2019 USD per watt fell 45 percent between 1990 and 2000, 58 percent between 2000 and 2010, and 81 percent between 2010 and 2019). The model predicts exactly this pattern: an initial carbon policy raises expected profits in the abatement sector, inducing entry, which intensifies competition and compresses prices. The initial abatement price in the model (2.5 times the final-goods price) eventually falls below 1 after 2040 under a carbon-tax-only policy. The paper notes the solar sector&amp;rsquo;s trajectory was partly driven by government subsidies in several countries, consistent with the model&amp;rsquo;s policy recommendation.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-main-shock-processes-in-the-model-and-what-do-impulse-response-functions-reveal"&gt;Q10. What are the main shock processes in the model and what do impulse response functions reveal?&lt;/h3&gt;
&lt;p&gt;Five structural shocks are estimated: TFP (productivity), government spending, CO2 emissions, firm entry (innovation), and temperature. All are AR(1) processes. Estimated AR(1) coefficients: productivity 0.949, government spending 0.867, CO2 emissions 0.940, firm entry 0.592, temperature 0.181 — so temperature shocks are nearly serially uncorrelated at annual frequency. Generalized impulse response functions (computed at 2019 state variables, averaged over 500 draws) show: (1) A positive productivity shock raises output and worsens emissions, stimulating abatement-sector entry and reducing the abatement price. (2) A positive CO2 emissions shock triggers a sharp abatement effort and firm entry, but depresses output by almost 5 percent in the short run. (3) A government spending shock (demand shock) raises final-good production, worsens emissions, but crowds out abatement — abatement effort and firm numbers fall 5 percent and 1.1 percent respectively. (4) A firm-entry shock raises firm numbers by nearly 10 percent at peak, reducing abatement prices and encouraging abatement effort without increasing emissions. (5) A temperature shock depresses output by more than 6 percent initially, reducing emissions and abatement effort, and shrinking the abatement sector while pushing abatement prices up.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-three-ipcc-aligned-scenarios-used-in-the-projections-and-how-do-they-differ"&gt;Q11. What are the three IPCC-aligned scenarios used in the projections and how do they differ?&lt;/h3&gt;
&lt;p&gt;The three scenarios correspond to SSP1–1.9, SSP2–4.5, and SSP3–7.0. (1) Below +2 degrees C (SSP1–1.9): carbon neutrality by 2060, followed by negative emissions (up to -10 Gt by 2100). Requires the carbon tax to rise to approximately $480 per ton by 2080. Abatement cost reaches 3.4 percent of GDP in 2060. This is the scenario used for the policy experiments. (2) Below +3 degrees C (SSP2–4.5): carbon neutrality delayed to shortly after 2100. Carbon tax rises gradually to $300 per ton by 2100. Abatement cost rises to 0.5 percent of GDP in 2050 and 1.2 percent by 2100. Detrended output falls to -3 percent by 2060. (3) +4 degrees C (SSP3–7.0): no policy, laissez-faire. Emissions peak at 57 Gt in 2060 and 70 Gt in 2100. Temperature rises approximately 4 degrees C by 2100. Damages reach 4 percent of GDP per year by 2100. Detrended output decreases from 3 percent to -1 percent by 2050 and -3 percent by 2100 due to climate damage alone.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-main-policy-implications-and-their-scope-conditions"&gt;Q12. What are the main policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central implication is that carbon tax revenues should not be recycled to households as lump-sum transfers (the conventional approach in environmental economics) but should instead be used to subsidize entry and operation in the abatement goods sector. The welfare-maximizing split is 60 percent to startups and 40 percent to incumbents. This reduces the cumulative GDP loss from $258 trillion to approximately $138–141 trillion by 2060, saving roughly $120–123 trillion total ($2.9 trillion per year on average). Scope conditions: (1) The result is conditional on the below-2-degree Paris scenario — less stringent emissions targets require lower carbon taxes and generate smaller transition costs, so the absolute gain from subsidies would be smaller. (2) The policy must be announced credibly in advance (2019 in the simulation) so that firms adjust expectations and entry decisions. (3) The model abstracts from capital, cross-country heterogeneity, sector-level differences, and physical risks from climate change. (4) Stochastic uncertainty about future shocks is not incorporated into the policy optimization (extended-path solution collapses uncertainty around the deterministic path). The authors suggest future work should evaluate the optimal policy accounting for stochastic climate and economic risks (following Cai and Lontzek 2019).&lt;/p&gt;
&lt;h3 id="q13-how-does-the-paper-relate-to-prior-e-dsge-and-iam-literature-and-what-is-novel"&gt;Q13. How does the paper relate to prior E-DSGE and IAM literature, and what is novel?&lt;/h3&gt;
&lt;p&gt;The paper positions itself relative to two literatures. First, integrated assessment models (IAMs) originating with DICE (Nordhaus 1992, 1994): IAMs provide long-run analysis but lack microfounded expectations and uncertainty. Second, E-DSGE models (Fischer and Springborn 2011; Heutel 2012; Angelopoulos et al. 2013; Golosov et al. 2014; Annicchiarico and Di Dio 2015, 2017; Diluiso et al. 2021): these have microfoundations and handle short-run dynamics well but typically operate in a linearized, stationary framework unsuited for long-run climate trends. Some prior E-DSGE work includes endogenous entry (Annicchiarico et al. 2018; Shapiro and Metcalf 2021) but focuses on short-run analysis or specific country (U.S.) settings. The paper&amp;rsquo;s novelties are: (1) Merging DICE with a BGM-style endogenous market structure for the abatement sector in a unified framework suitable for long-run analysis; (2) Nonlinear estimation of the E-DSGE model using the extended-path plus inversion-filter approach — the authors claim this is the first attempt to estimate a nonlinear E-DSGE with both environmental and macroeconomic trends; (3) Distinguishing intensive and extensive margins of abatement-sector adjustment and optimizing the subsidy split between them; (4) Computing present-value subsidy multipliers for climate policy.&lt;/p&gt;
&lt;h3 id="q14-what-are-the-main-limitations-and-caveats-acknowledged-by-the-authors"&gt;Q14. What are the main limitations and caveats acknowledged by the authors?&lt;/h3&gt;
&lt;p&gt;The authors acknowledge several limitations. (1) Capital is excluded from the production function to keep the model tractable given the focus on the abatement goods sector and endogenous entry. (2) The model is a world aggregate with no cross-country heterogeneity; a multicountry model would be needed to study distributional effects across nations. (3) The policy analysis is conditional on the below-2-degree scenario and does not account for uncertainty about future economic and climate conditions — the extended-path method does not incorporate stochastic uncertainty in the forward-looking path. (4) The analysis does not account for the positive benefits of avoided physical risk from climate change (reduced damages in alternative scenarios are noted but not attributed to subsidy policy per se). (5) Non-CO2 radiative forcing is modeled as a simple deterministic path, which simplifies the climate dynamics. (6) The comparison with DICE via second moments rather than formal model selection criteria (since the DICE version has one fewer observable and one fewer shock) limits the formal identification of the endogenous entry mechanism. (7) The model does not include labor market frictions, nominal rigidities, or financial frictions, all of which could affect transition dynamics.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Abatement goods sector&lt;/strong&gt;: In this paper, the sector producing intermediate inputs (abatement goods) purchased by final-goods firms to reduce their CO2 emissions. The sector is initially immature and highly concentrated, with high barriers to entry that prevent competition and keep abatement prices above the price of the final good. The paper models this sector with endogenous firm entry following Bilbiie, Ghironi, and Melitz (2012), distinguishing between incumbents (intensive margin) and startups (extensive margin).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Transition risk&lt;/strong&gt;: In this paper, the macroeconomic cost — in terms of GDP loss, employment diversion, and abatement expenditure — of implementing climate policy (specifically a carbon tax path) to achieve net-zero emissions by 2060. Transition risk is distinct from physical risk (climate damage to productivity); the paper focuses exclusively on transition risk and does not account for avoided physical risk when evaluating policy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous market structure&lt;/strong&gt;: The property that the number of firms (varieties) in the abatement goods sector is not fixed but responds endogenously to expected future profits, sunk entry costs, and exit shocks. Following Bilbiie et al. (2012), the paper models a free-entry condition where households create startups until the marginal cost of entry (sunk cost) equals the expected discounted value of future profits. This endogeneity allows the model to capture how carbon taxes and subsidies affect abatement-sector competition and prices over time.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive margin vs. extensive margin (abatement sector)&lt;/strong&gt;: The intensive margin refers to adjustment by existing (incumbent) abatement firms — increasing production of current varieties when demand rises. The extensive margin refers to the creation of new firms (startups) that introduce additional varieties. The paper shows these margins respond differently to subsidy design: incumbent subsidies have immediate price effects but crowd out entry; startup subsidies have delayed effects but generate lasting competitive pressure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extended-path solution method&lt;/strong&gt;: A numerical method (Fair and Taylor 1983; Adjemian and Juillard 2014) for solving nonlinear rational-expectations models with stochastic growth trends. In each period, agents are surprised by current shocks but expect future shocks to be zero on average (consistent with rational expectations). The method provides accurate solutions while accounting for model nonlinearities, and is combined with an inversion filter to form the likelihood function for Bayesian estimation. It is used here instead of standard log-linearization, which would be invalid under unbalanced growth dynamics.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Subsidy multiplier (present value)&lt;/strong&gt;: The ratio of the discounted cumulative GDP gain (or consumption gain) to the discounted cumulative subsidy expenditure over a given horizon, in the spirit of fiscal multipliers (Feve and Sahuc 2017; Leeper et al. 2017). In this paper, these multipliers measure the efficiency of redirecting carbon-tax revenues to abatement-sector subsidies. GDP multipliers exceed 2.0 through 2035 because the competition-enhancing effect of startup subsidies lowers abatement prices economy-wide, amplifying the direct expenditure impact.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Damage function&lt;/strong&gt;: The function Phi(T_t) = 1/(1 + aT_t^2) in the TFP equation, where T_t is the surface temperature anomaly and a is a calibrated damage parameter taken from DICE-2016R2. It captures the reduction in total factor productivity caused by climate change. The function implies damages of 4 percent of GDP per year by 2100 under the laissez-faire scenario (approximately 4 degrees C warming), and less than 1 percent under the below-2-degree scenario.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inversion filter&lt;/strong&gt;: A computationally efficient method for evaluating the likelihood function of a nonlinear dynamic model (Fair and Taylor 1983; Guerrieri and Iacoviello 2017; Atkinson et al. 2020). Instead of particle-filter simulation, it analytically recovers the sequence of structural shocks by inverting the observation equations for a given set of initial conditions and parameter values. Combined with the extended-path solution, it allows Bayesian estimation of the nonlinear E-DSGE model on world data.&lt;/p&gt;</description></item><item><title>Expecting Floods: Firm Entry, Employment, and Aggregate Implications</title><link>https://macropaperwarehouse.com/papers/expecting-floods-firm-entry-employment-and-aggregate-implications/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/expecting-floods-firm-entry-employment-and-aggregate-implications/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper studies how the &lt;em&gt;expectation&lt;/em&gt; of rising flood risk — distinct from realized flood events — reshapes where firms locate, where workers live and how much they work, and what this implies for U.S. aggregate output. The motivation is climate-driven: roughly 6 million Americans lived within a 100-year flood zone in 1998, rising to 13 million by 2018, and FEMA floodplains are projected to grow about 45% by century&amp;rsquo;s end. Prior work largely studied actual floods or housing-price effects; this is among the first to examine firm entry and employment responses to anticipated risk.&lt;/p&gt;
&lt;p&gt;Data and design: The authors digitize FEMA Special Flood Hazard Zone maps (historic Q3 maps tied to 1998 Flood Insurance Rate Maps, and 2018 National Flood Hazard Layer), measuring flood risk as the share of land area within flood zones at the county and ZIP-code (ZCTA) level. Average flood-zone share rose 1.5 percentage points from 1998 to 2018, with a 20-pp increase at the 90th percentile of ZIP-level changes. Firm entry/exit, employment, population and county real GDP come from Census Business Dynamics Statistics, ZIP Codes Business Patterns, and BEA; actual flood events come from the Dartmouth Flood Observatory. The baseline specification is a two-period (1998, 2018) fixed-effects regression with county (or ZCTA) fixed effects, state-by-year fixed effects, demographic/economic controls (female labor share, manufacturing share, population density, China import-penetration change), and a control for actual flooded area.&lt;/p&gt;
&lt;p&gt;Main reduced-form findings: A one-standard-deviation (7-percentage-point) increase in flood risk over 1998-2018 reduced firm entry by 1.2%, employment by 1.2%, population by 0.8% (smaller than employment, implying both relocation and labor-supply margins), and real GDP by 2.4%. Firm exits also &lt;em&gt;declined&lt;/em&gt; with higher risk (smaller magnitude), reflecting reduced business dynamism. A county at the 90th percentile of risk increase saw a 3.3% drop in firm entry. ZIP-level estimates are similar. An IV using the interaction of rest-of-state risk change with local geo-climatic conditions (rainfall, temperature, evaporation) yields comparable magnitudes (entry -1.2%, employment -1.4%, GDP -2.2%); a placebo (1990-1998 outcomes) test is insignificant. In sharp contrast, actual flood &lt;em&gt;events&lt;/em&gt; had negligible effects on entry, exit, employment and population, but a one-SD (0.4) increase in flooded-area share lowered real GDP by 0.2% in the same year, driven by current-year shocks (lagged effects negligible).&lt;/p&gt;
&lt;p&gt;Model and quantification: The authors build a spatial-equilibrium model (McFadden 1978 location choice, Krugman 1980 monopolistic competition) with M = 2,772 counties (96% of 2018 GDP), σ = 5, exit rate κ = 0.08. Flood risk operates through three channels: direct damage, an employment channel (relocation + endogenous labor supply), and a love-of-variety channel (fewer firms). Damage parameters are disciplined by reduced-form evidence (δ = 0.005, δκ = 0.003) and Barrage (2020) (η = 0.002); labor-supply elasticities φL = 1.55, φM = 0.83 are set by indirect inference targeting employment and population responses. Non-targeted moments (output, entry, exit) match the data.&lt;/p&gt;
&lt;p&gt;Counterfactuals: Eliminating 2018 flood risk shows it reduced aggregate output by 0.52% (employment -0.31%, firm entry -0.30%, welfare -0.51%). Decomposition: direct damage -0.11% (21%), labor relocation 0%, labor supply -0.33% (63%), variety -0.08% (15%) — so about 80% of the loss is expectation-driven and 20% direct damage. Effects are highly unequal: top-5% and top-1% counties (by output loss) lost 7.9% and 13.9% of output. A projected 4.5% rise in at-risk properties (2020-2050) would cut output 0.12%. Extensions (entry costs in goods, interregional trade, capital and land) yield somewhat larger losses (0.57%, 0.62%, 0.67%). Policy implication: counting only direct damages badly understates disaster costs and the social cost of carbon, because firms and workers rationally adjust to anticipated risk.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The core design is a two-period (1998 and 2018) fixed-effects regression of log outcomes (firm entry, exit, employment, population, real GDP) on the share of land in FEMA flood zones, absorbing locality fixed effects (time-invariant characteristics like industry composition), state-by-year fixed effects (statewide growth/business cycles), demographic/economic controls, and a control for actual flooded area. The main threat is measurement error in FEMA risk maps: some underlying data are outdated, and political-economy incentives lead politicians and homeowners to resist map updates to avoid higher insurance premiums, so designations may reflect politics rather than true risk. A second threat is omitted local economic trends correlated with both risk and outcomes. The authors address measurement error with a Bartik-type IV (rest-of-state average risk change interacted with own geo-climatic features — satellite temperature, cumulative rainfall, evaporation), controlling for cumulative past flooded area. IV estimates are close to the fixed-effects ones (entry -1.2%, employment -1.4%, GDP -2.2%), with first-stage KP F-statistics around 63-66. A placebo/pre-trend test (regressing 1990-1998 changes on 1998-2018 risk changes, following Goldsmith-Pinkham et al. 2020) yields small, insignificant coefficients, arguing against omitted-trend confounding.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically-and-in-the-model"&gt;Q2. What are the main mechanisms, and how are they distinguished empirically and in the model?&lt;/h3&gt;
&lt;p&gt;Three channels: (1) direct damage — realized floods lower firm productivity and firm survival; (2) employment channel — anticipated risk lowers real wages/amenities, prompting out-migration and reduced labor supply per household; (3) love-of-variety — fewer firms enter, reducing the variety component of welfare/output. Empirically, the authors distinguish &lt;em&gt;flood risk&lt;/em&gt; (long-run anticipation) from &lt;em&gt;flood events&lt;/em&gt; (short-run realization) by estimating both: risk hits entry/employment/population strongly while events do not, but events hit current-year GDP (productivity) while risk hits it more through adjustment. In the model, direct damages are calibrated from the actual-flood GDP and exit responses (δ, δκ); the employment and variety channels are separated in the counterfactual by sequentially allowing population shares, then labor supply, then variety to respond. The decomposition attributes -0.11% to direct damage, ~0% to labor relocation (offsetting in- and out-migration), -0.33% to labor supply, and -0.08% to variety.&lt;/p&gt;
&lt;h3 id="q3-why-does-population-fall-less-than-employment-and-why-do-firm-exits-decline"&gt;Q3. Why does population fall less than employment, and why do firm exits decline?&lt;/h3&gt;
&lt;p&gt;Employment falls 1.2% while population falls only 0.8% for a one-SD risk increase, implying the response is not purely relocation — remaining households also reduce labor supply. This motivates introducing a positive labor-supply elasticity φL alongside migration elasticity φM, capturing &amp;lsquo;immobile labor&amp;rsquo; (as in Autor et al. 2013) where some workers cut hours rather than move. Firm exits decline with higher risk even though floods mechanically raise closures, because higher risk deters entry so much that the stock of firms shrinks, lowering the base of firms that can exit — reflecting reduced business dynamism rather than greater firm survival.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented"&gt;Q4. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Large regional dispersion. While national output fell 0.52%, the top-5% and top-1% counties by output loss lost 7.9% and 13.9% of output respectively (the abstract describes top-5% losses of 7-14%). The hardest-hit counties — coastal and riverine areas in southern and eastern regions (e.g., Cape May NJ, Marion County FL, Sharkey County MS) — lost population, labor supply per household, and firms (top-1% counties: -6.1% population, -4.7% labor supply per household, -10.8% firms). Conversely, mildly affected counties (some Midwestern) were &amp;lsquo;winners,&amp;rsquo; gaining in-migration, more firm entry, and higher labor supply per worker. For the 2020-2050 projection, direct damages play a &lt;em&gt;smaller&lt;/em&gt; relative role (12% vs 21% for 2018) because projected risk increases are more positively correlated with regional productivity, amplifying aggregate adjustment effects.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;(1) Controlling vs. not controlling for actual flooded area leaves risk estimates stable. (2) ZIP-code-level regressions exploiting finer spatial variation give similar magnitudes (establishments -0.233, employment -0.240, payroll -0.221). (3) Restricting to counties with available Q3 (1998) FEMA maps gives qualitatively similar, slightly larger estimates (Appendix Table A.2); the authors conservatively use baseline estimates for calibration. (4) IV estimation and (5) placebo pre-trend tests as above. (6) Lagged flood shocks (Appendix A.4) have negligible effects, confirming floods act through current-year productivity. (7) Model non-targeted moments (output, entry, exit) match data, and model-data correlations of regional GDP, population, emp-to-pop ratio, and firm count are near unity. (8) The implied regional-population-to-real-wage elasticity φM(1+φL) ≈ 2.1 lies within the 1.1-2.5 range from Fajgelbaum et al. (2018).&lt;/p&gt;
&lt;h3 id="q6-what-model-extensions-are-explored-and-how-do-results-change"&gt;Q6. What model extensions are explored and how do results change?&lt;/h3&gt;
&lt;p&gt;Four extensions, all yielding somewhat larger output losses than the 0.52% baseline: (1) entry costs paid partly/fully in final goods rather than labor — with α=1 the loss is 0.57%, because final-goods prices respond more to risk than wages; (2) interregional trade with traded/nontraded sectors — requires a larger labor-supply elasticity (φL=1.72) to match data, giving a 0.62% loss; (3) capital (mobile, rented at constant global rate) and land (fixed, congestion force) in production — 0.67% loss, since risk also lowers the capital-to-labor ratio (by 0.34%) as capital becomes relatively more expensive, outweighing land congestion (small land share). The authors read the modest size of these differences as evidence the simplified baseline captures the key forces.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It contributes to climate-spatial-economics work (Costinot et al. 2016, Desmet et al. 2021, Alvarez &amp;amp; Rossi-Hansberg 2021, Rudik et al. 2021). Closest are three flood-aggregate studies: Desmet et al. (2021) on coastal-flooding costs via migration and local technology investment; Balboni (2019) on infrastructure misallocation under sea-level risk; Lin et al. (2021) on coastal housing construction. Differences: prior work focuses mainly on coastal land inundation from sea-level rise, whereas this paper uses historic flood-zone designation maps capturing overall flood risk and studies production damage rather than land loss; and it reconciles structural estimates with reduced-form evidence showing firm/worker responses to &lt;em&gt;risk&lt;/em&gt; differ from responses to &lt;em&gt;actual floods&lt;/em&gt;. Relative to Kocornik-Mina et al. (2020) (satellite-nightlight evidence that floods reduce output transiently), this paper confirms the short-run finding but shows risk has larger, longer-run effects via behavioral adjustment. It relates to Hino &amp;amp; Burke (2020) (same risk data; floods cut property values 1-2%), interpreting housing-price effects as amenity changes; their estimate implies a 0.3-0.6% utility loss, comparable to the paper&amp;rsquo;s calibrated amenity loss of 0.2%.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central implication is that evaluations counting only direct flood damages substantially understate true costs, since about 80% of the 0.52% 2018 output loss comes from expectation-driven adjustments (labor supply, migration, fewer firms) rather than the 20% direct damage. Direct damages (-0.11%) match FEMA&amp;rsquo;s ~$17B/year (~0.1% of GDP) estimate, validating the model&amp;rsquo;s lower bound. Policies addressing climate damage — and estimates of the social cost of carbon — should incorporate firms&amp;rsquo; and workers&amp;rsquo; long-run general-equilibrium adjustments. Scope conditions: the analysis is U.S.-specific (chosen for systematic flood-risk data), uses establishments as &amp;lsquo;firms,&amp;rsquo; abstracts from flood insurance (justified by near-actuarially-fair pricing evidence) and from explicit housing, treats unmapped areas as zero-risk, and assumes observed FEMA designations are the risk signal agents act on despite measurement error. The authors note the approach generalizes to other natural disasters.&lt;/p&gt;
&lt;h3 id="q9-what-are-notable-caveats-or-limitations"&gt;Q9. What are notable caveats or limitations?&lt;/h3&gt;
&lt;p&gt;GDP data do not capture variety/welfare changes, so the love-of-variety channel matters for welfare but is invisible in GDP-based estimates. The amenity parameter η is not directly estimated but imported from Barrage (2020) (output-to-utility damage ratio ~3); the authors note η has little effect on national productivity impact because amenity mostly drives offsetting migration. Labor supply is assumed fixed before shocks (micro-founded by job-search frictions). Flood insurance and housing are not modeled explicitly. Risk is measured by flood-zone land share, which is converted to flood probabilities {rm} via a regression of 2015-2019 actual flooded shares on 2018 zone shares. The two-period long-run design limits dynamics, and counties without FEMA maps are assigned zero risk.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Flood risk vs. flood events&lt;/strong&gt;: The paper sharply separates anticipated flood risk (the share of local land in FEMA Special Flood Hazard Zones, a long-run signal firms/workers observe and act on) from realized flood events (the share of area actually flooded in a given year, from Dartmouth data). Risk drives firm-entry and employment relocation; events drive transient productivity/GDP losses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Expectation effects (vs. direct damages)&lt;/strong&gt;: Output losses arising because firms and workers rationally adjust location, entry, and labor supply in anticipation of flood risk — comprising the employment and variety channels. In 2018 these accounted for about 80% (the employment channel 0.33% plus variety 0.08% of the 0.52% loss), four times the 20% from direct physical damage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Employment channel&lt;/strong&gt;: In the model, the mechanism by which higher flood risk lowers real wages and amenities, inducing both out-migration (relocation, ~0% net aggregate effect due to offsetting regions) and reduced labor supply per household (the dominant -0.33% component), governed by elasticities φM (migration) and φL (labor supply).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Love-of-variety channel&lt;/strong&gt;: The output/welfare loss from fewer firms entering under higher risk, operating through the CES variety term (agglomeration force 1/(σ-1)). It reduced 2018 output by 0.08% and matters for welfare but is not captured in GDP data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Direct damage channel&lt;/strong&gt;: The component of flood losses from realized floods lowering firm productivity (parameter δ=0.005) and destroying a fraction of firms (δκ=0.003) plus amenity loss (η=0.002), calibrated from the short-run actual-flood reduced-form estimates; it caused a 0.11% output decline in 2018 (21% of the total).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Indirect inference calibration&lt;/strong&gt;: The simulated-method-of-moments procedure (Gouriéroux &amp;amp; Monfort 1996) used to set labor-supply elasticities φL=1.55 and φM=0.83: running the same 1998-vs-2018 panel regressions on model-generated data and choosing elasticities so model employment and population responses to flood risk match the empirical coefficients.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Immobile labor&lt;/strong&gt;: Following Autor et al. (2013), the model feature that some households respond to local flood risk by reducing labor supply rather than relocating, which is why employment falls more (1.2%) than population (0.8%) and motivates a positive labor-supply elasticity φL.&lt;/p&gt;</description></item><item><title>Firm dynamics, monopsony, and aggregate productivity differences</title><link>https://macropaperwarehouse.com/papers/firm-dynamics-monopsony-and-aggregate-productivity-differences/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/firm-dynamics-monopsony-and-aggregate-productivity-differences/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; Firms are larger and grow faster over the life cycle in high-income countries, while labor markets in poorer countries are less competitive (employers hold more wage-setting power). The paper asks how important employer labor market power (monopsony) is for explaining cross-country differences in firm dynamics and aggregate productivity. The novelty is that beyond the standard static misallocation-of-workers channel, monopsony also distorts &lt;em&gt;selection into entrepreneurship&lt;/em&gt; and &lt;em&gt;productivity-enhancing technology adoption&lt;/em&gt;, potentially making the losses larger than prior static estimates suggest.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and setup.&lt;/strong&gt; Stylized facts come from the World Bank Enterprise Surveys (WBES), an establishment-level survey of non-agricultural, non-financial private firms with at least 5 full-time permanent employees, covering more than 90 countries from 2006 to 2021, merged with World Development Indicators GDP per capita (2017 constant USD). The estimation sample restricts to countries that ever had GDP per capita above 25,000 USD and to manufacturing firms with non-missing sales/workers/material/capital data, yielding 37,096 firm-year observations across 31 middle- and high-income countries (poorest: Kazakhstan, 19,615 USD in 2009; richest: Ireland, 91,791 USD in 2020). Local labor markets are defined as location-industry (2-digit ISIC v3.1) pairs. The model is a dynamic general-equilibrium neoclassical-monopsony model with occupational choice (entrepreneur vs. wage worker), endogenous productivity investment, and Card-et-al.-style taste-for-employer (amenity) differentiation that gives firms wage-setting power. It is calibrated to the Netherlands (GDP per capita 54,275 USD; median wage markdown 1.301, implying firm-level labor supply elasticity 3.318) via method of simulated moments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main quantitative findings.&lt;/strong&gt; Empirically, moving from poorer to richer countries in the sample, average firm age triples from 11 to nearly 30 years; annualized firm growth rises ~1.6 percentage points per year per doubling of GDP per capita; the share of firms doing R&amp;amp;D more than doubles (from ~15% to &amp;gt;40%); product innovation rises from 20% to 80% and process innovation from 20% to 50%; and median wage markdowns fall (from ~2.25 at 25,000 USD GDP per capita — workers paid ~55% below marginal product — to ~1.25 at 60,000 USD — paid 20-25% below). The calibrated model matches a right-skewed firm-size distribution, life-cycle growth, employer turnover, age distribution, and R&amp;amp;D share (sum of squared deviations between empirical and simulated moments = 1.7%). In counterfactuals raising the markdown from 1.2 to 3, average firm growth shrinks by more than half (from ~150% to ~50%), average firm size falls from ~60 to ~45 employees, the innovating share halves (from ~40% to ~25%), and average firm productivity is ~20% higher in competitive markets. Differences in wage markdown alone account for &lt;strong&gt;25%&lt;/strong&gt; of observed cross-country TFP variation (model TFP std dev 0.051 vs. data 0.201), and &lt;strong&gt;no less than 11%&lt;/strong&gt; across robustness checks. In a Netherlands-vs-Greece decomposition, about &lt;strong&gt;85%&lt;/strong&gt; of the model-implied TFP gap is attributable to lower technology adoption, ~9% to distorted selection into entrepreneurship, and ~6% to static employment reallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanisms and implications.&lt;/strong&gt; Labor market competition acts as a “skill-biased” force favoring high-productivity firms through three channels: (i) static labor reallocation toward high-productivity, low-amenity firms; (ii) improved selection into entrepreneurship (low-productivity high-amenity agents stop being able to profitably attract workers as ϵL rises); and (iii) higher returns to innovation. The policy implication is that raising labor market competition in less-developed economies could yield substantial productivity gains, and that prior static studies understate the cost of monopsony because they omit the dynamic investment/selection channels.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identificationcalibration-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification/calibration strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The model is calibrated to the Netherlands using a mix of externally set and internally estimated (method-of-simulated-moments) parameters. Externally: model period = 1 year; σν (Gumbel scale) normalized to 1; β = 0.961 (4% annual rate); δw = 0.025 (40-year working life); revenue elasticity of labor ξ = 0.333 (estimated via control function in Section 2); labor supply elasticity ϵL = 3.318 backed out from median markdown 1.301 via ϵL = 1/(µ−1). Six parameters {c_f, c_x, p_i, p_n, σ_z, σ_a} are estimated by MSM. The markdown itself is a key input and is estimated as the ratio of marginal revenue product of labor to wage, with revenue elasticity ξ from a standard control-function approach. Threats: the markdown estimate drives the whole quantitative exercise; the WBES sample is truncated at firms with ≥5 employees (biasing toward larger firms), addressed by re-estimating with imputed moments; and the cross-country counterfactual attributes all variation in ϵL to labor market power while holding all other parameters at Netherlands values, so other cross-country differences are not separately identified.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-mechanisms-and-how-are-they-distinguished-quantitatively"&gt;Q2. What are the three mechanisms and how are they distinguished quantitatively?&lt;/h3&gt;
&lt;p&gt;(1) Static labor allocation: lower competition raises marginal factor cost only for sufficiently high-productivity firms, reallocating employment toward less-productive, lower-paying employers. (2) Selection into entrepreneurship: when ϵL is low, amenities matter more for profits, letting low-productivity high-amenity agents profitably self-select into entrepreneurship. (3) Technology adoption: returns to innovation increase with ϵL, so weak competition lowers the share of firms investing. They are distinguished via a decomposition that sequentially fixes policy functions at benchmark levels: ~6% of the TFP loss is from employment allocation alone, ~85% from the distortion to innovation policy, and ~9% from distorted selection into entrepreneurship.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-across-firms-is-documented"&gt;Q3. What heterogeneity across firms is documented?&lt;/h3&gt;
&lt;p&gt;Firms differ in entrepreneurial productivity z and amenity a. Average revenue product of labor rises with productivity and falls with amenities, and this dispersion is much steeper under weak competition: the elasticity of APL with respect to productivity is 0.31 in the baseline (Netherlands) vs 0.79 in the counterfactual (Greece), and with respect to amenities -0.28 vs -0.81. High-productivity, low-amenity firms face the biggest barriers in less-competitive markets and stay inefficiently small; low-productivity, high-amenity firms are propped up. Innovation distortion is concentrated among high-productivity firms.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run-and-what-do-they-show"&gt;Q4. What robustness checks are run and what do they show?&lt;/h3&gt;
&lt;p&gt;Four main checks, each reported as the share of cross-country TFP variation explained (data std dev 0.201): (1) Productivity-amenity correlation — allowing entrants to draw correlated (z,a) with σ_za = 0.296 (matching Sockin 2024’s 0.622 wage-satisfaction correlation) lowers explained variation to ~15% (model std dev 0.030), because correlation reduces scope for reallocation. (2) Costs in terms of labor instead of final goods (per Klenow and Li 2025) gives ~22% (std dev 0.044). (3) Imputed firm-level moments covering all firms (not just ≥5 employees) gives ~14% (std dev 0.028). (4) Over-identified alternative identification using size/age/R&amp;amp;D shares and annualized growth gives ~11% (std dev 0.023). The headline range is therefore 25% baseline, no less than 11% across checks.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It builds on static monopsony cost estimates: Berger et al. (2022, eliminating US labor market power raises average wage 48%, welfare +6% of lifetime consumption); Armangüé-Jubert et al. (2025, labor market power explains 15% of GDP-per-capita gap over development); Deb et al. (2022, less competition lowered US low/high-skill wages 12% and 11%); Amodio et al. (2025b, eliminating monopsony in Peru raises earnings 26%); Bachmann et al. (2022, monopsony caused a 10% aggregate productivity loss in East Germany). Its contribution is to add the entrepreneurial-selection and innovation channels, yielding larger losses than static studies, and to bridge the monopsony-cost literature with the misallocation literature (Restuccia-Rogerson, Guner et al., Hsieh-Klenow).&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Raising labor market competition (higher firm-level labor supply elasticity) improves allocative efficiency, selection into entrepreneurship, and innovation, raising firm growth and aggregate productivity. Scope conditions: the quantitative results apply to middle- and high-income countries (sample restricted to those ever above 25,000 USD GDP per capita); the 25% headline depends on the assumption that initial productivity and amenities are independent (falls to ~15% under positive correlation); and the decomposition attributing 85% to innovation is specific to the Netherlands-vs-Greece comparison. The model treats labor supply elasticity differences as the sole varying parameter, so the counterfactuals isolate the labor-market-power channel rather than reproducing total cross-country income gaps.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-netherlands-vs-greece-comparison-specifically"&gt;Q7. What is the Netherlands-vs-Greece comparison specifically?&lt;/h3&gt;
&lt;p&gt;Greece has roughly half the GDP per capita of the Netherlands (29,000 vs 54,000 USD) and much weaker competition (wage markdown 2.623 vs 1.301, labor supply elasticity 0.616 vs 3.318). In the Greece counterfactual, average firm size is 26 vs 59 employees, life-cycle growth 84.5% vs 153%, average age 22.5 vs 30 years, and R&amp;amp;D investing share 18% vs 41%. Labor market competition differences explain 29% of the firm-size gap, 27% of the firm-age gap, and 74% of the R&amp;amp;D-share gap between the two countries.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-model-get-right-that-was-not-targeted"&gt;Q8. What does the model get right that was not targeted?&lt;/h3&gt;
&lt;p&gt;The firm size and age distributions are not targeted yet are matched: in the data ~57.6% of firms have &amp;lt;20 employees and ~6.2% have &amp;gt;100; ~60% of firms are under 30 years old and ~10% over 60. The estimated parameters imply investing firms are 15% more likely to grow (p_i=0.649 vs p_n=0.499); innovation and operating costs equal ~43% and ~8% of average incumbent profits respectively; standard errors are small, indicating informative moments.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Firm Heterogeneity, Market Power and Macroeconomic Fragility</title><link>https://macropaperwarehouse.com/papers/firm-heterogeneity-market-power-and-macroeconomic-fragility/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/firm-heterogeneity-market-power-and-macroeconomic-fragility/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Ferrari and Queirós ask why US recoveries have become progressively slower and argue that rising firm heterogeneity and market power — well-documented long-run trends — can substantially increase the probability that a moderate aggregate shock triggers a quasi-permanent slump rather than a transitory recession. They call this probability macroeconomic fragility.&lt;/p&gt;
&lt;p&gt;The theoretical framework is an RBC model with oligopolistic (Cournot) competition, endogenous firm entry, and elastic capital and labor supply (GHH preferences). The economy consists of many product markets; within each market, firms with heterogeneous idiosyncratic TFP compete in quantities, with the marginal firm earning zero net profit. A central complementarity drives the results: more competition raises factor shares and factor prices, which expands factor supply, which in turn allows more firms to enter, sustaining high competition. This complementarity can generate multiple stochastic steady-states — a high-competition, high-output regime and a low-competition, low-output regime.&lt;/p&gt;
&lt;p&gt;Two forces increase fragility by shrinking the basin of attraction around the high steady-state. First, a mean-preserving spread (MPS) in idiosyncratic TFP: the dominant firm expands market share, factor shares fall (market-power effect), the factor price index drops, and smaller firms approach their exit threshold — requiring only a smaller shock to trigger cascading exit. Second, rising fixed production costs: the unstable steady-state shifts toward the high steady-state, narrowing the gap and making downward transitions more likely.&lt;/p&gt;
&lt;p&gt;The model is calibrated three times — to match COMPUSTAT moments in 1975, 1990, and 2007 — varying only the log-normal standard deviation of idiosyncratic productivity (λ = 0.182, 0.213, 0.232) and the fixed cost parameter (c × 10⁻³ = 0.351, 0.691, 0.751). The fixed-to-total-cost ratio in COMPUSTAT rises from 21.9% in 1975 to 31.7% in 1990 to 36.9% in 2007; the standard deviation of log revenues rises from 1.59 to 1.91 to 2.04.&lt;/p&gt;
&lt;p&gt;The quantitative results are stark. The 1975 economy has a unimodal ergodic distribution (one stable steady-state); the 1990 and 2007 economies are bimodal (two stable steady-states). When subjected to the same TFP shock sequence (εt = −σε for four quarters), output falls 4.0% after five quarters in the 1975 economy, 5.1% in 1990, and 5.9% in 2007; after 100 quarters, the 2007 economy remains 6.3% below pre-shock output, against 3.0% for 1990 and 1.3% for 1975. For a larger shock (εt = −2σε for six quarters), only the 2007 economy transitions permanently to the low steady-state, with output 12.5% below trend after 100 quarters. The minimum shock required to trigger a downward transition is 6.84σε for the 1990 economy but only 1.62σε for the 2007 economy. In Monte Carlo simulations, the probability of a recession exceeding 10% of output over a 40-quarter window is 1.7% in 1975, 12.4% in 1990, and 19.6% in 2007. In expectation, the 2007 economy experiences such a recession every 70 years, the 1990 economy every 95 years, and the 1975 economy every 380 years.&lt;/p&gt;
&lt;p&gt;Applying the 2008–09 TFP shocks to the 2007-calibrated model generates a persistent deviation from trend: output is 12.1% below trend by 2019, investment 14.4% below, and hours 9.8% below — closely matching the data (14.2%, 14.7%, and 5.5% respectively). The same shocks applied to the 1975 and 1990 economies produce no permanent transition; by 2040 the 1975 (1990) economy is only 1.5% (4.7%) below trend.&lt;/p&gt;
&lt;p&gt;Cross-industry evidence corroborates the mechanism. Using US Census and BLS data on 791 six-digit NAICS industries, the authors find that a 1 percentage point higher pre-crisis four-firm concentration ratio (CR4) in 2007 is associated with 1.8–1.9 percentage points lower employment growth, 2–3 percentage points lower net firm entry, and a larger decline in the labor share between 2007 and 2016. These qualitative and quantitative patterns are matched by simulated cross-industry regressions from the model.&lt;/p&gt;
&lt;p&gt;On policy, an entry subsidy that eliminates fixed-cost barriers for the approximately 11.8% of markets with positive fixed costs can prevent downward transitions and yields a welfare gain of roughly 10% in consumption-equivalent terms in the 2007 economy. A revenue subsidy applied to all firms achieves welfare gains between 30% and 50% for a 20% subsidy rate, acting as a steady-state selection device by shifting probability mass from the low to the high competition regime. These gains are nonlinear: even a 5% revenue subsidy yields roughly a 20% welfare gain in the 2007 economy. The gains are in line with Edmond et al. (2023), who find welfare costs of markups up to 50%.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-models-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the model&amp;rsquo;s identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper is primarily theoretical and quantitative rather than identification-based in the econometric sense. The causal claim — that rising firm heterogeneity and fixed costs increase macroeconomic fragility — comes from two sources: (1) analytic comparative statics (Propositions 4–6) that formally show fragility rises with a mean-preserving spread on TFP or with fixed costs, and (2) calibration counterfactuals where the 1975, 1990, and 2007 economies face the same shock sequence but differ only in λ and c. The cross-industry regressions are reduced-form and subject to standard endogeneity concerns — pre-crisis concentration could be correlated with industry-specific demand shocks coinciding with 2008. The authors partially address this by including pre-crisis growth trends as controls and sector fixed effects, but do not use an instrumental variable for concentration.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-core-mechanism-linking-firm-heterogeneity-to-fragility-and-how-is-it-distinguished-from-steady-state-multiplicity"&gt;Q2. What is the core mechanism linking firm heterogeneity to fragility, and how is it distinguished from steady-state multiplicity?&lt;/h3&gt;
&lt;p&gt;The mechanism runs through factor markets. When idiosyncratic TFP dispersion rises (MPS), the dominant firm expands market share and charges a higher markup, depressing the aggregate factor share (Proposition 4). This reduces the factor price index and real wages, contracting labor supply. Marginal firms, already earning near-zero profits, move closer to their exit threshold. A smaller aggregate shock suffices to push them out, triggering cascading exit, a further collapse in competition, a further fall in factor prices, and a self-reinforcing transition to the low steady-state. Fragility is distinct from multiplicity: the existence of two steady-states is a necessary but not sufficient condition for fragility. Fragility specifically measures the size of the basin of attraction around the high steady-state from below — how large a shock is needed to trigger a downward transition. An economy can have two steady-states but be highly resilient if the basin is wide.&lt;/p&gt;
&lt;h3 id="q3-what-roles-do-the-three-model-channels-endogenous-market-structure-oligopolistic-markups-elastic-factor-supply-play-quantitatively"&gt;Q3. What roles do the three model channels (endogenous market structure, oligopolistic markups, elastic factor supply) play quantitatively?&lt;/h3&gt;
&lt;p&gt;The authors isolate each channel by shutting it down one at a time and comparing output volatility (Table 8). In the baseline, the standard deviation of log output is 0.063 and autocorrelation is 0.975. Fixing the number of firms (removing the endogenous market structure channel, leaving only elastic factor supply) reduces output standard deviation to 0.035, accounting for 55% of baseline volatility. Replacing oligopoly with monopolistic competition (constant markups, love-for-variety active) recovers 0.049 — approximately 78% of baseline — implying the endogenous markup channel accounts for about one-fourth of total amplification. The love-for-variety channel accounts for another approximately one-fourth. Crucially, all three alternative models exhibit unimodal ergodic distributions, confirming that all three channels are jointly required to generate steady-state multiplicity and the model&amp;rsquo;s nonlinear amplification.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented-and-how-does-it-motivate-the-models-calibration"&gt;Q4. What heterogeneity is documented and how does it motivate the model&amp;rsquo;s calibration?&lt;/h3&gt;
&lt;p&gt;Rising US firm heterogeneity is documented along three dimensions: (1) standard deviation of log revenues (sales) for COMPUSTAT firms, rising from 1.59 in 1975 to 1.91 in 1990 to 2.04 in 2007; (2) the average ratio of fixed (SG&amp;amp;A) to total costs (fixed + COGS), rising from 21.9% in 1975 to 31.7% in 1990 to 36.9% in 2007; (3) sales-weighted average markups for public firms rising from 1.28 in 1975 to 1.37 in 1990 to 1.46 in 2007 (from De Loecker et al., 2020). These moments are the calibration targets for the time-varying parameters λ and c. The structural parameters (elasticities of substitution σI = 1.46 and σG = 11.50) are time-invariant and calibrated jointly to the markup levels across the three years.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-papers-account-of-the-great-recession-differ-from-other-slow-recovery-theories"&gt;Q5. How does the paper&amp;rsquo;s account of the Great Recession differ from other slow-recovery theories?&lt;/h3&gt;
&lt;p&gt;Most related theories attribute slow recovery to (1) the zero lower bound on interest rates and constrained monetary policy (Christiano et al., 2015; Eggertsson et al., 2019; Guerrieri and Lorenzoni, 2017), (2) endogenous TFP decay through R&amp;amp;D decisions (Anzoategui et al., 2019; Bianchi et al., 2019; Queralto, 2020), or (3) declining firm entry per se (Clementi and Palazzo, 2016). Ferrari and Queirós instead argue the 2008 shock was not unusually large — the same shock does not cause a permanent transition in the 1975 or 1990 economies — but rather that the US economy had become structurally more fragile over the preceding decades due to rising concentration and fixed costs. The closest related model is Schaal and Taschereau-Dumouchel (2018), who also use coordination failures among oligopolistic firms to generate multiple steady-states. The key contribution of Ferrari and Queirós relative to that work is the explicit role of cross-sectional firm heterogeneity in determining the probability of transitions, and the empirical documentation that rising heterogeneity preceded the crisis.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-cross-industry-empirical-results-in-detail"&gt;Q6. What are the cross-industry empirical results in detail?&lt;/h3&gt;
&lt;p&gt;The dataset covers 791 six-digit NAICS industries from the US Census, SUSB, and BLS, with the concentration variable defined as CR4/CR50 (top-4 share scaled by top-50 share). Key results: (1) Employment: a 1 pp higher CR4/CR50 in 2007 is associated with 1.77–1.89 pp lower annualized employment growth between 2007 and 2016 (significant at 1%); robust to controlling for pre-crisis employment trends and sector fixed effects. (2) Payroll: similarly negative coefficient of approximately −0.041 on log payroll growth. (3) Net firm entry: a 1 pp higher concentration is associated with 2–3 pp lower post-crisis net entry. (4) Labor share: a negative relationship between 2007 concentration and the change in industry labor share between 2008 and 2016 (coefficient approximately −0.031, significant at 10%). All results are mirrored qualitatively and quantitatively in simulated cross-industry regressions from the model: concentrated markets in the model experience 5.4% larger drops in employment, 3.7% higher firm exit, and 1.1% larger decline in labor share.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-and-extensions-are-reported"&gt;Q7. What robustness checks and extensions are reported?&lt;/h3&gt;
&lt;p&gt;Several extensions and checks are noted: (1) An alternative shock — fluctuations in the fraction of industries with positive fixed costs (xc) rather than TFP shocks — also replicates the medium-run behavior of the US economy, with output falling roughly 15% on impact and remaining −18% below trend in the long run; the cross-sectional implications are unchanged. (2) The 1990 recession counterfactual: applying 1990–1991 recession shocks to the 1990 economy produces no permanent transition, but the same shocks applied to the 2007 economy do, confirming that fragility rather than shock size drove the 2008 outcome. (3) Factor-price-dependent fixed costs: Ferrari and Queirós (2022) show steady-state multiplicity is preserved when fixed costs depend on factor prices. (4) Varying M: results are unchanged for M = 50 and M = 100 potential firms per market. (5) The cross-industry regressions are robust across multiple specifications including controls for the number of firms in 2007, pre-crisis growth, and sector fixed effects (Appendix B.7).&lt;/p&gt;
&lt;h3 id="q8-what-are-the-models-aggregate-predictions-for-labor-share-profit-share-and-markups-post-2008-and-how-do-they-compare-to-data"&gt;Q8. What are the model&amp;rsquo;s aggregate predictions for labor share, profit share, and markups post-2008, and how do they compare to data?&lt;/h3&gt;
&lt;p&gt;Between 2007 and 2016, the model predicts (Table 9): a 0.4 pp decline in the aggregate labor share (data: −2.9 pp decline; the model explains approximately 14% of the total decline, or 17% accounting for the pre-crisis trend); a 0.9 pp increase in the profit share (data: +3.2 pp; model explains 30% of the trend deviation); a 3.7 point increase in sales-weighted markups for COMPUSTAT firms (data: +14.2 points; model explains 26% of the total increase and 58% of the deviation from the pre-crisis trend). The model also predicts a persistent fall in the number of firms in markets with positive fixed costs of 13.4 log points, compared to the observed 15.1 log point decline in the number of US firms with at least one employee. The model understates the magnitude of all these changes, but correctly signs and persists them, consistent with its role in providing a partial explanation.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper studies two interventions: (1) An entry subsidy covering a fraction τf of fixed costs for markets with c &amp;gt; 0 (roughly 11.8% of all markets). A 5% entry subsidy is sufficient to eliminate the welfare costs associated with multiplicity in the 2007 economy; higher subsidies improve allocation within the high steady-state. An entry subsidy large enough to prevent downward transitions yields approximately 10% welfare gain in consumption-equivalent terms. The effect is highly targeted and quantitatively modest per-dollar because only 11.8% of markets are affected. (2) A revenue subsidy τR applied to all firms, equivalent to a fraction of revenues subsidized. Even a 5% revenue subsidy generates approximately 20% welfare gain in the 2007 economy by shifting probability mass from the low to the high competition regime. A 20% revenue subsidy yields gains between 30% and 50% in the 1990 and 2007 economies. The gains are nonlinear in the economies with multiple steady-states, and much smaller in the 1975 economy, which has only one steady-state. A revenue tax has asymmetric large welfare costs in the 1990 economy (which has large output gaps between regimes) relative to the 2007 economy (smaller gap but higher transition probability). The welfare gains come from two sources: reducing static markup distortions and reducing the dynamic cost of transitions (quasi-permanent slumps).&lt;/p&gt;
&lt;h3 id="q10-what-caveats-and-limitations-does-the-paper-acknowledge"&gt;Q10. What caveats and limitations does the paper acknowledge?&lt;/h3&gt;
&lt;p&gt;The authors are explicit about several limitations. First, the model lacks sunk entry costs: all entry decisions are static, which may understate hysteresis and overstate the responsiveness of exit to shocks. Introducing sunk costs with oligopolistic competition poses a computational challenge (20^10 partial equilibria for M=20 and 10 values per firm). Second, idiosyncratic productivities are time-invariant, ruling out Schumpeterian creative destruction within the model. Third, the model features only one-sided market power (product markets only); recent work on labor-market oligopsony could interact with the mechanism. Fourth, the model has no monetary policy channel; the interaction between monetary policy and endogenous market structure is left for future research. Fifth, the model explains only a fraction of the observed post-2008 declines in the labor share (14–17%), profit share (30%), and markup levels (26% of total, 58% of trend deviation), suggesting complementary mechanisms are at work.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-characterize-the-relationship-between-the-great-moderation-and-rising-fragility"&gt;Q11. How does the paper characterize the relationship between the Great Moderation and rising fragility?&lt;/h3&gt;
&lt;p&gt;The paper directly addresses the apparent tension between the Great Moderation (declining aggregate output volatility from 1980 to 2007) and the model&amp;rsquo;s prediction of rising fragility over the same period. The resolution is that aggregate output volatility is the product of exogenous TFP shock volatility and endogenous amplification. If exogenous TFP shocks became less volatile over time (a plausible claim, attributed to demographic shifts and the rising share of low-volatility service industries), then aggregate volatility could have declined even as endogenous amplification increased. Fragility, as defined in the paper, is about the probability of large discrete transitions, not about the variance of the ergodic distribution around a single steady-state. An economy can exhibit lower volatility on average while being more prone to catastrophic (quasi-permanent) downturns.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Macroeconomic Fragility&lt;/strong&gt;: The probability of long slumps, formally measured as the proximity of the high stable steady-state to the preceding unstable steady-state (χ = KU/K*). A higher χ means a smaller negative shock is sufficient to trigger a permanent downward transition. Fragility is distinct from steady-state multiplicity (which is necessary but not sufficient) and distinct from stability (which measures the full basin of attraction in both directions).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Competition-Factor Supply Complementarity&lt;/strong&gt;: The positive feedback loop through which more competitive product markets generate higher factor shares and factor prices, inducing higher labor and capital supply, which in turn allows more firms to enter and compete. This complementarity is the structural foundation for multiple steady-states in the model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mean-Preserving Spread (MPS) on Idiosyncratic TFP&lt;/strong&gt;: An increase in cross-firm productivity dispersion that leaves the average unchanged. In the model&amp;rsquo;s context, an MPS raises aggregate TFP (allocative efficiency effect as output shifts to high-productivity firms) but lowers the factor share and factor price index (market power effect as concentration increases), and shrinks the stable steady-state&amp;rsquo;s capital level while raising the unstable steady-state&amp;rsquo;s capital level — thereby increasing fragility.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Low Competition Trap&lt;/strong&gt;: The low stable steady-state in which the economy becomes trapped following a transition from the high steady-state. Characterized by fewer active firms, higher markups, lower factor shares, lower capital stock, and lower output relative to the high steady-state. In the 2007 calibration, the two steady-states are approximately 21% apart in output terms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous Market Structure&lt;/strong&gt;: The model feature whereby the number of active firms in each product market is determined endogenously by a free-entry condition: the marginal firm exactly breaks even (net profits equal fixed costs). This makes the number of firms — and hence the degree of competition, markups, and factor shares — respond endogenously to aggregate shocks and capital accumulation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Factor Price Index (Θ)&lt;/strong&gt;: A composite of the wage and rental rate representing the minimum cost of one unit of output for a firm with unit productivity. In the model, Θ equals the product of the aggregate factor share and aggregate TFP. It serves as a sufficient statistic for both factor prices and the competitive environment, decreasing with higher firm heterogeneity (via lower factor shares) and increasing with more firms (via higher competition).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Great Deviation&lt;/strong&gt;: The paper&amp;rsquo;s term (following Hall, 2011) for the persistent and widening gap between actual US output and its pre-2007 trend following the 2008–09 recession. In the data, real GDP per capita was 14.2% below its pre-crisis trend as of 2019Q1, a deviation far larger and more persistent than in any prior postwar recession. The paper&amp;rsquo;s model rationalizes this as a transition to the low steady-state.&lt;/p&gt;</description></item><item><title>From Population Growth to TFP Growth</title><link>https://macropaperwarehouse.com/papers/from-population-growth-to-tfp-growth/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/from-population-growth-to-tfp-growth/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks how the well-documented slowdown in labor-force growth affects aggregate total factor productivity (TFP) growth, a question that prior work on business dynamism had left unanswered. The authors build a general-equilibrium business-dynamics model that embeds two engines of productivity growth: innovation by young entrants (a step-size improvement over the leading-productivity frontier, in the spirit of Romer 1990 and Aghion-Howitt 1992) and steady productivity growth by mature leading businesses. Population (labor-force) growth determines the demographic composition of the business stock, because the number of firms must grow in proportion to the labor force along any balanced growth path (BGP). A slower labor force therefore shifts the firm distribution toward older incumbents.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central theoretical result is a &amp;ldquo;sufficient statistic&amp;rdquo; for whether slower population growth reduces TFP growth: the employment-size growth rate of surviving old businesses, which converges to the ratio gS/gX (the productivity growth of leading businesses divided by average economy-wide productivity growth). If gS/gX &amp;lt; 1 — i.e., old firms&amp;rsquo; productivity grows more slowly than the economy average — then a lower labor-force growth rate raises the share of old firms and drags down aggregate productivity growth. Both the sign and the magnitude of the effect are characterized in closed form.&lt;/p&gt;
&lt;p&gt;The model is calibrated to U.S. and Japanese establishment data (Business Dynamics Statistics; Economic Census and Establishment/Enterprise Census), targeting the life-cycle profiles of exit rates, average employment size by age, and the employment growth rate of surviving businesses, with the reference period 1980–1999. The U.S. labor-force growth rate used in calibration is 1.67 percent per year (average 1980–1999); Japan&amp;rsquo;s is 0.72 percent. A key calibrated quantity is gS: 1.060 for the U.S. and 1.030 for Japan, reflecting the faster decline in the size of surviving old establishments in Japan relative to the U.S. The benchmark model adds entry congestion (parameter ϕ = 0.55, taken from Karahan, Pugsley and Sahin 2024) and spillovers from young to old firms&amp;rsquo; productivity growth (γ = 0.342, estimated from BDS data using venture capital investment as an IV).&lt;/p&gt;
&lt;p&gt;Main quantitative findings across BGPs: In the U.S., the projected decline in labor-force growth from approximately 2.59 percent (1970–1980) to 0.26 percent (2050–2060) implies a long-run reduction in TFP growth of approximately 0.3 percentage points. In Japan, the decline from approximately 1.86 percent (1950–1960) to −0.97 percent (2050–2060) — a drop of more than 3 percentage points — implies a long-run reduction in TFP growth of approximately 0.6 percentage points. These effects are substantially attenuated when congestion and spillovers are removed: the U.S. effect falls from 0.30 to 0.19 percentage points and the Japan effect falls from 0.63 to 0.41 percentage points in the simplest model, so roughly 65 percent of the benchmark effect is attributable to the core mechanism alone.&lt;/p&gt;
&lt;p&gt;For the transition analysis, the model accounts for approximately 49.7 percent of the observed U.S. TFP growth slowdown between 1980–1999 and 2000–2019 (an observed decline of 0.184 percentage points, model-explained 0.091 percentage points). In Japan, the model explains approximately 24.2 percent of a larger observed slowdown of 0.451 percentage points (model: 0.109 pp). A critical feature of the dynamics is that TFP growth responds sluggishly to population growth changes. Two transitional counterbalancing forces explain this: (1) a &amp;ldquo;level-vs-growth&amp;rdquo; effect — on impact, a higher share of older (larger and more productive) firms temporarily raises productivity growth in levels even while it lowers the growth rate in the long run; and (2) a &amp;ldquo;labor-reallocation&amp;rdquo; effect — fewer entrants means less labor in the innovation sector and more in production, temporarily raising the production-sector labor share and boosting measured TFP growth. Both effects fade as the economy converges to the new BGP.&lt;/p&gt;
&lt;p&gt;Looking forward, the expected further decline in TFP growth from population aging is -0.05 to -0.06 percentage points for the U.S. between 2020 and 2100 (benchmark, without incorporating forecasts), and -0.14 to -0.17 percentage points for Japan over the same horizon. When BLS/CAO forecasts for labor-force growth through 2060 are incorporated, these magnitudes rise to -0.07 to -0.08 pp (U.S.) and -0.24 to -0.34 pp (Japan) between 2020 and 2100. Cross-sectional IV regressions using lagged state birth rates as instruments confirm that a 1-percentage-point change in labor-force growth maps to approximately a 0.1 to 0.2 percentage-point change in labor productivity growth across U.S. states, consistent with model predictions. Local projections using U.S. state data 1977–2019 show that the dynamic pattern in data (initial positive then negative response of productivity growth to a labor-force shock) mirrors the model&amp;rsquo;s transitional dynamics closely.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-core-theoretical-result-and-what-is-the-sufficient-statistic"&gt;Q1. What is the paper&amp;rsquo;s core theoretical result, and what is the &amp;lsquo;sufficient statistic&amp;rsquo;?&lt;/h3&gt;
&lt;p&gt;The main result (Lemma 4) states that if the employment-size growth rate of surviving old businesses is negative — equivalently, if gS/gX &amp;lt; 1 — then an increase in the labor-force growth rate raises average productivity growth, and vice versa. The &amp;lsquo;sufficient statistic&amp;rsquo; is gS/gX, the ratio of old-firm productivity growth to economy-wide average productivity growth. This ratio asymptotically equals the employment growth rate of surviving old firms in a BGP (Lemma 3). Lemma 5 further shows that the magnitude of the effect is increasing in how fast old firms&amp;rsquo; size shrinks, i.e., larger when gS/gX is further below 1. This means the calibration of the life-cycle profile of surviving business growth is the decisive input for the quantitative results.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-growth-engines-in-the-model-and-how-do-they-interact"&gt;Q2. What are the two growth engines in the model and how do they interact?&lt;/h3&gt;
&lt;p&gt;The first engine is innovation by new entrants: innovators choose a step size g relative to the average leading-firm productivity frontier χ, paying convex research costs. The free-entry condition ties the step size to structural parameters (research cost slope and entry cost), making g* constant in equilibrium. The second engine is the exogenous (in the benchmark) or endogenous (in extensions) productivity growth of leading businesses at rate gS per period. Both engines operate simultaneously: gX is determined by a weighted average of these two sources, where the weight on the old-firm engine equals their share in the firm distribution. Population growth affects this weight by determining the number of new entrants relative to incumbents.&lt;/p&gt;
&lt;h3 id="q3-what-identification-strategy-is-used-in-the-empirical-validation-and-what-are-the-threats"&gt;Q3. What identification strategy is used in the empirical validation and what are the threats?&lt;/h3&gt;
&lt;p&gt;Two empirical strategies are used. First, local projections (Jordà 2005) using U.S. state-level data 1977–2019 regress the change in labor productivity growth over horizons i = 0 to 8 years on the change in labor-force growth, controlling for seven lags of each variable and a quadratic time polynomial. This establishes that the dynamic pattern in the data mirrors the model-predicted non-monotonic response (initial positive effect, then negative and significant effects at 2–5 years). Second, cross-sectional IV regressions for U.S. states average 2004–2024 data and use the lagged state birth rate (pushed back 20 years) as an instrument for labor-force growth, with controls for initial GDP per capita and state population. The main threat is reverse causality: workers may relocate to states with higher expected productivity growth. The authors note the IV addresses this by using birth rates from 20 years prior. A further threat acknowledged is knowledge spillovers across states, which would bias the local-projection coefficient downward.&lt;/p&gt;
&lt;h3 id="q4-what-does-the-paper-say-about-the-role-of-entry-congestion-and-innovation-spillovers"&gt;Q4. What does the paper say about the role of entry congestion and innovation spillovers?&lt;/h3&gt;
&lt;p&gt;Entry congestion modifies the free-entry condition to make entry costs rise with the ratio of entrants to population (with elasticity ϕ = 0.55). This means that when population growth slows and fewer entrants arrive, entry costs fall, which discourages innovation intensity (lower g*), adding a second channel through which slower population growth lowers TFP growth. Innovation spillovers allow the productivity growth of leading businesses (gS) to respond positively to lagged aggregate productivity growth (with elasticity γ = 0.342, estimated via IV). When population growth slows and productivity growth falls, spillovers to incumbents also fall, amplifying the total effect. Together, these features explain roughly 35 percent of the benchmark effect beyond what the core mechanism delivers alone: the U.S. effect rises from 0.19 pp (no congestion, no spillovers) to 0.30 pp in the benchmark.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-robustness-checks-on-the-bgp-results"&gt;Q5. What are the robustness checks on the BGP results?&lt;/h3&gt;
&lt;p&gt;Five alternative productivity processes are considered. Case 1 is a standard two-state AR(1), Case 2 allows transition probabilities to depend on age, Case 3 uses deterministic productivity growth by type (high and low) with age-dependent transitions, Case 4 is the benchmark (asymmetric absorbing high-productivity state with tenure-dependent productivity history), and Case 5 cuts the productivity jump θ in half. All five deliver similar qualitative results, with the long-run U.S. effect ranging from -0.15 to -0.22 percentage points compared to -0.19 in the benchmark. The AR(1) specification (Case 1) yields the smallest effect because it misses the growth of young and old businesses in the data. Endogenous exit is examined in a separate extension: the exit rate declines further when population growth falls (amplifying the old-firm share effect), but this is nearly exactly offset by higher innovation incentives from longer business horizons, resulting in very small net change. Endogenous innovation by leading businesses is also explored and found to amplify the result at low population growth rates (making the effect nonlinear and potentially larger in future decades), but its impact at observed historical ranges is modest.&lt;/p&gt;
&lt;h3 id="q6-how-do-the-transitional-dynamics-differ-from-the-bgp-comparison-and-why"&gt;Q6. How do the transitional dynamics differ from the BGP comparison, and why?&lt;/h3&gt;
&lt;p&gt;The BGP comparison provides the long-run effect of a permanently different population growth rate on TFP growth. The transition shows that convergence to this new BGP is very slow — taking more than 20 years to reach the new steady-state share of young businesses after a step decline in population growth. This slowness is driven by two counterbalancing forces. The level-vs-growth effect: on impact, a lower entry rate raises the share of larger, more productive older firms, which temporarily boosts the level of productivity growth even as the long-run growth rate falls (because young firms have lower productivity levels despite faster productivity growth). The labor-reallocation effect: fewer entrants mean less labor in the innovation sector, reallocating workers to production, which temporarily raises the production-employment share and therefore measured TFP growth. As a result, the model accounts for 49.7 percent of the U.S. TFP growth slowdown between 1980–1999 and 2000–2019, not the full long-run 0.30 pp effect. The sensitivity analysis shows that lower sS, lower β, or higher gS all speed up convergence.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-karahan-pugsley-and-sahin-2024-and-hopenhayn-neira-and-singhania-2022"&gt;Q7. How does this paper relate to Karahan, Pugsley and Sahin (2024) and Hopenhayn, Neira and Singhania (2022)?&lt;/h3&gt;
&lt;p&gt;Both prior papers show that slower labor-force growth reduces business dynamism by generating a startup deficit and shifting the firm age distribution toward older incumbents. They share the basic Hopenhayn (1992) firm-dynamics structure with this paper. The key distinction is that those papers focus on entry rates, exit rates, employment concentration, and labor market dynamics as outcomes, whereas Inokuma and Sanchez focus on TFP growth. As a validation exercise, this paper shows its model also reproduces the decline in U.S. business dynamism (entry rate, exit rate, share of young establishments) when fed the trend in labor-force growth.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-peters-and-walsh-2022"&gt;Q8. How does this paper relate to Peters and Walsh (2022)?&lt;/h3&gt;
&lt;p&gt;Peters and Walsh (2022) also studies population growth and productivity. Their framework builds on Klette and Kortum (2004) and emphasizes scale effects, variety expansion, market concentration, and markups, abstracting from firm life-cycle dynamics. This paper instead builds on Hopenhayn (1992) and focuses on how innovation intensity varies with firm age. The two mechanisms are complementary: the life-cycle mechanism in this paper would add 56 percent to the productivity growth decline found in Peters and Walsh (Peters and Walsh find approximately 0.23 pp per 1 pp decline in population growth, almost all from varieties; Inokuma and Sanchez find 0.13 pp per 1 pp for the U.S., so the combined effect would be roughly 0.36 pp).&lt;/p&gt;
&lt;h3 id="q9-what-heterogeneity-is-documented-in-the-paper"&gt;Q9. What heterogeneity is documented in the paper?&lt;/h3&gt;
&lt;p&gt;The most important heterogeneity is between the U.S. and Japan. Japan&amp;rsquo;s establishments exhibit a much flatter size profile by age (the ratio of employment in establishments 29+ years to age-1 establishments is 1.5 in Japan versus 3.5 in the U.S.) and a sharper decline in the size of surviving old establishments, yielding a calibrated gS of 1.030 for Japan versus 1.060 for the U.S. This implies a larger sufficient statistic |1 - gS/gX| for Japan and therefore a larger elasticity of TFP growth to population growth: 0.6 pp effect for Japan versus 0.3 pp for the U.S. over their respective projected population growth declines. Within the model, the two types of firms (laggard and leading) have different survival rates (sS &amp;gt; sU), different productivity levels (leading firms are roughly 200 vs 10 employees on average), and different exit dynamics (laggards face much higher exit rates, especially when young).&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper does not focus on policy prescriptions, but the implied lesson is that policies affecting the entry rate of new firms — or the productivity life-cycle of mature incumbents — are the primary levers for mitigating the TFP drag from aging populations. Because the effect operates through firm-age composition, any policy that encourages new business formation (lowering entry costs, relaxing congestion) would partially offset the demographic headwind. The scope conditions are important: the main result holds under a perfectly elastic supply of new businesses, constant entrant innovation intensity, and exogenous survival/productivity profiles. Congestion and spillovers amplify the mechanism. When exit is endogenous, competing forces nearly cancel, so the result is robust. The direction of the effect depends critically on gS &amp;lt; gX (i.e., old firms&amp;rsquo; productivity growing more slowly than average), which is empirically verified for both the U.S. and Japan. If the sufficient statistic were positive (gS &amp;gt; gX), slower population growth would raise TFP growth.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-paper-say-about-scale-effects-and-how-they-interact-with-the-life-cycle-mechanism"&gt;Q11. What does the paper say about scale effects and how they interact with the life-cycle mechanism?&lt;/h3&gt;
&lt;p&gt;In a CES variety model (as in Peters and Walsh 2022), gTFP = g_tilde_X + (1/(sigma-1)) * gN, adding a direct scale effect where slower population growth reduces the number of varieties and TFP directly. Calibrating sigma = 4 (consistent with Jones 2022), this implies a 0.33 pp TFP decline per 1 pp population growth decline from the variety channel. The life-cycle mechanism in this paper adds 0.13 pp for the U.S. and 0.22 pp for Japan per 1 pp decline. Thus the two mechanisms together would imply a 0.46 to 0.55 pp decline per 1 pp of population growth slowdown — 30 to 60 percent larger than the variety channel alone.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-level-vs-growth-effect-and-how-does-it-arise"&gt;Q12. What is the &amp;rsquo;level-vs-growth&amp;rsquo; effect and how does it arise?&lt;/h3&gt;
&lt;p&gt;When population growth slows suddenly, the entry rate falls and fewer young firms enter. This means the firm pool immediately becomes more skewed toward older, larger, more productive incumbents. On impact, this raises the average level of productivity in the economy (because old firms have higher levels, even if slower growth rates). This temporarily boosts the growth rate of average productivity in the short run, even though in the long run the effect is to lower TFP growth (because old firms&amp;rsquo; productivity growth rate gS is below gX). This transient positive effect on TFP growth counterbalances and delays the long-run decline, contributing to the sluggish response.&lt;/p&gt;
&lt;h3 id="q13-what-role-does-the-discount-factor-and-household-preferences-play-in-the-results"&gt;Q13. What role does the discount factor and household preferences play in the results?&lt;/h3&gt;
&lt;p&gt;The household problem involves standard intertemporal optimization with risk aversion ε = 2 and discount factor β = 0.96. These parameters enter the speed of convergence in the transition: lower β increases the speed of convergence (sensitivity analysis shows β has an elasticity of -4.212 for convergence speed). Along the BGP, household preferences determine the interest rate through the Euler equation and affect the capital share α-tilde, which varies across BGPs. The paper notes that d(alpha-tilde)/d(gM) is likely negative, meaning that lower population growth also reduces the capital share, amplifying the effect on TFP growth, though extreme parameter values could reverse this.&lt;/p&gt;
&lt;h3 id="q14-what-are-the-data-sources-and-what-moments-are-targeted-in-calibration"&gt;Q14. What are the data sources and what moments are targeted in calibration?&lt;/h3&gt;
&lt;p&gt;For the U.S.: establishment-level data from the Business Dynamics Statistics (BDS), spanning 1978 onwards; labor force data from BLS Current Population Survey (1949–2019) and Lebergott (1966) for 1900–1948; TFP from Penn World Table 10.0; venture capital investment from PwC/CB Insights MoneyTree. For Japan: establishment data from the Establishment and Enterprise Census (1981–2006) and Economic Census (2009–2021); labor force from Statistics Bureau of Japan; TFP from PWT 10.0. Calibration targets 32 moments for the U.S. (31 life-cycle bars plus average productivity growth) and 20 for Japan. The targeted moments are the exit rate by establishment age (with equal weighting), the average employment size profile by age, and the growth rate of surviving establishments by age. Ten parameters are jointly estimated to minimize the distance between model-implied and data moments.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Sufficient statistic (gS/gX)&lt;/strong&gt;: The employment-size growth rate of surviving old businesses, which asymptotically equals the ratio of old-firm productivity growth (gS) to economy-wide average productivity growth (gX). This single ratio determines both the sign (if less than 1, slower population growth reduces TFP growth) and the magnitude (the faster gS/gX falls below 1, the larger the effect) of population growth&amp;rsquo;s impact on productivity growth along balanced growth paths.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Leading versus laggard businesses&lt;/strong&gt;: The paper&amp;rsquo;s two-type firm classification. Laggard businesses start with productivity θ·χ·g (below the frontier), grow at a flat rate, and face high exit rates; they can transition to the leading group with age-dependent probability λ_a. Leading businesses begin at or above the frontier (productivity χ·g at entry), grow at constant rate gS per period, and face lower exit rates. The share of leading versus laggard firms — and the speed at which laggards transition — determines the life-cycle productivity profile that is central to the sufficient statistic.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Level-vs-growth effect&lt;/strong&gt;: A transitional counterbalancing force: when population growth slows, fewer young (small, low-productivity-level) firms enter, immediately raising the average level of productivity in the firm pool and temporarily boosting measured productivity growth, even though the long-run effect is negative. The short-run level gain outweighs the long-run growth-rate loss, delaying the TFP growth decline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor-reallocation effect&lt;/strong&gt;: A second transitional counterbalancing force: lower entry rates reduce the number of workers employed in innovation (research and development) activities, reallocating them to goods production. This increase in the production-sector labor share temporarily raises measured TFP growth. Like the level-vs-growth effect, it fades as the economy converges to the new balanced growth path.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Entry congestion&lt;/strong&gt;: An extension to the free-entry condition in which the per-entrant cost rises with the ratio of the entry rate to population growth (with elasticity ϕ = 0.55). When population growth slows, congestion costs fall, reducing the incentive to invest in high-step-size innovation, thus providing a second channel through which slower population growth reduces TFP growth beyond the core composition channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Innovation spillovers&lt;/strong&gt;: A mechanism by which the productivity growth of already-leading businesses (gS) responds positively to lagged aggregate productivity growth gX (with estimated elasticity γ = 0.342). This link means that when population growth slows and gX falls, mature firms also grow more slowly, amplifying the initial effect. Calibrated using OLS and IV (venture capital investment as instrument) regressions of old-establishment productivity growth on aggregate past productivity growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Balanced growth path (BGP) comparison&lt;/strong&gt;: The primary analytical exercise: comparing steady-state TFP growth rates across economies that differ only in their constant labor-force growth rate. This isolates the long-run equilibrium effect, abstracting from the transitional dynamics that counteract the decline in the short run. The BGP effect is larger than what is observed during any historical transition window because of the slow convergence.&lt;/p&gt;</description></item><item><title>Global Value Chains and Labor Standards: The Race-to-the-Bottom Problem</title><link>https://macropaperwarehouse.com/papers/global-value-chains-and-labor-standards-the-race-to-the-bottom-problem/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/global-value-chains-and-labor-standards-the-race-to-the-bottom-problem/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Im and McLaren (2025) ask whether globalization induces governments to weaken labor standards for workers — the so-called &amp;ldquo;race to the bottom&amp;rdquo; (RTB) hypothesis. The question has high stakes: advocates point to events such as the 1,136-worker Rana Plaza factory collapse in Bangladesh (2013) and to India&amp;rsquo;s deregulation campaign after 2014 (associated with approximately 6,500 workplace deaths in 2015–2020) as evidence that competition for global capital systematically erodes safety and working conditions. The paper builds a stylized many-country equilibrium model of labor-market integration adapted from the Grossman and Rossi-Hansberg (2008) tasks framework. Output requires a continuum of tasks z in [0,1], performable in any of N countries; labor requirements per task follow a Weibull distribution (shape parameter nu &amp;gt; 0), independently across tasks and countries. Working conditions (kappa_i) enter the cost function multiplicatively — better conditions reduce worker productivity at the relevant margin. Utility is separable in wages and conditions with both components strictly concave, and Assumption 1 (x&lt;em&gt;xi&amp;rsquo;(x) and x&lt;/em&gt;mu&amp;rsquo;(x) strictly decreasing) ensures conditions are normal goods and second-order conditions hold. The unregulated equilibrium task allocation is equivalent to CES cost minimization with elasticity of substitution 1/(1-rho) &amp;gt; 1, rho = nu/(1+nu). Governments set minimum standards non-cooperatively in Nash equilibrium.\n\nThe paper&amp;rsquo;s results fall into two conceptually distinct categories. &amp;ldquo;Globalization in the large&amp;rdquo; (autarky vs. open economy): whether standards are market-determined or government-set, integrating two previously autarkic countries raises labor standards in both (Proposition 1). Under autarky, market and government-optimal conditions coincide — all costs of better standards are borne domestically. Under trade, wages rise (income channel: conditions are a normal good), and governments gain a terms-of-trade incentive: tightening kappa_i makes domestic effective labor scarcer and shifts part of the cost onto foreign consumers, inducing government standards to strictly exceed market standards. Formally, for each country i: autarky level = market level under autarky &amp;lt; market level under integration &amp;lt; government level under integration.\n\n&amp;quot;Globalization at the margin&amp;quot; with symmetric countries (Proposition 2): as more identical countries join (N increasing), both market-set and government-set standards rise monotonically. The terms-of-trade motive does not vanish because each country specializes in an increasingly narrow value-chain slice, retaining market power regardless of N. Government standards exceed market standards for every N &amp;gt;= 2 and grow strictly with N — a race to the top — and are shown to be above the social optimum because each country externally imposes part of its improvement costs on others.\n\n&amp;quot;Globalization at the margin&amp;quot; with a North-South structure (Proposition 3): when Southern host countries (i = 2,&amp;hellip;,N) have perfectly correlated productivity draws (close substitutes for one another), the result reverses for N &amp;gt; 2. Integration of two countries initially raises Southern standards via both channels. But as additional similar Southern competitors join, competition depresses Southern wages and erodes both the income-based demand for better conditions and the terms-of-trade motive (unilateral tightening redirects demand to competitors without cost-shifting benefit). Both market and government standards fall monotonically as N rises beyond 2. As N approaches infinity, both converge to autarky levels. Critically, however, for any finite N, Southern standards remain strictly above their autarky levels — the race to the bottom, even when operative, never fully materializes while integration is incomplete. The efficiency implication is counter-intuitive: government-set standards are inefficiently strict under GVCs because each country over-provides standards by externalizing costs onto trading partners.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-models-formal-structure-and-how-does-it-generate-tractable-results"&gt;Q1. What is the model&amp;rsquo;s formal structure and how does it generate tractable results?&lt;/h3&gt;
&lt;p&gt;The model adapts Grossman and Rossi-Hansberg (2008). Output requires a unit measure of tasks; labor requirement for task z in country i is A_i * a^i_z, where A_i = bar_A_i * kappa_i, so working conditions raise unit labor costs. Each a^i_z is drawn Weibull(nu, 1) independently. A result (adapted from Anderson et al. 1987, applied by Artuç and McLaren 2015) is that the cost-minimizing task allocation is equivalent to minimizing cost with a CES aggregate of national effective labor supplies, with elasticity of substitution 1/(1-rho) and rho = nu/(1+nu). This reduces the multi-dimensional problem to a standard CES factor-demand problem, yielding closed-form wage equations and tractable Nash equilibrium characterizations.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-channels-driving-globalization-in-the-large-raising-standards-above-autarky"&gt;Q2. What are the two channels driving &amp;lsquo;globalization in the large&amp;rsquo; raising standards above autarky?&lt;/h3&gt;
&lt;p&gt;Two reinforcing channels. First, the income channel: integration raises real wages (gains from specialization), and since working conditions are a normal good under Assumption 1 (utility sufficiently concave), demand for better conditions rises. Second, the terms-of-trade channel: tightening kappa_i makes domestic effective labor more expensive and scarcer; part of the resulting cost increase is borne by foreign consumers and workers via the unit cost identity rather than solely by domestic workers. This cost-shifting gives governments an incentive to tighten standards beyond what the unregulated market sets. The mechanism is formally analogous to the policy externalities in Bagwell and Staiger (2001) and the terms-of-trade motive in Chau and Kanbur (2006), though the latter has no value chains.&lt;/p&gt;
&lt;h3 id="q3-why-does-the-terms-of-trade-motive-for-over-regulation-persist-even-as-the-number-of-symmetric-countries-approaches-infinity"&gt;Q3. Why does the terms-of-trade motive for over-regulation persist even as the number of symmetric countries approaches infinity?&lt;/h3&gt;
&lt;p&gt;As more countries join, each specializes in an increasingly narrow slice of the value chain in which it has comparative advantage. This deepening specialization preserves market power: the wage derivative dw_1/d_kappa_1 converges to a limit proportional to rho*w/kappa (strictly greater than the pure autarky productivity effect -w/kappa) rather than to zero. So even in the limit with infinitely many symmetric countries, each country retains some terms-of-trade gain from tightening its standard, and government standards keep rising above market standards.&lt;/p&gt;
&lt;h3 id="q4-under-what-precise-conditions-does-the-race-to-the-bottom-result-hold"&gt;Q4. Under what precise conditions does the race-to-the-bottom result hold?&lt;/h3&gt;
&lt;p&gt;The RTB result (Proposition 3) requires that competing host countries be close substitutes for one another. The paper operationalizes this with the extreme case of perfectly correlated productivity draws across Southern countries (a^i_z = a^2_z for all i &amp;gt;= 2 and all tasks z). Under this structure, as N increases from 2 onward, Southern market and government standards fall monotonically toward autarky levels. The mechanism: competition among near-identical countries means unilateral tightening of kappa_2 redirects Northern demand to competitors without generating a terms-of-trade gain for Country 2, so the wage falls and conditions deteriorate. The RTB thus requires high substitutability among competitors, not just trade openness.&lt;/p&gt;
&lt;h3 id="q5-does-the-race-to-the-bottom-ever-drive-standards-below-autarky-levels"&gt;Q5. Does the race to the bottom ever drive standards below autarky levels?&lt;/h3&gt;
&lt;p&gt;No. Proposition 3 parts (i) and (ii) establish that for any finite N &amp;gt;= 2, both market-set and government-set standards in Southern countries remain strictly above their autarky levels. The race is toward (but never below) the autarky benchmark. Only in the limit as N approaches infinity do standards converge to the autarky level (Proposition 3, part iii). For any realistic finite degree of globalization, even the worst-case RTB scenario leaves standards strictly above autarky.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-efficiency-implication-of-nash-equilibrium-government-set-standards"&gt;Q6. What is the efficiency implication of Nash equilibrium government-set standards?&lt;/h3&gt;
&lt;p&gt;Government-set standards under GVCs are inefficiently strict. Each government maximizes domestic welfare ignoring the cost its tightening imposes on foreign consumers and workers. Because tightening kappa_i raises costs partly borne abroad, each government over-provides standards relative to the global social optimum. This is a race to the top that generates a negative international externality — the mirror image of the usual RTB externality. The implication is that international coordination, if it occurred, would likely reduce Nash equilibrium standards toward the optimum, not raise them further.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-papers-setting-differ-from-prior-theoretical-work-on-the-race-to-the-bottom"&gt;Q7. How does the paper&amp;rsquo;s setting differ from prior theoretical work on the race to the bottom?&lt;/h3&gt;
&lt;p&gt;Prior RTB models (Chau and Kanbur 2006; Felbermayr et al. 2012; Chen and Dar-Brodeur 2020) model countries competing for export markets — competing to sell goods to a common importer — rather than competing to host tasks in global value chains. The current paper frames globalization as an increase in the number of countries that can supply tasks to a common production process, a qualitatively different competitive margin. Prior work also largely takes the degree of globalization as fixed, while this paper explicitly traces out effects as N changes. The distinction between similar versus different competitors as a determinant of the direction of the RTB is also new. The companion paper Im and McLaren (NBER WP 31363) extends the framework to collective-bargaining rights with an empirical component.&lt;/p&gt;
&lt;h3 id="q8-what-heterogeneity-is-documented-and-what-does-it-imply"&gt;Q8. What heterogeneity is documented and what does it imply?&lt;/h3&gt;
&lt;p&gt;The paper develops two polar cases of country heterogeneity: (1) symmetric countries with independent productivity draws — produces a race to the top as N rises; (2) North-South structure with correlated (identical) Southern productivity draws — produces a race to the bottom as N rises beyond 2. The contrast is the central result: the direction of the marginal effect of globalization on standards depends on the degree of substitutability among competing host countries. The authors connect this to observed patterns — Korean firms relocating only to East Asian affiliates (similar countries) when domestic minimum wages rose, and Chan and Ross (2003) noting that competition is &amp;lsquo;most vicious not between North and South, but among nations of the South.&amp;rsquo;&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The core implication is that trade restrictions justified by RTB concerns lack general theoretical support — globalization relative to autarky always raises standards. However, the model validates a targeted RTB concern: when a country faces competition from many similar low-wage countries (e.g., Mexico competing with China in labor-intensive sectors), standards can erode relative to the peak reached under limited integration. The appropriate response in that case is to integrate with structurally different partners (as Mexico did via NAFTA with the US) rather than restrict trade. Since Nash equilibrium standards already exceed the global optimum, international agreements that ratchet standards up further could be welfare-reducing. The paper explicitly cautions that causation is hard to establish in the Mexico-China-NAFTA example, treating it as suggestive illustration rather than proof.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-main-limitations-and-threats-to-the-conclusions"&gt;Q10. What are the main limitations and threats to the conclusions?&lt;/h3&gt;
&lt;p&gt;The paper is entirely theoretical; no empirical test is conducted for working conditions (the authors cite data scarcity as the reason, having a companion empirical paper on collective-bargaining rights instead). Key assumptions include: (a) Weibull, independent task-productivity draws (ensure tractability but are untested); (b) working conditions always reduce productivity at the margin (rules out the many cases where safety improvements also raise output — e.g., Alfaro-Ureña et al. 2021 find no productivity effect of responsible sourcing in Costa Rica, suggesting the trade-off assumption is plausible but not universal); (c) citizen activism, which empirically affects labor standards (Harrison and Scorse 2010; Koenig and Poncet 2019, 2022), is abstracted away; (d) the model has a single final good and no intermediate goods trade beyond the task-allocation interpretation, limiting applicability to multi-sector settings.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Labor standards (kappa_i)&lt;/strong&gt;: In the paper&amp;rsquo;s specific sense, the quality of working conditions that (i) raise worker utility holding wages fixed and (ii) increase unit labor costs for employers. Explicitly restricted to improvements that involve a trade-off — e.g., safety provisions, clean bathrooms, break times — excluding complementary improvements that raise both utility and productivity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Globalization in the large&lt;/strong&gt;: The paper&amp;rsquo;s term for the comparison of any open-economy equilibrium (N &amp;gt;= 2 countries integrated) against autarky. Result: labor standards are always strictly higher in the open economy whether market-set or government-set, because income rises and the terms-of-trade motive activates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Globalization at the margin&lt;/strong&gt;: The paper&amp;rsquo;s term for the effect on labor standards of adding one more country to an already-integrated economy (increasing N by 1). This effect is ambiguous: it raises standards when new entrants are dissimilar (symmetric model) and lowers them when new entrants are similar (North-South model).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Terms-of-trade effect (labor-standards channel)&lt;/strong&gt;: The mechanism by which tightening a country&amp;rsquo;s labor standard (raising kappa_i) reduces domestic effective labor supply, raises the relative price of domestic tasks, and shifts part of the cost improvement onto foreign consumers and workers. This creates an incentive for governments to set standards above the market level and above the global social optimum — producing standards that are too strict from an efficiency standpoint.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Normal good (working conditions)&lt;/strong&gt;: The property implied by Assumption 1 (both x&lt;em&gt;xi&amp;rsquo;(x) and x&lt;/em&gt;mu&amp;rsquo;(x) strictly decreasing in x) that workers&amp;rsquo; marginal valuation of working conditions relative to wages is higher at higher income levels. This ensures that any source of income gains — including gains from trade — mechanically raises equilibrium demand for better working conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Race to the top&lt;/strong&gt;: The paper&amp;rsquo;s characterization of the symmetric-countries equilibrium: as N increases, both market-set and government-set labor standards rise monotonically, because market power persists through value-chain specialization and the terms-of-trade motive remains strong. Government standards also exceed the social optimum, making this over-regulation an externality imposed on trading partners.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Race to the bottom (conditional)&lt;/strong&gt;: The result in the North-South model where additional similar Southern host countries erode Southern labor standards as N rises beyond 2. The race is toward autarky levels but never below them for finite N. The RTB requires high substitutability among competing host countries and does not hold as a general consequence of globalization.&lt;/p&gt;</description></item><item><title>Health Sector Structural Change</title><link>https://macropaperwarehouse.com/papers/health-sector-structural-change/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/health-sector-structural-change/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;The paper investigates why the U.S. health-services sector has simultaneously experienced a tripling of relative prices since 1948 and a rise in the personal consumption expenditure (PCE) share from under 5% in the late 1940s to 19.6% by 2022. The authors attribute this structural transformation to three candidate drivers: (1) rising relative health-sector markups, (2) unbalanced technological change (differential TFP growth rates across sectors), and (3) changes to the composition of demand from population aging and improving health-investment efficiency.&lt;/p&gt;
&lt;p&gt;The paper proceeds in two stages. First, a growth-accounting decomposition uses a two-sector Dixit-Stiglitz monopolistically competitive model to identify growth rates of relative markups and health-sector TFP directly from sectoral input/output data (NIPA PCE, BEA Fixed Asset Tables, Penn World Tables). The non-health capital intensity is set at 0.40 (from Horenstein and Santos 2019) and health-sector capital intensity at 0.26 (from Donahoe 2000). Because health is more labor-intensive than non-health (alpha_h &amp;lt; alpha_c), GE effects on input prices actually dampen relative price growth. In the baseline decomposition (1954-2019), average annual relative markup growth is estimated at 1.6%, with cumulative growth of approximately 186%. When allowing for a time-varying non-health labor share, relative markup growth rises to 2.0% annually and 255% cumulatively. Average annual health-sector TFP growth is 0.3% (baseline) and 0.2% (time-varying labor share), compared to 0.7% for the non-health sector per Penn World Tables. If no markup growth is assumed, the implied health-sector TFP growth falls to -1.3% annually, implying a 56.3% cumulative decline from 1954 to 2019, which the authors regard as implausible in light of observed healthcare advances. Across all four decomposition exercises, GE effects consistently dampen rather than amplify relative price growth, indicating that demand-side composition shifts from aging play at most a minor role in driving prices.&lt;/p&gt;
&lt;p&gt;Second, the paper builds and calibrates a full general-equilibrium overlapping-generations model (calibration period 1960-2015, in 5-year intervals) with endogenous survival probabilities following Hall and Jones (2007), monopolistic competition, and a PAYG social security system. The model is calibrated to match five time series: relative health price, life expectancy, health expenditure share, capital share in health production, and labor share in health production. The baseline GE model additionally fits the non-targeted decline in average GDP growth rates well. In the baseline calibration, health-sector TFP is estimated to have grown at 0.3% annually from 1950-1970, accelerating to 0.8% (1975-1980), 1.3% (1985-1995), and 1.5% thereafter — faster than the non-health sector’s 0.7% after the mid-1970s. These GE-corrected estimates exceed those from partial-equilibrium exercises because the growth-accounting approach fails to account for factor-input endogeneity; the true GE path requires health-sector TFP to outpace non-health TFP to reconcile observed relative price growth with the magnitude of markup increases.&lt;/p&gt;
&lt;p&gt;Counterfactual simulations isolate each channel. When only demand effects operate (population growth and health-investment efficiency improvements), relative prices rise by only 6.4% compared to 131% in the predicted baseline, and the health share of expenditure rises by 0.004 percentage points versus 0.171 in the baseline — confirming the minor role of aging and demand-composition change. Rising markups alone reproduce nearly all relative price growth but drive expenditure shares up via price rather than quantity increases. Unbalanced TFP growth (with health-sector TFP growing faster post-1975) contributes to real output expansion in the health sector, partially drives up the expenditure share through quantities, supports GDP growth, and — by raising the real value of health services — sustains life-expectancy gains. By 2050, the baseline calibration projects health-sector markups to be approximately 6 times non-health-sector markups if the estimated 1.7% average annual markup growth continues.&lt;/p&gt;
&lt;p&gt;The policy implication is direct: market concentration — documented by HHI levels exceeding 2,500 in the majority of U.S. metropolitan areas, with 19% of MSAs having a single monopolistic hospital provider in 2017 — is the primary driver of rising relative health prices. Antitrust enforcement and policies encouraging technology adoption would together address price growth without sacrificing the real productivity gains that have driven longevity improvements. However, welfare analysis of such policies requires distinguishing between curbing care-provider market power versus pharmaceutical/equipment-manufacturer market power, the latter involving R&amp;amp;D investment incentives that the current aggregate model cannot disentangle.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-relative-markup-growth-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy for relative markup growth, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Relative markup growth is identified from the growth-accounting expression derived from a two-sector Dixit-Stiglitz model: the growth rate of the health-services share of aggregate consumption can be decomposed into relative markup growth, non-health TFP growth (from Penn World Tables), growth in sectoral capital and labor inputs (from BEA and NIPA), and aggregate consumption growth. Taking the capital intensity of the non-health sector as given (alpha_c = 0.40 from Horenstein and Santos 2019) and the data series as known, relative markup growth is backed out residually without requiring knowledge of health-sector TFP or alpha_h. Key threats: (1) the assumption that wages are equalized across sectors (the paper documents supporting evidence in Supplemental Appendix B.6); (2) the constancy of alpha_c, though a time-varying labor-share extension relaxes this; (3) the Dixit-Stiglitz framework abstracts from market selection and endogenous concentration, so markups are characterized as symmetric representative-firm markups rather than firm-distribution markups; (4) the Horenstein and Santos (2019) alternative markups from Compustat cover only publicly traded firms and may understate aggregate markup growth before the 1980s corporatization wave, biasing downward their markup-growth estimates and biasing upward implied TFP-growth estimates for that period.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The three mechanisms are: (1) rising relative markups (supply-side pricing power), (2) unbalanced TFP growth (sector-differential productivity), and (3) changing demand composition (aging and health-investment efficiency). In the partial-equilibrium growth-accounting stage, the three are separated algebraically in equation (8): relative price growth equals relative markup growth plus a GE-effect term (which captures input-price ratio variation and thus embeds demand composition effects) plus relative TFP variation. In the full GE counterfactual stage, channels are separated by switching them off one at a time (fixing gN = gz = gζj = 0 for demand; fixing µt at its 1955 level for markups; fixing gAc = gAh = 0 for TFP), and by activating only one channel at a time. Table 3 presents the counterfactual outcomes for five targeted moments (relative price growth, life-expectancy change, health expenditure share change, capital and labor input shares) under each scenario.&lt;/p&gt;
&lt;h3 id="q3-what-does-the-paper-find-about-health-sector-tfp-growth-and-how-does-this-revise-the-literature"&gt;Q3. What does the paper find about health-sector TFP growth, and how does this revise the literature?&lt;/h3&gt;
&lt;p&gt;The standard view (Triplett and Bosworth 2004; Bates and Santerre 2013) treats health as a ‘cost-disease’ sector with near-zero or negative TFP growth. The paper challenges this: in the baseline partial-equilibrium decomposition, health-sector TFP grows at 0.3% per year on average (1954-2019), compared to -1.3% per year in a model that ignores markup growth entirely. In the full GE model, health-sector TFP growth is higher still — 0.3% (1950-1970), 0.8% (1975-1980), 1.3% (1985-1995), and 1.5% thereafter — eventually exceeding the non-health sector’s 0.7% annual rate. The authors argue this upward revision is correct: partial-equilibrium exercises omit GE feedback effects through factor-input reallocation, and prior studies that did not account for rising markups mechanically attributed all relative price growth to slow TFP growth, biasing health-sector TFP estimates downward.&lt;/p&gt;
&lt;h3 id="q4-what-role-does-population-aging-and-demand-composition-change-play-and-what-is-the-channel"&gt;Q4. What role does population aging and demand-composition change play, and what is the channel?&lt;/h3&gt;
&lt;p&gt;Demand composition changes (population aging and improvements in health-investment efficiency ztζjt) have only a minor role. In GE, such changes can affect input prices (r/w) and thereby health prices only if the health sector uses a different capital intensity than the non-health sector (alpha_h ≠ alpha_c); the elasticity of relative price with respect to the input-price ratio is (alpha_h - alpha_c), which is negative since health is more labor-intensive. This means demand effects actually dampen rather than amplify relative price growth. In the counterfactual where only demand effects operate, relative prices rise by only 6.4% (versus 131% in the predicted baseline from 1960-2015), and the health expenditure share increases by only 0.004 percentage points (versus 0.171 in the predicted baseline). Demand effects do, however, significantly affect life expectancy: shutting them off while allowing only markups produces declining life expectancy, illustrating that income growth and health-investment efficiency improvements are central to longevity gains.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented"&gt;Q5. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;The model features age heterogeneity in three dimensions: (1) age-specific health elasticity θj (how much health expenditure converts to health status); (2) age-specific health output intensity φj; (3) age-specific health-investment productivity ζjt, borrowed from Hall and Jones (2007). These parameters allow older individuals to have lower elasticities of health status with respect to health expenditure, matching the empirical regularity that older patients benefit less per dollar spent on health care. The paper also documents heterogeneity in the sub-components of the health PCE aggregate: over time, prescription drugs and medical appliances have declined in their relative contribution to aggregate health price increases, while hospital services have increased in their relative contribution, consistent with Cooper et al. (2019) on hospital pricing power. Across calibrations, health-sector TFP growth rates vary across four eras, reflecting the different pace of productivity improvements over time.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The authors conduct four different decomposition exercises in the partial-equilibrium stage: (a) baseline with constant non-health capital intensity; (b) time-varying non-health labor share; (c) using Horenstein and Santos (2019) markups from Compustat for publicly traded firms; (d) zero relative markup growth as an extreme baseline. In Supplemental Appendix C.2 they also invert the identification: they set health-sector TFP growth to values from the literature (-0.6% to 0.4% per year) and back out alpha_h, obtaining values between 0.25 and 0.38, consistent with the externally calibrated 0.26. Five full GE calibrations correspond to the five decomposition assumptions. Model fitness is assessed via RMSE across the five targeted moments; the baseline calibration fits best. An untargeted validity check against observed average GDP growth rates over 5-year intervals further supports the baseline model. Results from alternative calibrations’ counterfactuals are presented in Supplemental Appendix D.6.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The closest antecedents are: (1) Horenstein and Santos (2019), who attribute rising U.S. relative health prices to markups and price wedges using Compustat data; the present paper both uses their results and critiques them for under-coverage of non-publicly-traded firms. (2) Hall and Jones (2007), who model health investment, endogenous survival, and the demand side; the present paper embeds their survival technology into a two-sector GE model and adds the supply-side markup and TFP structure. (3) Fonseca et al. (2021, 2023), who account for the rise in health expenditure and cross-country health price differences; the present paper complements them by jointly modeling prices and quantities in a structural change framework. (4) Zhao (2014), who asks why health expenditure shares have risen from a demand side; this paper explores the supply-side (markup and TFP) counterpart. (5) Cost-disease literature (Baumol 1967; Triplett and Bosworth 2004): the paper directly challenges the ‘cost disease’ narrative by showing health-sector TFP is positive and — once GE and markup effects are controlled for — possibly faster than the rest of the economy. Distinctive contributions include the joint treatment of relative prices and real output quantities in structural change, the full GE calibration with endogenous population aging, and the explicit separation of health-care-quantity TFP from health-investment efficiency (the ztζjt composite).&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary policy implication is that antitrust enforcement targeting market concentration in health services is the most direct lever for reducing relative price growth, since markup growth is almost entirely responsible for rising relative prices. A secondary policy recommendation is to encourage technology adoption in the health sector to sustain the high TFP growth that has benefited consumers through output expansion and life-expectancy improvements. The authors caution, however, that the model uses a broad definition of the health-services sector (encompassing care providers, pharmaceutical companies, and equipment manufacturers), and welfare implications differ sharply depending on whether policies target care-provider pricing power versus pharmaceutical/equipment pricing power, the latter involving R&amp;amp;D investment incentives. The model cannot disaggregate the sources of health-sector productivity growth, so the precise antitrust strategy requires further research. Additionally, the paper abstracts from 2020 short-term fluctuations and focuses on long-run structural change, so findings are most relevant for secular policy rather than cyclical interventions.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-role-of-unbalanced-tfp-growth-for-gdp-and-life-expectancy"&gt;Q9. What is the role of unbalanced TFP growth for GDP and life expectancy?&lt;/h3&gt;
&lt;p&gt;Counterfactual simulations reveal that unbalanced TFP growth — which in the baseline calibration favors the health sector after the mid-1970s — supports aggregate GDP growth. In the counterfactual where TFP growth is turned off (both sectors), GDP grows more slowly because the main remaining driver of income growth is exogenous population growth. The panel (f) of Figure 7 shows that GDP growth is slower without unbalanced TFP variation. For life expectancy, the absence of TFP growth causes life expectancy to rise until the 1980s then stagnate (purple line, panel (b) of Figure 7), since rising income is needed to purchase longevity gains through health investment. The interaction between income growth from TFP and the endogenous demand for health investment is central: health services function as a luxury good in the model, so income growth drives up the quantity demanded and thus survival rates.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-paper-find-about-the-current-level-and-trajectory-of-relative-markups"&gt;Q10. What does the paper find about the current level and trajectory of relative markups?&lt;/h3&gt;
&lt;p&gt;In the baseline calibration, health-sector markups were approximately 1.18 times non-health-sector markups in 1955. By 2010, this ratio had risen to approximately 3.2. The time-varying labor-share model implies even faster growth, from 1.09 in 1955 to 3.9 by 2010. Horenstein and Santos (2019) markups (slowest) go from 1.10 in 1955 to 3.04 in 2010, still a 176% increase. Under the baseline calibration projecting continued markup growth at 1.7% annually, health-sector markups would reach approximately 6 times non-health-sector markups by 2050. These projections are corroborated by micro evidence: HHI for managed care exceeds 2,500 in all California counties (Tawil and DiGiorgio 2022); national MSA-level hospital-bed HHI rose from 5,426 in 2007 to 5,808 in 2017; and 19% of MSAs had a single monopolistic provider in 2017 (Johnson and Frakt 2020).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Relative markup&lt;/strong&gt;: The ratio of the health-sector markup (price over marginal cost in a Dixit-Stiglitz monopolistically competitive equilibrium) to the non-health-sector markup; variation in this ratio is identified from sectoral input/output data and is almost entirely responsible for rising relative health-services prices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Unbalanced technical change&lt;/strong&gt;: Differential rates of TFP growth across the health and non-health sectors; in models with homothetic preferences and identical factor intensities, relative prices move inversely with relative TFP, but in the paper’s GE setting with different capital intensities the relationship is modified by GE input-price effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Health-investment efficiency (ztζjt^θj)&lt;/strong&gt;: An age-specific and time-varying composite productivity term governing how effectively a dollar of health-services expenditure (hjt) converts into improved health status and survival probabilities; it captures environmental, behavioral, and knowledge-based factors orthogonal to health-sector TFP (Aht), and is borrowed from Hall and Jones (2007).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GE (general equilibrium) effect&lt;/strong&gt;: In the price-decomposition framework, the term (αh − αc)(gLh,t − gKh,t) capturing how changes in the economy-wide capital-labor ratio — driven by demographic change, markup growth, and TFP changes — feed back into relative sector input prices and thereby into relative health prices; because αh &amp;lt; αc, this effect consistently dampens relative health-price growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cost disease&lt;/strong&gt;: The Baumol (1967) hypothesis that labor-intensive sectors like health services experience slow TFP growth, causing their relative prices to rise as economy-wide wages grow; the paper challenges this characterization by showing health-sector TFP growth is positive and, once GE and markup effects are controlled for, exceeds that of the non-health sector after the mid-1970s.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Corporatization of health services&lt;/strong&gt;: The historical transition of health-services providers from not-for-profit and public-sector organizations to for-profit investor-owned corporations (including private equity-backed systems), which the paper argues has driven the increase in aggregate health-sector markups and whose timing explains why Compustat-based markup estimates from the 1970s understate long-run markup growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous population / survival rate&lt;/strong&gt;: In the model, survival probabilities are functions of individual health-services expenditure (hjt) and health-investment efficiency; this makes population aging partly endogenous to health-sector pricing and productivity, linking structural change in health to aggregate life-expectancy dynamics and GDP growth within a unified OLG framework.&lt;/p&gt;</description></item><item><title>How Costly Are Cartels?</title><link>https://macropaperwarehouse.com/papers/how-costly-are-cartels/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/how-costly-are-cartels/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Moreau and Panon ask how much cartels cost the aggregate economy — in terms of both total factor productivity and welfare — and find the losses are considerably larger than the received wisdom from Harberger (1954) would suggest. The paper&amp;rsquo;s motivation is the mounting evidence that markups are large and growing, combined with a near-total absence of macroeconomic quantification of collusion as one micro-origin of those markups.&lt;/p&gt;
&lt;p&gt;The empirical foundation is an original firm-level database for France covering the period 1994–2007, assembled by scraping all written decisions of the French Competition Authority (ADLC). The final dataset contains 174 cartels and more than 1,000 firms before matching. These cartel records are merged to administrative balance-sheet and income-statement data covering the universe of French firms (BRN and RSI regimes). Key facts documented: average cartel duration is 4.5 years (median 3 years); average cartel size is 6.3 members (median 4); cartels are prevalent across construction, manufacturing, wholesale, retail, and transportation. Crucially, cartel members are empirically shown to be dramatically larger than non-members even within narrowly defined 4-digit industries — roughly 1,900% more sales, a market share premium of 4 percentage points, 1,150% more employment, and 37% higher labor productivity. Firms within a cartel are also substantially more homogeneous in productivity than the overall within-industry distribution: the interquartile productivity ratio across cartel members is only 1.4-to-1, versus 2-to-1 across all non-cartel firms in the same industry.&lt;/p&gt;
&lt;p&gt;The theoretical framework extends the static heterogeneous-firm oligopoly model of Atkeson and Burstein (2008) by introducing collusion microfounded via the cross-ownership framework of O&amp;rsquo;Brien and Salop (1999). A single collusion-intensity parameter κ ∈ [0,1] governs how much each cartel member internalizes the profits of other members. When κ = 0 the model reduces to competitive Cournot oligopoly; when κ = 1 all cartel members jointly maximize profits. In equilibrium, markups rise with firm market share, generating endogenous markup dispersion. Adding collusion causes cartel members to face a lower effective demand elasticity — their own market share augmented by the weighted market shares of co-conspirators — and to charge supracompetitive markups (overcharges). Critically, the effect of cartels on aggregate productivity is theoretically ambiguous: the output contraction of colluding firms redirects demand toward non-colluding firms. If the cartel is composed of the largest (most productive) firms, demand shifts toward less productive non-members, reducing productivity. If the cartel is composed of the least efficient firms, demand shifts toward large non-members, potentially improving allocation.&lt;/p&gt;
&lt;p&gt;The model is calibrated to match six moments from French data in 2007 — aggregate markup, cartel overcharge, the slope of the inverse-markup-on-HHI regression, the median number of firms per sector, the median number of cartel members, and the distribution of relative sales. The key calibrated parameters are: within-sector elasticity of substitution ρ = 10.19; across-sector elasticity η = 1.86; collusion intensity κ = 0.79. The cartel overcharge target is set to 10%, consistent with the OECD benchmark used by antitrust authorities and with Laborde (2021).&lt;/p&gt;
&lt;p&gt;Main quantitative findings (baseline calibration, cartels composed of top producers):&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Eliminating all cartels raises aggregate TFP by 1.1%.&lt;/li&gt;
&lt;li&gt;The productivity cost of markups with respect to the efficient allocation is 70% higher in the model with collusion (3.67%) than in the calibrated competitive oligopoly (2.16%), because collusion generates additional markup dispersion on top of the dispersion inherent in firm heterogeneity.&lt;/li&gt;
&lt;li&gt;Eliminating cartels brings the economy 30% closer to the efficient allocation.&lt;/li&gt;
&lt;li&gt;The aggregate markup falls by approximately 1.5 percentage points when cartels are eliminated.&lt;/li&gt;
&lt;li&gt;Consumption-equivalent welfare gains from eliminating cartels equal 2%.&lt;/li&gt;
&lt;li&gt;Larger cartels (market share above median) account for roughly 80% of the productivity gains; dismantling only large cartels yields a 0.88% TFP gain and 1.97% consumption-equivalent welfare gain; smaller cartels yield 0.23% TFP and 0.54% welfare.&lt;/li&gt;
&lt;li&gt;Umbrella pricing — non-cartel members raise their markups because the cartel&amp;rsquo;s higher prices provide cover — dampens aggregate gains quantitatively but only slightly: fixing non-members&amp;rsquo; markups yields 1.14% productivity gain versus 1.11% in the benchmark.&lt;/li&gt;
&lt;li&gt;Reducing collusion intensity from κ = 0.79 to κ ≈ 0.4 (roughly a 50% reduction) still generates TFP gains of 0.54% and welfare gains of 0.85%, demonstrating that tougher antitrust enforcement at the intensive margin (forcing cartels to soften, not dissolve) yields substantial gains.&lt;/li&gt;
&lt;li&gt;These estimates are one order of magnitude above Harberger&amp;rsquo;s (1954) 0.1% dead-weight loss estimate; the paper shows this discrepancy arises because Harberger uses sectoral data and near-unit demand elasticities, both of which suppress markup dispersion within sectors.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The paper&amp;rsquo;s scope conditions are explicit: results reflect the static cost of cartels; dynamic effects (entry deterrence, innovation incentives) are acknowledged but not quantified; only domestic, detected cartels are covered, so estimates likely understate the true cost; the channel through geographic markup dispersion is excluded.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-primary-identification-strategy-and-what-are-its-main-limitations"&gt;Q1. What is the paper&amp;rsquo;s primary identification strategy, and what are its main limitations?&lt;/h3&gt;
&lt;p&gt;The paper does not rely on a natural experiment or difference-in-differences design. Instead, it uses a structural calibration approach: a heterogeneous-firm oligopoly model with collusion is calibrated to match French data moments, and the cost of cartels is computed as the difference between the calibrated cartel equilibrium and a counterfactual competitive Nash-Cournot equilibrium. The main threats to this strategy are: (1) the sample of cartels consists only of detected cartels, which may not be representative of the latent population — discovered cartels could be either more or less severe than undiscovered ones; (2) no firm-level price data are available, so markups cannot be estimated directly; (3) the counterfactual is a calibrated competitive model rather than an empirically observed post-cartel state; (4) the model abstracts from entry and exit, which may dampen or amplify the true gains from cartel dissolution.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-through-which-cartels-affect-aggregate-productivity-and-how-are-they-distinguished"&gt;Q2. What are the main mechanisms through which cartels affect aggregate productivity, and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;Two channels operate simultaneously. First, the direct price effect: cartel members raise markups above the competitive level (overcharges), reducing their output. In the presence of markup dispersion, this disproportionately contracts output from high-markup (high-productivity) firms, increasing misallocation. Second, the demand reallocation effect: as cartel members contract output and raise prices, non-cartel members gain market share and increase their markups via the umbrella pricing mechanism. The net effect on productivity depends on which firms gain market share. When cartels consist of top producers, reallocation goes toward less productive non-members, reducing aggregate TFP. When cartels consist of the least efficient firms, reallocation goes toward larger non-members, potentially improving allocation. The two channels are not empirically separated in the data; rather, the model disentangles them analytically and then disciplines the net effect via calibration to observed cartel overcharges.&lt;/p&gt;
&lt;h3 id="q3-why-do-the-authors-assume-cartels-are-composed-of-the-most-productive-firms-and-what-is-the-evidence-for-this"&gt;Q3. Why do the authors assume cartels are composed of the most productive firms, and what is the evidence for this?&lt;/h3&gt;
&lt;p&gt;The assumption is motivated by three pieces of evidence. First, empirical regressions on the matched administrative data show that cartel members within their 4-digit industries have roughly 1,900% more sales, 1,150% more employment, and 37% higher labor productivity than non-members. Second, firms within a cartel are much more homogeneous than the overall within-industry distribution: the interquartile productivity ratio within a cartel is 1.4-to-1, versus approximately 2-to-1 for all non-cartel firms in the same industry, and the 90-10 ratio is 1.7-to-1 within a cartel versus over 4-to-1 across the industry. Third, only the top-producer composition assumption, combined with a collusion intensity κ = 0.79, can generate a cartel overcharge of 10% consistent with the calibration target. All other composition configurations (least efficient, all-inclusive, random top-10%) yield either implausibly small overcharges or implausibly large ones.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-umbrella-pricing-effect-and-how-large-is-it-quantitatively"&gt;Q4. What is the umbrella pricing effect and how large is it quantitatively?&lt;/h3&gt;
&lt;p&gt;Umbrella pricing refers to the mechanism by which cartel members&amp;rsquo; higher prices raise the sectoral price index, allowing non-cartel members to expand output and raise their own markups without reducing their market share. Proposition 1 of the model shows that collusion increases the markups of all firms — cartel and non-cartel — with non-cartel members experiencing markup increases that are larger for larger non-members. Quantitatively, when non-cartel members are held to fixed markups (so the umbrella effect is turned off), the aggregate TFP gain from eliminating cartels rises from 1.11% to 1.14% — a difference of 0.03 percentage points, or less than 3% of the total effect. The welfare effect is similarly small: 2.01% versus 2.00%. The umbrella pricing channel thus dampens aggregate gains but is quantitatively minor.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-in-cartel-effects-is-documented"&gt;Q5. What heterogeneity in cartel effects is documented?&lt;/h3&gt;
&lt;p&gt;Three dimensions of heterogeneity are explored. First, cartel size matters: large cartels (those with cumulated market share above the median) account for roughly 80% of the aggregate TFP gain from eliminating all cartels (0.88 percentage points out of 1.11%), while small cartels account for only 0.23 percentage points. Second, cartel composition is critical: top-producer cartels amplify misallocation, all-inclusive cartels generate very large overcharges and dramatically higher misallocation, least-efficient-firm cartels barely affect allocation, and random-top-10% cartels can slightly improve allocation. Third, collusion intensity matters monotonically: across the range κ = 0.1 to κ = 0.4, TFP gains from elimination fall from 0.99% to 0.54%, and welfare gains fall from 1.70% to 0.85%.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run-and-how-do-the-results-change"&gt;Q6. What robustness checks are run, and how do the results change?&lt;/h3&gt;
&lt;p&gt;The paper runs six main robustness experiments, all recalibrating the model: (1) Alternative overcharge target of 15% (versus 10% baseline): requires κ = 1.28, yields TFP gains of 1.63% and welfare gains of 2.77%. (2) Low aggregate markup target M = 1.1: TFP gain of 1.37%, welfare gain of 2.07%. (3) High aggregate markup target M = 1.3: TFP gain of 0.90%, welfare gain of 1.96%. (4) Bertrand rather than Cournot competition: TFP gain of 0.55%, welfare gain of 1.35% — smaller because Bertrand generates less markup dispersion, though the reduction in distance to the efficient allocation is larger (39%). (5) Heterogeneous κ across cartels drawn from a truncated normal with four variance levels: TFP gains range from 0.84% to 1.11% and welfare gains from 1.53% to 1.99%, close to the benchmark of 1.11% and 2.00%. (6) The cartel screen regression yields an estimated κ of 0.70 from data on colluding firms, close to the calibrated benchmark of 0.79.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-model-generate-a-cartel-detection-screen-and-what-does-it-find"&gt;Q7. How does the model generate a cartel detection screen, and what does it find?&lt;/h3&gt;
&lt;p&gt;The model&amp;rsquo;s equilibrium first-order conditions imply a regression of a cartel member&amp;rsquo;s labor share (a proxy for the inverse markup under log-linear production) on its own market share and the total cartel market share. The ratio of the estimated coefficient on cartel market share to the sum of both coefficients recovers the collusion intensity κ. Running this regression on the sample of detected cartel firms, the authors find a coefficient on own market share of -0.53 and an intercept of 0.70, both significant at 1%. Adding the cartel joint market share, its coefficient is negative and significant at 1%; the estimated κ from this specification is 0.70, close to the benchmark of 0.79. Results are qualitatively robust to including year fixed effects, though estimates become slightly noisier.&lt;/p&gt;
&lt;h3 id="q8-how-do-the-authors-explain-the-large-discrepancy-with-harberger-1954"&gt;Q8. How do the authors explain the large discrepancy with Harberger (1954)?&lt;/h3&gt;
&lt;p&gt;Harberger&amp;rsquo;s classic estimate of the deadweight loss from monopoly is approximately 0.1% of GDP. The authors show that their model can reproduce estimates close to this when (a) the model is aggregated to the sectoral level, eliminating within-sector markup dispersion — in that case, the TFP gain from eliminating cartels falls to 0.08%; or (b) demand elasticities are set close to unity as in Harberger&amp;rsquo;s sectoral data — the TFP gain falls to 0.24%. The key reason for the discrepancy is that Harberger&amp;rsquo;s framework suppresses both the within-sector dispersion of markups (which in the baseline model amplifies allocative losses) and the endogenous markup response to market share changes (which is large when ρ is substantially greater than 1). Using disaggregated firm-level data and calibrated high-within-sector elasticities restores the large estimated costs.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper implies that antitrust enforcement against horizontal price-fixing cartels can yield aggregate TFP gains of 1.1% and welfare gains of 2% in consumption-equivalent terms — figures the authors describe as conservative, because (i) the estimate is static (no dynamic gains from entry or innovation effects are included), (ii) only domestic detected cartels are captured and international cartels are excluded, (iii) geographic markup dispersion is abstracted from, and (iv) the calibration uses a conservative overcharge target of 10%. Importantly, the gains from targeting the intensive margin (forcing cartels to reduce overcharges rather than dissolving them entirely) are also substantial: a 50% reduction in κ still yields 0.54% TFP and 0.85% welfare gains. The results further imply that industrial policy and trade liberalization reforms that ignore competition enforcement may be partially undermined if new market power enables cartelization. The scope condition most critical to the quantitative magnitude is cartel composition: results depend on cartels being composed of top producers; the sign and magnitude of productivity effects can flip for alternative compositions. The authors also note that if cartels spur long-run innovation (through higher profits), their static welfare cost estimates would overstate the net social cost.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-differ-from-edmond-midrigan-and-xu-2022-and-baqaee-and-farhi-2020"&gt;Q10. How does this paper differ from Edmond, Midrigan, and Xu (2022) and Baqaee and Farhi (2020)?&lt;/h3&gt;
&lt;p&gt;Edmond et al. (2022) and Baqaee and Farhi (2020) quantify the total welfare and productivity cost of markups relative to the efficient allocation — the gap between the current economy (with all its markup dispersion from firm heterogeneity) and the first-best. Moreau and Panon instead isolate the cost of one specific, policy-relevant source of excess markup dispersion — collusion — by computing the gap between the cartel equilibrium and the competitive (but still imperfect) Nash-Cournot equilibrium. They also show that competitive oligopoly models of the Edmond et al. type understate the total misallocation cost of markups by approximately 70% when cartels are present and composed of top producers, because competitive models are calibrated to match the same aggregate markup data but attribute all markup dispersion to firm heterogeneity rather than to collusion. The papers are thus complementary: Edmond et al. bound the full cost of all markup distortions, while Moreau and Panon bound the portion attributable to cartels and amenable to competition enforcement.&lt;/p&gt;
&lt;h3 id="q11-what-caveats-and-limitations-do-the-authors-acknowledge"&gt;Q11. What caveats and limitations do the authors acknowledge?&lt;/h3&gt;
&lt;p&gt;The authors flag several important limitations. (1) The analysis is static: dynamic effects — including entry deterrence by cartels, barriers to exit for inefficient firms, and the innovation-competition relationship — are not modeled. The relationship between competition and innovation is hump-shaped (Aghion et al., 2005), so cartels could in principle spur or dampen innovation; the authors treat their estimates as an upper bound if cartels raise innovation. (2) Only detected French domestic cartels are in the sample; international cartels (investigated by the European Commission) and undetected cartels are excluded, likely causing understatement of total costs. (3) The selection of detected cartels is non-random: the direction of bias from using only discovered cartels is unclear — discovered cartels may be unusually large (biasing costs upward) or undiscovered large cartels may exist (biasing costs downward). (4) The model abstracts from geographic markup dispersion and from vertical arrangements across industries. (5) The model has no entry or exit of firms, which could amplify or dampen transition dynamics. (6) Firm-level prices are unavailable, so markups cannot be directly measured and must be inferred from the model or from labor shares.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Collusion intensity parameter (κ)&lt;/strong&gt;: A scalar in [0,1] that governs the weight each cartel member assigns to co-conspirators&amp;rsquo; profits when choosing output. When κ = 0, behavior is competitive Cournot; when κ = 1, members jointly maximize aggregate cartel profits. In the baseline calibration κ = 0.79, chosen to match a 10% median cartel overcharge in French data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cartel overcharge&lt;/strong&gt;: The percentage difference in cartel members&amp;rsquo; average markups between the cartel equilibrium and the competitive Nash-Cournot equilibrium. Computed as the median overcharge across cartels in the model. In the baseline calibration it is 10%, consistent with the OECD benchmark and Laborde (2021). The overcharge increases with both collusion intensity (κ) and the cartel&amp;rsquo;s total market share.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Umbrella pricing&lt;/strong&gt;: The mechanism by which a cartel&amp;rsquo;s higher prices raise the sectoral price index, enabling non-cartel members to expand demand, gain market share, and charge higher markups than they would in the absence of the cartel. In the model, umbrella pricing implies that the introduction of collusion increases the markups of all firms in cartelized sectors, not just cartel members; quantitatively, the effect dampens but does not reverse the aggregate productivity gains from cartel dissolution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distance to efficient allocation&lt;/strong&gt;: The ratio of the productivity gain from eliminating cartels (Acartel → Acomp) to the total productivity gain from eliminating all markup dispersion (Acomp → Aeff or equivalently from Acartel → Aeff). In the baseline, eliminating cartels reduces this distance by 30%, meaning cartels are responsible for roughly 30% of the gap between the actual economy and the first-best efficient allocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous markups (size-related)&lt;/strong&gt;: In the Atkeson-Burstein framework embedded in this model, a firm&amp;rsquo;s equilibrium markup is a harmonic average of within- and between-sector demand elasticities weighted by the firm&amp;rsquo;s own market share. More productive firms endogenously hold larger market shares and thus face lower demand elasticities, charging higher markups. Collusion further distorts this by augmenting the effective market share with co-members&amp;rsquo; shares, yielding supracompetitive overcharges.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cartel composition&lt;/strong&gt;: The identity of firms within a cartel — specifically, where they sit in the within-industry productivity distribution. The paper shows this is the single most important determinant of whether cartels amplify or dampen aggregate misallocation. Empirically, discovered French cartels are composed of the largest, most productive firms (nearly 1,900% more sales than non-members), and this is the only composition configuration that can match observed 10% overcharges in the calibrated model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive versus extensive margin of cartel policy&lt;/strong&gt;: The extensive margin refers to whether a cartel exists (zero versus positive κ); the intensive margin refers to the degree of collusion among existing cartel members (high versus low κ). The paper shows both margins are quantitatively important: breaking down all cartels (extensive margin) yields 1.11% TFP gain, while halving κ without dissolution (intensive margin) yields 0.54% TFP gain and 0.85% welfare gain.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cartel screen&lt;/strong&gt;: A regression of cartel members&amp;rsquo; labor shares on their own market share and the joint cartel market share, derived directly from the model&amp;rsquo;s equilibrium first-order conditions. The collusion intensity κ can be recovered as the ratio of the joint market share coefficient to the sum of both market share coefficients. Applied to French data on detected cartel firms, this screen yields κ̂ = 0.70, close to the calibrated value of 0.79.&lt;/p&gt;</description></item><item><title>Illuminating the Global South</title><link>https://macropaperwarehouse.com/papers/illuminating-the-global-south/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/illuminating-the-global-south/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Satellite nighttime lights (luminosity) are the dominant remote-sensing proxy for local economic conditions in low-income countries, yet their accuracy at fine spatial scales and over time has remained contested. This paper by Chiovelli, Michalopoulos, Papaioannou, and Regan makes two linked contributions. First, it constructs a standardized, annual, global panel of nighttime lights from 1992 to 2023, integrating the legacy DMSP-OLS satellite series (1992–2013) with the higher-quality VIIRS series (2013–onward) after applying three adjustments to the noisier DMSP data: cross-sensor inter-calibration (following Li et al. 2020), top-coding correction (following Bluhm and Krause 2022, using a truncated Pareto distribution to replace pixels with Digital Number ≥ 55), and blooming correction (following Cao et al. 2019, modeling light spillover as spatial decay and subtracting predicted pseudo-light). VIIRS is then downgraded to DMSP-comparable units using an ensemble machine-learning method — extremely randomized trees trained on the single year of full overlap (2013) — yielding an out-of-sample RMSE of 1.50 versus 3.27 for the Li et al. sigmoid approach and 1.57 for the Nechaev et al. convolutional neural network; the F1 score for the binary lit/unlit classification is 0.72 versus 0.51 and 0.71 for those alternatives, with recall = 0.95 and precision = 0.58 against an actual lit-pixel share of only 8.6 percent globally. At the cross-country level — a sample of 173 countries — the adjusted series retains an elasticity of luminosity to GDP of approximately 0.85 and an R² around 0.9 in cross-section; for Africa specifically the elasticity is 0.7 and R² remains around 0.9. In long-difference panel regressions over 1992–2019, the luminosity-GDP elasticity is approximately 0.25–0.24, broadly consistent with Henderson et al. (2012)&amp;rsquo;s estimate of 0.30–0.33, while at the five-year panel frequency the elasticity is around 0.15–0.17. The second contribution is a systematic validation of the new series against multiple local development proxies across four low-income settings. Using 139 georeferenced DHS surveys from 34 African countries (gridcells of ~28km × 28km), the adjusted series yields cross-sectional coefficients of approximately 0.6 standard deviations for schooling, electricity access, and improved sanitation, and approximately 1 standard deviation for the composite wealth index, between lit and unlit gridcells; in within-gridcell panel regressions, the adjusted log-lights coefficient on schooling is approximately double that of the unadjusted series (~0.02 versus ~0.01), and lit/unlit panel coefficients are statistically significant only with the adjusted series — gridcells turning lit see schooling rise by ~0.05 standard deviations (~0.125 schooling years), wealth index rise by ~0.05 SD, and electricity access rise by ~0.05 SD. In Mozambique, using all post-civil-war censuses (1997, 2007, 2017) across 1,126 admin-4 localities, schooling and non-agricultural employment are at least 0.5 standard deviations higher in lit than unlit localities, equivalent to approximately 0.5 years of schooling and 10 percentage points of non-agricultural employment; within-locality changes in lights co-move significantly with schooling changes, with the difference in schooling gain between localities that turn lit versus stay unlit being about half a year even controlling for admin-3 fixed effects. In Indonesia, panel estimates for public goods across more than 60,000 PODES villages show the adjusted series yields a positive and significant coefficient on the composite wealth index while the unadjusted series yields a counterintuitively negative coefficient. In India, across more than 550,000 SHRUG villages and towns, the adjusted series consistently produces stronger cross-sectional and panel associations with non-farm, manufacturing, and services employment. A key empirical regularity across all settings is that the adjusted series outperforms the unadjusted one most sharply at finer spatial resolutions and in over-time (panel) comparisons, while at coarse aggregation levels (large administrative units or large grid squares) differences between the two series are minor, as spatial averaging attenuates measurement error in the unadjusted data too. Blooming correction delivers most of the improvement in the African context, where top-coding is rare (fewer than 2% of lit DMSP pixels in Africa approach the 63 DN ceiling). The paper also replicates three canonical studies — Michalopoulos and Papaioannou (2013) on precolonial ethnic institutions, Michalopoulos and Papaioannou (2014) on national institutions and split ethnic homelands, and Hodler and Raschky (2014) on regional favoritism — confirming that qualitative conclusions are robust to the data revision while documenting that the adjusted series sharpens several estimates, particularly those exploiting within-region over-time variation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper is a measurement and validation study rather than a causal identification exercise. Its core design is correlational: it regresses local development proxies on nighttime luminosity across gridcells and administrative units, conditioning on country-year fixed effects in cross-section and on unit fixed effects in panel regressions. The main threats are (a) reverse causation (luminosity and development are jointly determined), which the authors acknowledge but do not attempt to address — they are explicit that the goal is proxy validation, not causal estimation; (b) measurement error in both the luminosity variable and the development outcomes (DHS wealth index, census schooling, PODES public goods), which the paper addresses by comparing adjusted versus unadjusted luminosity series and interpreting attenuation bias reduction as evidence of improved measurement; (c) the binary transformation of luminosity (lit/unlit) produces non-classical measurement error — an explicit point drawn from econometric theory (Aigner 1973; Meyer and Mittag 2017) — which partly motivates the adjusted continuous series; and (d) spatial autocorrelation and systematic geographic patterns in prediction error, which the authors check by regressing prediction errors on latitude and longitude and find that the ERT-downgraded series reduces the latitude coefficient to 10% of its magnitude in the unadjusted VIIRS specification for log lights and to 35% for the lit indicator.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-dmsp-deficiencies-corrected-and-what-are-the-specific-methods-used"&gt;Q2. What are the three DMSP deficiencies corrected and what are the specific methods used?&lt;/h3&gt;
&lt;p&gt;Cross-sensor inter-calibration: DMSP data come from six satellites; Li et al. (2020) supply a cross-calibrated series using a second-order polynomial fitted on overlapping satellite years, which the paper adopts as its &amp;lsquo;unadjusted&amp;rsquo; baseline. Top-coding: DMSP records 8-bit Digital Numbers (DN) 0–63, so radiance above a ceiling is truncated. Pixels with DN ≥ 55 are subject to &amp;lsquo;implicit&amp;rsquo; top-coding (averages of potentially top-coded sub-readings). The correction uses the radiance-calibrated (RC) vintage available for seven years, ranks the top-coded pixels by the RC series from the nearest year, then replaces them with &amp;lsquo;structural values&amp;rsquo; drawn from a truncated Pareto distribution with parameters α = 1.5, L = 55, H = 2000. Blooming: the DMSP sensor stretches edge pixels and can be spatially displaced up to 3 km, causing light spillover. Following Cao et al. (2019), pseudo-light pixels (PLPs) — lit pixels neighboring at least one dark pixel — are identified. An OLS regression of PLP light on the inverse-squared-distance weighted sum of neighbors&amp;rsquo; light within a 7 × 7 window is estimated separately for broad global regions. The predicted blooming contribution is subtracted from each lit pixel, negative residuals are set to zero, and a local 3 × 3 mean smoothing is applied. Globally, the blooming correction raises the share of unlit pixels from 92% to 95% in 1992 and from 88% to 91% in 2012.&lt;/p&gt;
&lt;h3 id="q3-how-is-viirs-downgraded-and-harmonized-with-dmsp-and-what-does-extremely-randomized-trees-mean"&gt;Q3. How is VIIRS downgraded and harmonized with DMSP, and what does &amp;rsquo;extremely randomized trees&amp;rsquo; mean?&lt;/h3&gt;
&lt;p&gt;Because VIIRS records 14-bit DN at 15-arc-second resolution with far superior sensor quality, it is not directly comparable to the 8-bit, 30-arc-second DMSP. The authors&amp;rsquo; preferred approach downgrades VIIRS to match the DMSP scale. They use an ensemble machine-learning method called &amp;rsquo;extremely randomized trees&amp;rsquo; (Geurts et al. 2006), a variant of random forests that, instead of choosing the best splits from the training sample, picks split thresholds randomly, which further reduces variance and improves computational efficiency. Features used to predict DMSP-like values from VIIRS include: pixel statistics (mean, median, min, max of the four VIIRS sub-pixels within each DMSP 30-arc-second cell), statistics of neighboring pixels within windows of 3, 4, 7, 9, 11, 13, 17, and 21 pixel widths, and regional dummies for broad world regions. The model is trained on 2013 (the one full year of DMSP-VIIRS overlap) and its out-of-sample performance is assessed by retraining on 2012 and predicting 2013. Four merged series are produced corresponding to the four versions of DMSP (unadjusted; blooming only; top-coding only; both). The authors&amp;rsquo; approach outperforms both the Li et al. (2020) sigmoid-function method (RMSE 3.27 globally vs. 1.50) and the Nechaev et al. (2021) CNN approach (RMSE 1.57), especially in the low-to-middle luminosity range most relevant for low-income countries.&lt;/p&gt;
&lt;h3 id="q4-what-development-proxies-are-used-in-validation-and-across-what-samples"&gt;Q4. What development proxies are used in validation and across what samples?&lt;/h3&gt;
&lt;p&gt;Africa (DHS, 34 countries, 139 surveys, ~28km × 28km gridcells): mean years of schooling (respondents aged 15–39), DHS composite household wealth index, share of households with improved sanitation, share with electricity connection. All outcomes are standardized to mean zero, SD one. Mozambique (Census 1997, 2007, 2017, 1,126 admin-4 localities): mean years of schooling (aged 15–39) and non-agricultural employment (aged 15–24 or 19–24). Indonesia (PODES village census waves 1996–2018, 60,000+ villages): binary measures for garbage disposal, toilet use, drinking water access, gas/electricity for cooking, paved roads, and counts of kindergartens, primary, middle, and secondary schools — aggregated into a first principal component (eigenvalue ~3.5, capturing ~1/3 of variance). India (SHRUG dataset, 550,000+ towns and villages, Population Censuses 1991/2001/2011, Economic Censuses 1990/1998/2005/2013): population count, total non-farm employment, manufacturing employment, services employment.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented"&gt;Q5. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Spatial resolution: adjusted series outperforms unadjusted most at fine resolutions (2×2 gridcell blocks, ~56km × 56km at the equator); at coarse levels (12×12 blocks, ~336km × 336km), both series yield similar coefficients, as spatial aggregation attenuates noise in the unadjusted series. Urban vs. rural: cross-sectional estimates are similarly significant in urban and rural DHS samples. Panel estimates are statistically significant only with the adjusted series; urban panel coefficients are consistently larger than rural ones, echoing Asher et al. (2021)&amp;rsquo;s India finding. The adjustment matters more in rural areas than in urban areas in cross-section. Local variation (spatial RDD / fine fixed effects): with unadjusted series, panel wealth-index coefficients are statistically indistinguishable from zero until spatial fixed effects cover areas at least 7×7 gridcells (~200km × 200km at equator); with the adjusted series, coefficients remain significantly positive at all fixed-effect sizes including the finest 2×2 blocks. Top-coding vs. blooming: most of the improvement in Africa derives from blooming correction; top-coding correction has minor impact because fewer than 2% of lit African DMSP pixels approach the DN ceiling. Country-ethnic homelands (large areas, avg. 25,547 km²): adjustments matter little because spatial averaging already reduces noise. Applications replication: the precolonial institutions result (Michalopoulos and Papaioannou 2013) is robust and essentially unchanged because the units are very large. The national-institutions-at-border result (Michalopoulos and Papaioannou 2014) is strengthened in within-ethnicity specifications (coefficient marginally significant at 90% with adjusted series vs. p ≈ 0.15 with unadjusted); capital-proximity heterogeneity is sharpened. The regional-favoritism result (Hodler and Raschky 2014) strengthens: the log-lights lagged-leader coefficient rises from 0.038 to 0.058, and the lit-probability coefficient rises from ~3 to ~7 percentage points.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-and-specification-variations-are-run"&gt;Q6. What robustness checks and specification variations are run?&lt;/h3&gt;
&lt;p&gt;The paper compares four luminosity series (unadjusted Li et al.; blooming only; top-coding only; both combined + VIIRS fusion) to isolate each correction&amp;rsquo;s contribution. It checks the luminosity-GDP nexus at annual, five-year, and long-difference frequencies. It examines seven African countries&amp;rsquo; co-evolution of the harmonized series with electrification share (Kenya, DRC, Ghana, Tanzania, Nigeria, Mozambique, and one other) and finds no discontinuity at the 2012/2013 DMSP-VIIRS transition year. Spatial aggregation robustness: coefficients are computed across aggregation blocks ranging from 2×2 to 12×12 gridcells, showing stability in cross-section (~0.18) and mild size dependence in panel (~0.075, slightly rising with coarser units). Local variation robustness: fixed effects of increasing spatial coverage (2×2 to 12×12 cells) are added while the outcome remains at the gridcell level. Results replicated for schooling and electricity access (Appendix Section B.2) beyond the primary wealth-index outcome. Confounding by latitude in the ML model is assessed via regressions of prediction errors on latitude and longitude with and without country fixed effects. Median regressions confirm the OLS elasticity estimates at the cross-country level. The India analysis is replicated for both towns (urban) and villages (rural) separately.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Henderson et al. (2012): pioneer the use of luminosity as a cross-country GDP proxy and estimate a long-difference elasticity of 0.30–0.33 across 188 countries; this paper estimates 0.25–0.24 over a comparable specification, consistent but slightly lower. Gibson et al. (2021): show that VIIRS is superior to DMSP but find weak GDP-lights correlations outside cities for the early DMSP period in China, Indonesia, and South Africa; this paper addresses the concern by adjusting DMSP and merging it with VIIRS. Asher et al. (2021): validate luminosity as a strong proxy in India and find stronger urban-luminosity links; this paper replicates and extends those findings to Africa, Mozambique, and Indonesia and shows the adjusted series strengthens the Asher et al. patterns. Chen et al. (2024): find strong cross-sectional but weak panel associations; this paper&amp;rsquo;s adjusted series substantially strengthens panel associations. Bluhm and Krause (2022): provide the top-coding correction method adopted here. Cao et al. (2019): provide the blooming correction method. Nechaev et al. (2021): propose a CNN-based DMSP-VIIRS fusion but apply it to the unadjusted DMSP; this paper outperforms their RMSE slightly (1.50 vs. 1.57) and improves on their F1 score (0.72 vs. 0.71), with greater advantage in low-light regions. Li et al. (2020): propose a sigmoid-based fusion calibrated for high-light pixels; this paper substantially outperforms it (RMSE 1.50 vs. 3.27) particularly in low-luminosity areas. The paper thus synthesizes and extends multiple strands: it unifies the corrections of Bluhm-Krause and Cao et al., pairs them with state-of-the-art ensemble ML fusion, and provides by far the most comprehensive multi-country, multi-context validation of the resulting series.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary policy implication is methodological: researchers studying development in low-income countries should use the adjusted and harmonized nighttime lights series rather than raw DMSP data, and should be especially careful at fine spatial scales (e.g., spatial regression discontinuity designs, granular village-level analyses) and in panel specifications. The gains from adjustment are largest precisely where applied development research is moving — toward local identification strategies and over-time variation. For practitioners and statistical agencies, the series provides a low-cost annual proxy for local economic conditions in environments with weak administrative data, particularly across sub-Saharan Africa, South Asia, and Southeast Asia. Scope conditions: (a) Correlations are far from perfect — binary lit/unlit classification misses much variation in the many-zeros low-income context. (b) At large aggregate units (admin-1, country-ethnic homelands), the adjustments yield minimal additional improvement since noise averages out. (c) The series does not resolve the fundamental limitation that most of sub-Saharan Africa remains unlit (98.4% of DMSP pixels in Africa in 1992), so it captures variation among already-lit areas better than the development gradient at the zero-light frontier. (d) Future research blending nighttime lights with daytime imagery (traffic, built structures) is flagged as a promising extension, though daytime data are often proprietary.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-main-findings-from-the-three-replication-exercises"&gt;Q9. What are the main findings from the three replication exercises?&lt;/h3&gt;
&lt;p&gt;Michalopoulos and Papaioannou (2013) — precolonial ethnic institutions and contemporary development: Replication across 682 country-ethnic homelands confirms that areas with higher precolonial political centralization (as measured by a 0–4 jurisdictional hierarchy index) have significantly higher contemporary luminosity, conditional on country constants and geographic controls. With the adjusted series, the unlit share among homelands rises from 24% to 29% (because blooming correction removes spurious light), but the coefficients on political centralization are still highly significant, somewhat smaller in magnitude, and similar qualitatively. The main conclusion is robust because the units are large and spatial averaging already reduces noise in the raw series. Michalopoulos and Papaioannou (2014) — national institutions and split-border ethnic development: Replication across 38,427 gridcells of 220 systematically partitioned ethnic homelands. Cross-sectional results show a one-point increase in the rule-of-law index (range −2.5 to 2.5) is associated with a ~10 pp higher probability of a gridcell being lit. The within-ethnicity coefficient drops by more than half (~0.025). With the adjusted series, this within-ethnicity coefficient is marginally significant at 90% versus a p-value of ~0.15 with unadjusted. Spatial RDD coefficients remain small and insignificant regardless of adjustment. Capital-proximity heterogeneity: the positive association between rule of law and luminosity is significant only for ethnically split groups where both portions are close to their respective capitals, and this finding is more precisely estimated with the adjusted series; the effect is nil far from capitals in both series. Hodler and Raschky (2014) — regional favoritism: Panel replication across 38,427 subnational regions in 126 countries, 1992–2009. The lagged-leader dummy coefficient (log lights specification) rises from 0.038 to 0.058 with the adjusted series. The linear-probability-model lit indicator rises from ~3 to ~7 percentage points. All specifications with the adjusted series are at least two standard errors above zero, matching or exceeding the precision of the original.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-limitations-and-caveats-acknowledged-by-the-authors"&gt;Q10. What are the limitations and caveats acknowledged by the authors?&lt;/h3&gt;
&lt;p&gt;First, the correlations between luminosity and development are &amp;lsquo;far from perfect&amp;rsquo; — the binary lit/unlit transformation in particular fails to capture the significant continuous variation in assets, education, and public goods across regions that are all formally &amp;rsquo;lit.&amp;rsquo; Second, bottom-coding (under-recording of low-light areas) is acknowledged but not corrected; no existing method addresses it, though the authors note that their corrections nonetheless improve elasticities even in rural African regions with very low light. Third, downgrading VIIRS to DMSP by construction sacrifices some of the VIIRS data quality; the long-difference VIIRS elasticity for Africa (0.4) shrinks to 0.35 in the downgraded series. Fourth, daytime satellite imagery and combinations with nighttime lights (Jean et al. 2016; Yeh et al. 2020; Rossi-Hansberg and Zhang 2025) can better capture local wealth but are often proprietary and not replicable in standard economic research. Fifth, the top-coding correction in Africa is minor because very few pixels approach the DN=63 ceiling (0.98–1.7% of lit pixels in 1992–2012), so the main African improvement comes from blooming; other regions with denser urban cores may benefit more from top-coding correction. Sixth, the cross-sensor inter-calibration step is taken &amp;lsquo;off-the-shelf&amp;rsquo; from Li et al. (2020) and further investigation of sensor calibration is left to future work.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Top coding (DMSP)&lt;/strong&gt;: The truncation of Digital Number values at the 8-bit ceiling of 63 in DMSP-OLS data, caused by sensor calibration for cloud detection. Pixels with DN ≥ 55 also suffer &amp;lsquo;implicit&amp;rsquo; top coding because they represent averages of multiple potentially top-coded sub-readings. The paper corrects this by replacing top-coded pixels with structural values drawn from a truncated Pareto distribution, using the radiance-calibrated DMSP vintage to rank pixels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Blooming (spatial spillover of light)&lt;/strong&gt;: A measurement artifact in DMSP data whereby light from bright pixels spills into neighboring dark areas due to the sensor&amp;rsquo;s imprecise spatial accuracy and possible displacement of up to 3 km. The paper identifies pseudo-light pixels (lit pixels adjacent to at least one dark pixel), models the spillover as an inverse-squared-distance weighted function of neighboring lights, and subtracts the predicted blooming from each lit pixel. This correction raises the global unlit pixel share from 92% to 95% in 1992.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extremely randomized trees (ERT)&lt;/strong&gt;: An ensemble machine-learning method used to downgrade VIIRS luminosity data to the DMSP scale. Unlike standard random forests that find the best split thresholds within a random feature subset, ERT selects split thresholds randomly, reducing variance and improving computational efficiency. The authors train it on pixel statistics (mean, median, min, max) and neighborhood statistics within windows of varying sizes to predict DMSP-like values for 2014 onward from VIIRS readings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Harmonized (adjusted + fused) luminosity series&lt;/strong&gt;: The authors&amp;rsquo; main output: an annual global panel of nighttime lights from 1992 to 2023 that applies inter-sensor calibration, top-coding correction, and blooming correction to DMSP data (1992–2013), then uses the ERT ensemble model to convert post-2013 VIIRS data into DMSP-comparable units, yielding four variants (unadjusted, blooming only, top-coding only, both corrections) merged into a continuous time series at 30-arc-second (~1 km²) resolution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pseudo-light pixels (PLPs)&lt;/strong&gt;: In the blooming correction procedure, PLPs are defined as lit pixels (DN &amp;gt; 0) that have at least one dark neighbor (DN = 0). They are the pixels most likely to contain spurious light from neighboring bright areas. PLP light values are regressed on the inverse-squared-distance weighted sum of surrounding pixels to estimate the blooming decay function.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DHS composite wealth index&lt;/strong&gt;: Used in the validation analysis as a local development proxy: a principal-component aggregation of household characteristics including roof quality and ownership of consumer assets, constructed by the Demographic and Health Surveys program across African countries. The paper standardizes this and other outcomes to mean zero and standard deviation one for cross-outcome coefficient comparisons.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Spatial RDD (regression discontinuity design) using nighttime lights&lt;/strong&gt;: As applied in Michalopoulos and Papaioannou (2014) and referenced throughout, a design that restricts estimation to gridcells within a narrow band (e.g., 50 km) of a political or administrative border to compare otherwise similar areas on opposite sides, using luminosity as the outcome. The paper notes that such fine-resolution, localized comparisons are exactly the setting where measurement error in the unadjusted DMSP series is most consequential and where the adjusted series yields the largest improvement.&lt;/p&gt;</description></item><item><title>Import Liberalization as Export Destruction? Evidence from the United States</title><link>https://macropaperwarehouse.com/papers/import-liberalization-as-export-destruction-evidence-from-the-united-states/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/import-liberalization-as-export-destruction-evidence-from-the-united-states/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; How does import liberalization affect a country&amp;rsquo;s &lt;em&gt;export&lt;/em&gt; performance and welfare? Economic theory (Graham 1923, Ethier 1982, Krugman 1984) shows the answer hinges on whether production exhibits increasing returns to scale at the sector level. Krugman (1984) argued that with scale economies, import protection can be export-promoting because a protected industry expands, exploits scale economies, becomes more productive, and exports more — so conversely import liberalization is &amp;ldquo;export destroying.&amp;rdquo; The paper turns this logic into an empirical test: the sign of the import-liberalization-to-export relationship discriminates between constant-returns and increasing-returns trade models. Researchers otherwise lack tools to choose between these model classes, yet the choice matters greatly for multi-sector trade policy analysis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model and data.&lt;/strong&gt; The authors build a multi-sector general-equilibrium gravity model generalizing Krugman (1980) to many countries/sectors with input-output linkages (as in Caliendo-Parro 2015). The model nests constant returns (Armington, σ→∞) and increasing returns. The &amp;ldquo;scale elasticity&amp;rdquo; is 1/(σ−1); the &amp;ldquo;output elasticity&amp;rdquo; of exports equals the trade elasticity (ε−1) times the scale elasticity, and is positive iff there are increasing returns. The empirical application exploits US Permanent Normal Trade Relations with China (PNTR), passed Oct 2000, which removed tariff-revocation uncertainty. Exposure is measured by Pierce-Schott&amp;rsquo;s NTR gap (log gap between non-NTR and NTR tariffs; mean 0.23, SD 0.13, range 0–0.59). Trade data are from CEPII BACI; the baseline sample covers exports from 23 OECD countries (including the US) to 141 importers across 444 NAICS goods industries, in long differences (1995–2000 pre-period vs 2000–07 post-period).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings.&lt;/strong&gt; Reduced-form: US export growth fell in higher-NTR-gap industries after PNTR. The raw Figure 1 slope is −0.51 (SE 0.057); a 10-log-point NTR-gap increase is associated with 5.0 log points lower annual export growth, and the NTR gap explains 18% of cross-industry variation. This is inconsistent with constant returns and implies increasing returns in US goods production. An offsetting &lt;em&gt;input cost effect&lt;/em&gt; (lower imported-input costs) raises exports: PNTR reduced 2007 exports by 13% more for a 75th- vs 25th-percentile NTR-gap industry, but raised them 20% more for a 75th- vs 25th-percentile input-cost-shock industry; net effects range from −18% (Cigarettes) to +56% (Automobiles). A structural IV (NTR gap instrumenting output growth) yields an output elasticity of 0.74 (SE 0.41, preferred column).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantitative GE results.&lt;/strong&gt; Calibrating the output elasticity to 0.821 (matching the −0.10 conditional NTR-gap effect; trade elasticity set to 5), PNTR raised aggregate US exports/GDP by 3.2%, decomposed into −1.8% real market potential (export destruction), +2.4% input cost, and +2.7% foreign demand. Aggregate export growth is 28% larger with scale economies than without, because scale economies make the input-cost effect almost five times stronger (2.4% vs 0.5%). Exports nevertheless declined in the most exposed sectors (Textiles &amp;amp; Leather, Other Manufacturing), shifting US comparative advantage away from high-NTR-gap sectors. Welfare: PNTR raised US real income 0.068% (real expenditure 0.087%); gains are ~30% smaller than under constant returns because a negative specialization effect (−0.15%) offsets a larger ACR openness gain (0.22%). Chinese gains exceed US gains tenfold.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-theoretical-test-and-why-does-the-sign-of-the-import-liberalization-to-export-relationship-identify-returns-to-scale"&gt;Q1. What is the core theoretical test and why does the sign of the import-liberalization-to-export relationship identify returns to scale?&lt;/h3&gt;
&lt;p&gt;From the bilateral trade equation, the elasticity of exports to output equals the output elasticity (ε−1)/(σ−1), which is strictly positive iff there are increasing sector-level returns. Under constant returns (Proposition 1), conditional on foreign demand and domestic input costs, import liberalization does not affect exports (α1=0). Under increasing returns (Proposition 2), import liberalization shrinks domestic real market potential, lowers output, and — because productivity falls with output under scale economies — reduces exports to ALL destinations (α1&amp;lt;0), with the effect&amp;rsquo;s magnitude strictly increasing in the output elasticity. So estimating whether export growth falls in more-liberalized industries distinguishes the two model classes.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identification-strategy-and-its-main-threats"&gt;Q2. What is the identification strategy and its main threats?&lt;/h3&gt;
&lt;p&gt;A triple-difference: changes in US bilateral export growth by sector after PNTR relative to changes in other OECD exporters&amp;rsquo; growth, identified from the NTR gap interacted with Post and a US-exporter dummy. The estimating equation (12) uses importer-exporter-industry, importer-exporter-period, and importer-industry-period fixed effects to absorb importer demand, common-across-exporter technology shocks, and industry trends in supply capacity and trade costs. The NTR gap is plausibly exogenous because variation stems mostly from Smoot-Hawley (1930) non-NTR tariffs, unlikely related to economic conditions 70 years later; any endogeneity from NTR tariffs being higher in weak-growth industries would bias against finding a negative effect. Threat 1: unobserved US-specific technology shocks negatively correlated with the NTR gap not captured by input/skill/capital intensity controls. Addressed by re-estimating at HS 6-digit level with NAICS-industry-exporter-period fixed effects (Table 3), still finding negative effects. Threat 2: US-China competition in third markets — if PNTR shifted China&amp;rsquo;s export basket toward US-type products in high-NTR-gap industries. Tested by interacting with China&amp;rsquo;s market share (Table 4); the quadruple interaction is positive and insignificant, ruling this out.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-three-mechanisms-and-how-are-they-distinguished-empirically-and-quantitatively"&gt;Q3. What are the three mechanisms and how are they distinguished empirically and quantitatively?&lt;/h3&gt;
&lt;p&gt;(1) Real market potential / export destruction: import liberalization lowers the US price index, makes the domestic market more competitive, shrinks real market potential and output, and (under scale economies) cuts productivity and exports — identified by the negative α1 on the NTR gap. (2) Input cost effect: lower imported-input costs cut production costs and raise exports — identified by α2 on the input-output-weighted upstream NTR gap (CostShock), found negative and significant (lower input costs → higher exports). (3) Foreign demand effect: GE expansion of global demand and the trade-balance link between imports and exports — absorbed by fixed effects in the regression but recovered in the calibrated model&amp;rsquo;s decomposition (equation 16). In GE: −1.8% (market potential), +2.4% (input cost), +2.7% (foreign demand), netting +3.2%.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented"&gt;Q4. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Sector-level: the real market potential effect is negative in all goods sectors and stronger where the NTR gap is higher; the input cost effect is positively correlated with the NTR gap (due to heavy diagonal weight in the I-O table); the foreign demand effect is positive everywhere but uncorrelated with the NTR gap. Net exports/GDP rise in 12 of 15 goods sectors but fall in the highest-NTR-gap sectors — Textiles &amp;amp; Leather falls 22% (−32% market potential, +8.5% input cost, +4.6% foreign demand) and exports decline in 3 of the 4 highest-NTR-gap sectors. Under constant returns, by contrast, export growth is positive in all sectors and weakly POSITIVELY correlated with the NTR gap — qualitatively opposite. The correlation between sector-level export growth with vs without scale economies is insignificant (excluding Textiles &amp;amp; Leather) or significantly negative (including it).&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Appendix C checks robustness to: starting the post-period in 2001 instead of 2000; alternative NTR-gap definitions; aggregating exports across destinations; varying the exporter/importer/industry samples; allowing PNTR to affect domestic expenditure; and controlling for China import growth driven by non-PNTR shocks. An event study (equation 13, Figure 2) shows no NTR-gap/export relationship before 2000 and a negative one from 2001 until the 2007–08 financial crisis, ruling out pre-trends. The first-stage (Table 5) confirms higher-NTR-gap industries had lower OUTPUT growth (paralleling Pierce-Schott&amp;rsquo;s employment result). Alternative calibrations (Appendix D.5): without I-O linkages the market potential effect weakens but total export growth is roughly unchanged; allowing services scale economies raises US gains; combining Textiles &amp;amp; Leather with Other Manufacturing preserves results; using Bartelme et al. (2019) sector-varying elasticities still yields a negative specialization effect.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-output-elasticity-calibrated-and-how-does-it-compare-to-the-structural-estimate"&gt;Q6. How is the output elasticity calibrated and how does it compare to the structural estimate?&lt;/h3&gt;
&lt;p&gt;The output elasticity for goods is calibrated to 0.821 by matching the simulated NTR-gap effect to the −0.10 conditional reduced-form estimate (Table 2, column i), with services output elasticity set to zero and trade elasticity (ε−1) set to 5 (Head-Mayer 2014). This is below the value of 1 implied by Krugman (1980) or the Pareto-Melitz model but close to the Bartelme et al. (2019) mean of 0.83. It is reassuringly close to the independent structural IV estimate of 0.74 (SE 0.41). The simulated effect is decreasing in the output elasticity (consistent with Proposition 2 part ii) and rises sharply as the elasticity approaches one; the model has a unique solution for output elasticities below 0.95.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-welfare-decomposition-work-and-why-are-gains-smaller-with-scale-economies"&gt;Q7. How does the welfare decomposition work and why are gains smaller with scale economies?&lt;/h3&gt;
&lt;p&gt;Following Costinot-Rodríguez-Clare (2014), real-income gains decompose into an ACR term (changes in domestic expenditure share / trade openness) and a specialization term that exists only with scale economies (welfare from sectoral reallocation of employment, weighted by adjusted Leontief forward-linkage coefficients). With scale economies the ACR effect is +0.22% (vs +0.10% without), but it is more than offset by a −0.15% specialization effect, netting +0.068% real income — about 30% below the constant-returns gain. The specialization effect is negative because PNTR shifted resources toward services (weaker scale economies; goods output −0.55%, services +0.11%) and, more importantly per Appendix D.5, toward sectors with weaker FORWARD input-output linkages; cross-sectoral heterogeneity in scale economies alone contributes negligibly.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-relate-to-and-differ-from-closely-related-prior-work"&gt;Q8. How does this relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It extends Krugman (1984)&amp;rsquo;s partial-equilibrium oligopoly mechanism to a class of quantitative GE trade models (love-of-variety, external economies, Melitz-Pareto, or endogenous innovation — shown equivalent in Appendix A.3). Unlike prior scale-economy estimates (Antweiler-Trefler 2002, Lashkaripour-Lugovskyy 2018, Bartelme et al. 2019) and home-market-effect tests (Davis-Weinstein 2003, Costinot et al. 2019), it uses TRADE POLICY variation (not factor content, market size, or exchange rates) for identification and performs an ex-post policy analysis (echoing Goldberg-Pavcnik 2016). Relative to the PNTR/China-shock literature (Pierce-Schott 2016, Handley-Limão 2017, Autor-Dorn-Hanson 2013), it adds a new outcome — US EXPORTS and comparative advantage — and argues the &amp;lsquo;surprisingly swift&amp;rsquo; manufacturing decline would have been smaller absent scale economies. It complements Juhász (2018)&amp;rsquo;s infant-industry evidence (Napoleonic France) by quantifying the export-destruction cost while showing PNTR&amp;rsquo;s net effect on exports and welfare is positive. Dick (1994) tested the same hypothesis cross-sectionally for 1970 US data but found little support.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The findings support the existence of the scale-economies channel traditionally invoked to justify protection: pre-PNTR import protection shifted US comparative advantage toward the most-protected industries, and in the calibrated model targeted import protection CAN promote sector-level exports — but not under constant returns. However, the export-destruction effect is dominated, for most sectors and in aggregate, by export-promoting channels (input cost, foreign demand); total export growth is even greater WITH scale economies; and the negative specialization effect is more than offset by traditional gains from trade, so US gains from PNTR remain positive (+0.068% real income). Scope conditions: results rest on the calibrated output elasticity (0.821) and trade elasticity (5); the model assumes constant markups and full employment, so welfare excludes pro-competitive effects (Jaravel-Sager 2020, Amiti et al. 2020) and employment effects (Autor-Dorn-Hanson 2013); it studies a single liberalization episode; and the analysis cannot distinguish among alternative SOURCES of increasing returns. The authors stress accounting for scale economies (or their absence) is a prerequisite for correctly evaluating sector-level trade flows and welfare.&lt;/p&gt;
&lt;h3 id="q10-what-other-notable-findings-or-caveats-appear"&gt;Q10. What other notable findings or caveats appear?&lt;/h3&gt;
&lt;p&gt;PNTR is calibrated as a reduced-form openness shock (α5=0.43; equation 15), equivalent to a 13% average trade-cost reduction on US imports from China (SD 6.6% across industries) given trade elasticity 5 — matching Handley-Limão&amp;rsquo;s 13-percentage-point estimate. The calibrated economy has 12 economies and 24 sectors (15 goods). Chinese gains exceed US gains more than tenfold (because the US was much larger in 2000, so PNTR was a bigger shock to China), and China&amp;rsquo;s nominal wage rose 6.0% relative to the US, contributing to factor-price convergence. For comparison, Caliendo-Parro (2015) find NAFTA raised US welfare 0.08% and Fajgelbaum et al. (2020) find the Trump trade war cut US real income 0.04%. The model in changes is solved via exact hat algebra, holding each country&amp;rsquo;s trade deficit as a constant share of global value-added (which induces the positive import-export link in the foreign-demand term).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Labor Share, Markups, and Input-Output Linkages – Evidence from the U.S. National Accounts</title><link>https://macropaperwarehouse.com/papers/labor-share-markups-and-input-output-linkages-evidence-from-the-u.s.-national-accounts/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/labor-share-markups-and-input-output-linkages-evidence-from-the-u.s.-national-accounts/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;The paper asks why the U.S. labor share has declined over the postwar period, and whether rising markups or capital deepening (automation, falling capital prices) is the primary driver. The authors argue that the existing literature lacks consensus partly because micro-level studies weight producers by sales shares rather than Domar weights, which are gross-output-to-GDP ratios that correctly capture how sectoral changes propagate through the input-output structure. When intermediate inputs are themselves marked up by their producers and then re-marked-up by downstream firms (&amp;ldquo;double marginalization&amp;rdquo;), a modest sectoral markup increase is amplified into a substantially larger aggregate effect.&lt;/p&gt;
&lt;p&gt;The empirical framework is a two-sector (goods versus services) multisector extension of the Farhi-Gourio (2018) model with Cobb-Douglas production functions and monopolistic competition in the Dixit-Stiglitz tradition. The model is calibrated to three balanced-growth-path subperiods — 1957–1973, 1984–2000, and 2001–2016 — using U.S. NIPA data covering gross output, intermediate inputs, compensation, capital stocks, investment, and sectoral price-dividend ratios from Kenneth French&amp;rsquo;s data library. The unobservable user cost of capital, which is needed to separate normal capital returns from markups (factorless income), is backed out from the model&amp;rsquo;s Euler equation via the Gordon growth formula applied to sectoral price-dividend ratios and includes a risk premium.&lt;/p&gt;
&lt;p&gt;Main quantitative findings: The aggregate labor share fell 4.9 percentage points (pp) from 1957–1973 to 2001–2016. Aggregate markups rose 6.6 pp (from 1.072 to 1.138), more than either sector&amp;rsquo;s standalone increase, because double marginalization through input-output linkages amplifies sectoral markups into a larger aggregate effect. Sectoral gross-output markups rose approximately 3.8 pp in goods (1.039 to 1.077) and 3.3 pp in services (1.034 to 1.067). In the top-down counterfactual holding markups constant at their 1957–73 levels, the labor share falls only 0.5 pp instead of 4.9 pp — markups account for 4.4 pp of the total 4.9 pp decline. Holding labor output elasticities constant instead yields only a 2.2 pp decline; holding materials elasticities constant reduces the decline by 0.8 pp; holding structural change (sector output weights) constant causes the labor share to fall 7.7 pp — meaning structural reallocation to services offset 2.8 pp of the decline. In a bottom-up Taylor decomposition, the first-order direct effects of rising markups account for 5.0 pp and falling labor output elasticities account for 4.4 pp — together nearly twice the actual 4.9 pp decline, confirming the Grossman-Oberfield (2021) observation that individual candidate forces over-explain the total. The offsetting effects that reconcile the over-explanation are: (i) the interaction of falling goods-sector labor elasticities with structural change toward services (which have a higher and slightly rising labor elasticity) offsets 3.6 pp, and (ii) the interaction of rising markups with changing sector weights offsets a further 0.8 pp; the aggregate labor output elasticity αL barely changes (0.794 to 0.788) because capital deepening in goods (goods value-added labor elasticity fell from 0.807 to 0.700) is fully offset by reallocation to services (services value-added labor elasticity rose from 0.790 to 0.814). The final-output share of goods fell by more than half, from 0.460 to 0.194. Materials intensities rose in both sectors (goods non-intermediate factor share fell from 0.374 to 0.353; services from 0.625 to 0.572), amplifying double marginalization over time. The user cost of capital declined from roughly 13.8% to 12.3% in aggregate, driven by falling expected discount rates (from ~6.1% to ~3.4%), partially offset by rising depreciation rates. When IPP capital is excluded from NIPA measurement, the aggregate labor share declines by only 1 pp (from 0.743 to 0.732), consistent with Koh et al. (2021).&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s core implication for the debate is that forces concentrated in the goods sector — capital deepening, automation, globalization, declining union power — cannot account for the aggregate labor share decline because the goods sector shrank dramatically and structural change to services largely offsets goods-specific capital deepening. A credible candidate explanation must affect both goods and services with similar strength, and rising markups do: sectoral gross-output markups increased by similar amounts in both sectors (roughly 3.3–3.8 pp each), and input-output linkages amplify their aggregate impact substantially.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-model-structure-and-why-does-it-differ-from-one-sector-models"&gt;Q1. What is the model structure and why does it differ from one-sector models?&lt;/h3&gt;
&lt;p&gt;The model is a two-sector (goods, services) extension of Farhi-Gourio (2018) with Epstein-Zin preferences, Dixit-Stiglitz aggregation of varieties within each sector, Cobb-Douglas production in capital, labor, and intermediate inputs from both sectors, and sector-specific markups under monopolistic competition. The two-sector structure is essential because (i) labor shares differ substantially across sectors at any point in time, (ii) they evolve differently over time, and (iii) goods production is far more materials-intensive than services. A one-sector model cannot capture the double marginalization amplification, the input-output linkages between sectors, or the offsetting effects of structural change.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identification-strategy-for-separating-output-elasticities-from-markups"&gt;Q2. What is the identification strategy for separating output elasticities from markups?&lt;/h3&gt;
&lt;p&gt;A well-known challenge is that one must split the residual between capital&amp;rsquo;s normal return and pure profit (markup). The authors do not use micro production data. Instead they calibrate the user cost of capital from the model&amp;rsquo;s balanced-growth-path Euler equation: ρj is inferred from the Gordon growth formula applied to sectoral price-dividend ratios from Kenneth French&amp;rsquo;s data library. Given ρj, the depreciation-plus-capital-loss term δj + γQ is inferred from the sectoral investment-capital ratio. The markup then equals sectoral gross output value divided by the sum of all observed factor payments (labor compensation, materials costs) plus the imputed capital cost (user cost times capital stock). Output elasticities of each factor equal their respective cost shares in total factor payments, a standard Cobb-Douglas result.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-main-threats-to-identification-and-how-are-they-addressed"&gt;Q3. What are the main threats to identification and how are they addressed?&lt;/h3&gt;
&lt;p&gt;Three main threats are addressed. First, the price-dividend ratio (the key input for ρj) covers only listed corporations, not all private firms; the authors note listed firms account for about 60% of business capital, and Atkeson, Heathcote, and Perri (2025) find very similar rates of return using a broader measure. Second, the shift toward share repurchases rather than cash dividends may understate payout yield and overstate the fall in ρ, inflating markups; the authors rerun the model using Boudoukh et al. (2007) repurchase-adjusted yields and find aggregate markups still increase by 5 pp (vs. 7 pp in the baseline). Third, the balanced-growth-path assumption imposes constant ratios within each subperiod, which may be violated; the robustness exercise recalibrating with 2016 end-of-sample values yields nearly identical conclusions.&lt;/p&gt;
&lt;h3 id="q4-how-is-double-marginalization-measured-and-why-does-it-matter-so-much"&gt;Q4. How is double marginalization measured and why does it matter so much?&lt;/h3&gt;
&lt;p&gt;Double marginalization arises because approximately half of U.S. gross output value is materials costs, and those inputs are purchased from monopolistically competitive suppliers who charge a markup. When the downstream firm marks up its own price, it marks up the cost of already-marked-up inputs a second time. Formally, the aggregate markup exceeds any sectoral markup because intermediate goods get embedded in final goods through the Leontief inverse (Domar weights). The paper proves in Proposition 3 that aggregate markups are the same in gross-output and value-added models, but sectoral value-added markups are always larger than gross-output markups; this means taking simple cost- or revenue-weighted averages of sectoral value-added markups overstates the implied market power and misrepresents the channel. Materials intensities rose in both sectors over the sample, so double marginalization has itself increased over time, adding to the aggregate markup rise beyond what sectoral gross-output markups alone would imply.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented-across-sectors"&gt;Q5. What heterogeneity is documented across sectors?&lt;/h3&gt;
&lt;p&gt;Goods and services differ in three key respects that are quantified: (i) Materials intensity: goods gross-output materials share is roughly 0.60 versus 0.40 for services (2001-16 averages), making double marginalization far stronger in goods. (ii) Capital deepening: the goods value-added labor elasticity ˜αLg fell from 0.807 to 0.700 (-10.7 pp) while services ˜αLs rose from 0.790 to 0.814 (+2.3 pp). (iii) Domar weights: the goods Domar weight Φg fell from 1.018 to 0.572 while services Φs rose from 0.933 to 1.289, reflecting the shift of economic activity toward services. Despite these differences, sectoral gross-output markups increased by similar amounts in both sectors (3.8 pp goods, 3.3 pp services), which is the main reason markups can explain the aggregate decline while sector-specific capital deepening cannot.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Four sets of robustness exercises are presented. (1) Balanced-growth-path assumption: the model is recalibrated using only 2016 end-of-sample values for the final period; results are nearly unchanged, though markups come out slightly higher. (2) Dividend measurement: repurchase-adjusted payout yields from Boudoukh et al. (2007) are used; aggregate markups still increase by 5 pp rather than 7 pp, so the conclusion is unchanged though the magnitude is modestly smaller. (3) Missing capital — organizational capital: using Crouzet-Eberly (2021) estimates, including organizational capital reduces the markup level (from 1.138 to 1.092 in 2001-16) but barely changes the markup increase (from 4.6 pp to 3.9 pp). (4) Missing capital — land: industrial and commercial land values over 2002-16 averaged roughly $2 trillion versus a private non-real-estate capital stock of $16.6 trillion; eliminating the markup increase would require a 27% rise in the capital-output ratio, but land can provide at most a 12% increase even under extremely counterfactual assumptions. (5) Alternative user cost: using Barkai (2020)&amp;rsquo;s Aaa interest rate yields aggregate markups increasing from 1.101 to 1.151 (1984-2000 to 2001-16), similar to baseline. (6) IPP capital: excluding IPP from NIPA yields only a 1 pp labor share decline, consistent with Koh et al. (2021). (7) Intangible capital generally: including intangible capital in the NIPA does not change the importance of markups.&lt;/p&gt;
&lt;h3 id="q7-how-do-the-authors-reconcile-their-low-gross-output-markups-with-the-much-higher-firm-level-markups-found-by-de-loecker-eeckhout-and-unger-2020"&gt;Q7. How do the authors reconcile their low gross-output markups with the much higher firm-level markups found by De Loecker, Eeckhout, and Unger (2020)?&lt;/h3&gt;
&lt;p&gt;The reconciliation has two parts. First, weighting: De Loecker et al. use sales-weighted markups, whereas the model-correct weighting in this context is harmonic cost-weighting (Hasenzagl and Perez, 2023); cost-weighted markups in this paper grow only 5 pp (from 1.193 to 1.246) versus 21 pp for sales-weighted markups, substantially narrowing the gap. Second, returns to scale and fixed costs: the paper assumes constant returns to scale and no fixed costs, so all markup revenue is pure profit. Firm-level studies assume fixed costs exist, meaning their markups must cover both pure profits and overhead, so markups are mechanically larger. A fixed cost share of about 15% of production costs accounts for the remaining difference between the two estimates. Importantly, both approaches produce similar economic profit rates: this paper finds sales-weighted profit rates of 4.5% (1984-2000) rising to 6.6% (2001-16), similar to De Loecker et al.&amp;rsquo;s finding of profit rates rising from 1% in 1980 to 8% in 2016.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-role-of-structural-change-and-why-does-it-offset-capital-deepening-but-not-markups"&gt;Q8. What is the role of structural change, and why does it offset capital deepening but not markups?&lt;/h3&gt;
&lt;p&gt;Structural change — the reallocation of final-output expenditure shares away from goods toward services — acts as a natural counterweight when a factor depresses labor share only in the shrinking sector. Capital deepening (falling goods labor elasticity) is concentrated in goods; as goods&amp;rsquo; expenditure share fell from 0.460 to 0.194, the weight placed on goods in the aggregate labor share shrank, largely undoing the direct effect of capital deepening on aggregate labor share. The second-order interaction term in the bottom-up decomposition confirms this: the interaction of falling labor elasticities with changing sector weights offsets 3.6 pp. In contrast, markups rose by similar amounts in both goods and services, so there is no equivalent shrinking-sector effect to offset the markup increase; summing the direct markup effect (−5.0 pp) with the markup-weight interaction (+0.8 pp) gives approximately the full observed decline.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-treat-the-possibility-of-labor-monopsony-as-an-explanation"&gt;Q9. How does the paper treat the possibility of labor monopsony as an explanation?&lt;/h3&gt;
&lt;p&gt;The paper acknowledges that labor monopsony (markdowns over wages) could in principle produce a gap between price and marginal cost similar to product markups. However, the authors argue the evidence does not support a role for increasing markdowns in driving the aggregate labor share trend: Yeh, Macaluso, and Hershbein (2022) find large markdowns in manufacturing but no role for them in explaining the time series of manufacturing labor share; Deb et al. (2022), allowing for both markups and markdowns, attribute changes in the price-marginal-cost gap to markups; Kirov and Traina (2023) find similar evidence in manufacturing. The authors therefore interpret their factorless-income estimates as markups rather than markdowns.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The main policy implication is that explanations for the labor share decline should focus on product market power (markups) operating across both the goods and services sectors, not primarily on capital-deepening forces such as automation or falling capital prices. The paper does not directly propose policy remedies, but the results imply that policies targeting capital deepening or trade-induced displacement alone cannot fully explain or reverse the aggregate labor share trend. The analysis is scoped to the U.S. private economy excluding real estate, 1957–2016, and the two-sector decomposition. The authors acknowledge the NIPA-based approach cannot directly speak to the firm-level sources of increasing markups (market concentration, fixed costs, intangibles), leaving the microeconomic explanation for rising sectoral markups to future research. Extension to finer industry disaggregations and other countries is flagged as a direct next step.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-papers-contribution-relative-to-farhi-gourio-2018-and-karabarbounis-neiman-2014"&gt;Q11. What is the paper&amp;rsquo;s contribution relative to Farhi-Gourio (2018) and Karabarbounis-Neiman (2014)?&lt;/h3&gt;
&lt;p&gt;Farhi-Gourio (2018) is a one-sector model calibrated at the aggregate level; this paper extends it to two sectors with explicit input-output linkages, allowing the decomposition to distinguish sector-specific from aggregate forces and to quantify double-marginalization amplification. Karabarbounis-Neiman (2014) attributed the labor share decline primarily to falling relative prices of capital (capital deepening) driven by an elasticity of substitution between capital and labor exceeding one; this paper&amp;rsquo;s calibration finds that the aggregate output elasticity of labor barely changes (0.794 to 0.788), which is inconsistent with capital deepening as the dominant aggregate force, and notes that evidence from Herrendorf et al. (2015) and Oberfield-Raval (2021) suggests the elasticity of substitution is below one in most of the goods sector. Moreira (2022) also uses an input-output model but does not allow markups by intermediate producers, which this paper shows is quantitatively crucial.&lt;/p&gt;
&lt;h3 id="q12-why-does-the-paper-use-nipa-data-rather-than-firm--or-establishment-level-data"&gt;Q12. Why does the paper use NIPA data rather than firm- or establishment-level data?&lt;/h3&gt;
&lt;p&gt;NIPA data have four advantages in this context: (i) they cover all market activity rather than just publicly listed or large firms; (ii) they capture inter-sectoral input-output linkages that micro datasets lack; (iii) they include broad coverage of intangible assets (IPP) following the 1999 and 2013 revisions; and (iv) they respect standard accounting adding-up constraints, ensuring that sectoral forces aggregate consistently to the macro level. The NIPA-based calibration also has limited data requirements, making it feasible to extend the analysis back to the late 1950s and, potentially, to other countries. The main limitation is that NIPA data are available only at the two-sector level of aggregation for the full postwar period, due to the switch from SIC to NAICS classification in 1997 and the aggregated reporting of some items like proprietors&amp;rsquo; income.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Domar weight&lt;/strong&gt;: The ratio of a sector&amp;rsquo;s gross output value to aggregate final output (GDP). Unlike expenditure weights, Domar weights exceed one when summed across sectors because they capture how a sector&amp;rsquo;s output is both a direct contributor to final demand and an indirect contributor through its use as intermediate inputs elsewhere. The paper uses Domar weights as the correct aggregation weights for sectoral labor shares, showing that properly accounting for input-output linkages through these weights is essential for connecting sectoral forces to aggregate outcomes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Double marginalization&lt;/strong&gt;: The amplification of sectoral markups at the aggregate level that occurs because intermediate inputs are priced above marginal cost by their producers (first markup) and then purchased and re-priced above marginal cost by downstream firms (second markup). In this paper&amp;rsquo;s model, double marginalization causes the aggregate markup to exceed either sector&amp;rsquo;s standalone gross-output markup; with roughly half of U.S. gross output being materials cost, the amplification is quantitatively large (aggregate markups of 6.6 pp increase versus sectoral increases of only 3.3–3.8 pp).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gross-output markup&lt;/strong&gt;: The ratio of a sector&amp;rsquo;s gross output value to the sum of all factor payments at the gross-output level (capital user costs times capital stock, plus labor compensation, plus the cost of intermediate inputs from all sectors). Under perfect competition this ratio equals one; deviations above one represent market power. This differs from value-added markups, which divide value added by only capital and labor payments, and are therefore mechanically inflated in materials-intensive sectors via double marginalization even when gross-output markups are identical across sectors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Risky balanced growth path (RBGP)&lt;/strong&gt;: The equilibrium concept used for calibration, extending the standard balanced growth path to allow for rare disaster shocks (Farhi-Gourio). Along the RBGP, expected variables grow at constant rates, but occasional level shifts occur when the rare disaster shock materializes. This allows the model to have realistic risk premia embedded in the discount rate ρ while maintaining analytically tractable solutions; the calibration avoids modeling transitional dynamics and instead compares the RBGP parameters across sub-periods.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Output elasticity of labor (αL)&lt;/strong&gt;: The Cobb-Douglas coefficient on labor in the production function, equal under the paper&amp;rsquo;s calibration to each factor&amp;rsquo;s cost share in total factor payments. Changes in αL represent capital deepening (automation, falling capital prices) when αL falls because capital displaces labor in production. The key finding is that αL barely changes at the aggregate level (0.794 to 0.788) over the full 1957–2016 period because capital deepening in goods is offset by structural change toward services.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Factorless income&lt;/strong&gt;: Income that remains after subtracting payments to labor (at market wages) and payments to capital (at normal user cost rates) from gross output. In this model, factorless income equals markup revenue (the portion of output value above total factor payments). Rising factorless income / markups are the mirror image of the declining labor share when the output elasticity of labor does not change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural change (in this paper&amp;rsquo;s sense)&lt;/strong&gt;: The reallocation of final-output expenditure shares across goods and services sectors over time, captured by changes in the expenditure weights ϕj. The paper documents that the goods final-output share fell by more than half (from 0.460 to 0.194) over 1957–2016. Structural change acts as a counterweight to any force concentrated in the goods sector: as goods&amp;rsquo; weight shrinks, the aggregate labor share becomes more determined by services.&lt;/p&gt;</description></item><item><title>Long-Distance Trade and Long-Term Persistence</title><link>https://macropaperwarehouse.com/papers/long-distance-trade-and-long-term-persistence/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/long-distance-trade-and-long-term-persistence/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether the location of economic activity adapts to changes in the location of trading opportunities, or whether historical patterns of trade permanently fix where cities emerge and grow. The question is fundamental to economic geography: many large cities owe their origins to access to long-distance trade that has since moved on, yet the cities persist. The empirical context is the staggered liberalization of direct transatlantic trade across the Spanish Empire in the second half of the 18th century. Before the reform, a mercantilist system confined legal trade to four American ports (Cartagena de Indias, Callao, Portobello/Nombre de Dios, and Veracruz) and a single European port (Seville, then Cadiz). Following Spain&amp;rsquo;s defeat in the Seven Years&amp;rsquo; War, a sequence of decrees opened direct trade to an additional 40-plus ports between 1765 and the early 19th century. The reform was driven by European interstate competition and implemented from above, creating staggered, quasi-exogenous variation in transportation times to Europe across American cities.&lt;/p&gt;
&lt;p&gt;The empirical strategy is a difference-in-differences design. The author constructs a novel panel of 62 cities in Spanish America observed every 50 years from 1600 to 1850 (372 observations), plus a settlement-level panel of 53,581 grid-cell-decade observations for 1710-1810. The key treatment variable is the time-varying transportation time to Europe, computed via a directed network using maritime logbooks (282,322 daily entries from the CLIWOC 2.1 database, 1750-1855) to estimate wind-conditional sailing speeds, and land travel models based on slope, elevation, landcover, and postal routes. The reduction in transportation time ranged from 0 to 38.3 days across locations, with an average pre-reform time of 93.5 days and an average reduction of 7.7 days - economically significant, representing 0 to 40 percent of the baseline average.&lt;/p&gt;
&lt;p&gt;The paper documents four main empirical patterns, all within a city-and-time fixed-effects framework that absorbs time-invariant location fundamentals. First, the reform improved market integration: non-bullion Spanish imports from the Americas rose nearly fourfold after the 1778 decree (following no secular trend before 1765), and a commodity price ratio between Spain and Spanish America converged beginning in the second half of the 18th century, consistent with lower transportation costs facilitating arbitrage. Second, lower transportation times raised urban population. In the preferred specification, a one-day reduction in transportation time to Europe increases city population by approximately 2 percent over a 50-year period (baseline coefficient -0.023, significant at conventional levels, with the sign and approximate magnitude stable across specifications adding viceroyalty-by-year or country-by-year fixed effects and controls interacted with year indicators). Third, the effects are concentrated among smaller cities and in the fringe regions of the empire (Argentina, Chile, Venezuela, the Caribbean, etc.): for the fringe region the point estimate is -0.016, while for the colonial core (Mexico, Peru, Bolivia) the effect is statistically indistinguishable from zero. A ten-day reduction in transportation time raises the probability that a grid cell contains a settlement by approximately one percentage point (against a sample mean of 11 percent), suggesting the primary margin was growth of existing cities rather than expansion to new frontier areas. Fourth, the cross-sectional elasticity of contemporary (year 2000) population density to pre-reform (1750) population size is 0.592 overall, but falls to 0.369 for cities that experienced large reductions in transportation times, and rises to 0.866 for cities that experienced little change - consistent with the reform attenuating the persistence of pre-reform settlement patterns specifically where the trade shock was large.&lt;/p&gt;
&lt;p&gt;To interpret mechanisms and simulate long-term implications, the author calibrates a dynamic spatial general equilibrium model built on Allen and Donaldson (2022). The model features cities that differ in productivity, land endowments, and trade/migration costs, with agents living two periods, static and dynamic agglomeration economies (parameters a1 = 0.055 and a2 = 0.063 from the data), and Frechet-distributed migration preferences. Counterfactual exercises simulate the model forward 300 years. In the benchmark counterfactual, the average reduction in transportation costs increases urban population by 1.27 percent (25th/75th percentile: -0.06 to 1.34 percent), with a maximum city-level gain of 11.77 percent and a minimum of -0.2 percent. Effects in the fringe region average 1.9 percent population gain versus 0.26 percent in the core. Decomposition exercises show that: differences in location fundamentals (productivity and land endowments, A and H) account for part of the core-fringe differential (the gap falls from 1.64 to 1.11 percentage points when fundamentals are equalized); equalizing the pre-reform population size across cities leaves the gap nearly unchanged (1.64 to 1.65), suggesting dynamic agglomeration from historical size plays little role in driving the core-fringe difference; by contrast, equalizing the spatial incidence of the shock (the amount by which transportation times fell) closes the differential almost entirely (gap falls to 0.16 percentage points), indicating that the fringe was simply more restricted before the reform and thus received a larger shock. Trans-Atlantic migration is also an important channel: when trans-Atlantic migration is made prohibitively costly in the model, the average population effect falls to roughly 14.86 percent of the benchmark, indicating that migration from Europe amplified the effect of lower trade costs on city populations.&lt;/p&gt;
&lt;p&gt;The overarching conclusion is that economic geography is not fully path-dependent: where trading opportunities move, economic activity can follow - but this adaptation is conditional. Cities that had already accumulated large populations before the change in trading locations are insulated from reallocation, because their internal market size reduces their reliance on long-distance external trade. In less-developed fringes with smaller internal markets, however, the spatial distribution of activity is more malleable and adjusts substantially to the change in trading opportunities.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The identification strategy is a two-way fixed-effects difference-in-differences, exploiting cross-city variation in the change in transportation time to Europe induced by the staggered port-opening reform. City fixed effects absorb all time-invariant unobserved location fundamentals (agroclimatic characteristics, disease environment, natural harbors, etc.). Year fixed effects absorb common time-varying shocks. The key identifying assumption is parallel trends: absent the reform, population growth would have evolved similarly across cities with different exposure to the transportation-time reduction. Three main threats are addressed. First, selective port targeting: if policymakers chose to open ports in anticipation of their commercial potential, the reform would not be exogenous to growth trajectories. The author argues against this: historical accounts indicate reluctance to open the wealthiest colony (New Spain/Mexico) precisely because its prosperity might divert trade from other regions, and the reform was driven by European interstate competition (the Seven Years&amp;rsquo; War) rather than by American economic conditions. Second, confounding from contemporaneous administrative reforms: Bourbon-era reorganizations, new viceroyalties (Rio de la Plata and Nueva Granada), and ecclesiastical changes could coincide with the trade reform. The author addresses this by dropping cities in the two new viceroyalties (coefficients remain similar) and by including viceroyalty-by-year fixed effects. Third, the transportation network itself might endogenously reflect urban growth (roads built to connect growing cities). The author notes the transportation times are constructed from predetermined geographic characteristics (wind patterns, slope, elevation, landcover) and pre-existing postal routes, not from contemporaneous road-building. The dynamic pre-trends test (interacting the reform-induced change in transportation time with year indicators) shows no significant difference in population growth across differentially exposed cities before 1750, supporting the parallel trends assumption. An alternative synthetic control design yields qualitatively similar results (treatment effect of approximately 19 percent over one century). The author also estimates the model on the sub-sample of cities far from ports (distance above median) and finds similar coefficients, addressing concerns that the reform directly targeted port cities for reasons correlated with their growth.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically-and-in-the-model"&gt;Q2. What are the main mechanisms and how are they distinguished empirically and in the model?&lt;/h3&gt;
&lt;p&gt;Two principal mechanisms are proposed. The first is a trade-cost channel: lower transportation times reduce the iceberg cost of exporting to European markets, lowering the price index for traded goods in affected cities and raising real income, which attracts labor migration. This channel operates even if migration costs are unchanged. The second is a migration-facilitation channel: lower transportation times reduce information frictions and direct travel costs for migrants from Europe, lowering migration frictions as well as trade costs. The quantitative model distinguishes these by running counterfactuals with and without changes in migration frictions (keeping migration costs fixed at 1760 levels). In the benchmark, allowing migration frictions to fall alongside trade costs yields an average 1.27 percent population increase; fixing migration frictions yields 0.66 percent. This comparison indicates that lower migration frictions approximately double the population effect relative to the pure trade-cost channel. The model also distinguishes between trans-Atlantic migration (Spain to Americas) and intracolonial migration. When trans-Atlantic migration is shut off (migration costs set prohibitively high for Europe-America pairs), the average population effect falls to approximately 15 percent of the benchmark value, implying that trans-Atlantic migration is the dominant driver of the population response. A third dimension of heterogeneity concerns internal market size: in the partial-equilibrium analytics, the marginal impact of a reduction in the trade cost to Europe is attenuated in larger cities because a larger local market reduces the share of consumption sourced externally, making the price index less sensitive to external trade costs. This mechanism is consistent with the finding that the reform had statistically significant effects only in smaller cities and fringe regions.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented-and-what-explains-it"&gt;Q3. What heterogeneity is documented and what explains it?&lt;/h3&gt;
&lt;p&gt;The paper documents three main dimensions of heterogeneity. First, the colonial core (Mexico, Peru, Bolivia) versus the fringe (Argentina, Chile, Venezuela, Caribbean, Central America): the average effect in the fringe region is -0.016 per day of transportation time in the baseline city regressions (statistically significant), while the core coefficient is indistinguishable from zero. In the counterfactual model, the fringe shows a 1.9 percent average population gain against 0.26 percent in the core. Second, large versus small cities: the effects are larger and more precisely estimated for cities with below-median pre-reform population. Third, within the fringe, there is wide dispersion: the 25th/75th percentile population change in the model is -0.06 to 2.47 percent, with individual city gains up to 11.77 percent (most sizable in Buenos Aires and Caribbean ports) and losses up to -0.2 percent (most negative in cities whose relative economic centrality declined, such as Veracruz and Cartagena). The decomposition of the core-fringe differential shows: (a) location fundamentals (A and H, i.e. productivity and land endowments) explain part of the differential - equalizing fundamentals reduces the gap from 1.64 to 1.11 percentage points; (b) initial population size contributes little - equalizing pre-reform population shares barely moves the gap (from 1.64 to 1.65); (c) the spatial incidence of the shock explains most of the differential - equalizing the size of the transportation-cost reduction across all cities virtually eliminates the gap (to 0.16 percentage points), because the fringe was more trade-restricted before the reform and thus received a larger absolute reduction in transportation times.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper runs extensive robustness exercises. On the reduced-form side: (1) Dynamic event-study specifications show no significant pre-trends before 1750 for the full sample, sub-samples by city size and macro-region, and for settlements. (2) Synthetic control method: treating cities as a group, the synthetic control closely matches pre-reform population trends, with a divergence beginning in 1800 and an implied treatment effect of approximately 19 percent (0.338 log points) over a century, and the true treatment group has the highest post/pre-RMSE ratio relative to all placebo assignments. (3) Dropping outliers (cities outside the 5th-95th percentile of pre-reform population growth rates): coefficients remain around -0.018. (4) Weighting by population size: coefficients similar at around -0.024. (5) Spatial standard errors following Conley (1999): results hold. (6) Robustness value analysis following Cinelli and Hazlett (2020): a confounder would need to explain 16.9 percent of both outcome and treatment variation to fully account for the effect, and 7 percent to render it statistically insignificant - both larger than the combined R2 of observable fundamentals. (7) Interior cities only (distance to port above median): coefficient around -0.016, similar to baseline. (8) Dropping the viceroyalties of Nueva Granada and Rio de la Plata (formed in the 18th century): similar coefficients. (9) Estimating only through 1800 to exclude independence-era effects: point estimates similar for smaller cities. (10) Alternative transportation cost measures including a simple distance measure, showing qualitative robustness. On the model and counterfactual side: (1) Alternative values of the elasticity of substitution (sigma 3-7), Frechet shape parameter (theta 2-4), land expenditure share (1-mu: 0.4-0.6), and agglomeration parameters (a1 in [0.04, 0.07], a2 in [0.02, 0.07]) all yield qualitatively similar results. (2) Alternative trade-cost elasticity from Baum-Snow et al. (2018): similar. (3) Incorporation of national borders after independence (15 percent additional trade cost for cross-border flows): similar, somewhat larger effects. (4) Secular productivity improvements and secular declines in transportation costs (0.88 percent per year starting 1800 per Harley 1988): average effects similar or larger than baseline.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The paper connects to four main strands. First, the literature on history dependence in economic geography. Davis and Weinstein (2002) use WWII bombing shocks to show that Japanese cities return to their pre-shock size, highlighting persistence from locational advantages. Bleakley and Lin (2012) find that US portage sites retain elevated population density long after canals made them obsolete, a classic multiple-equilibria story. Redding, Sturm and Wolf (2010) exploit German division and reunification to show airports exhibit path dependence. Michaels and Rauch (2018) compare Roman and non-Roman cities in France, finding Roman legacy persists. This paper contributes by using a large-scale historical policy reform that changed the location of trading opportunities itself - controlling for time-invariant location fundamentals by construction - and showing that adaptation occurs but is contingent on initial urbanization levels. Henderson et al. (2018) use cross-country data to show locational advantages governing trade matter less in early developers (countries that developed under high transportation costs). This paper supports that cross-sectional finding and gives it a causal interpretation within a single institutional setting. Second, the literature on transportation costs and income. Frankel and Romer (1999) and Feyrer (2019) find large reduced-form effects. Pascali (2017) uses steamship diffusion and finds little aggregate effect except in countries with inclusive institutions - this paper focuses within countries (single institutional environment) and finds robust effects on the spatial distribution rather than aggregate national income. Third, the historical institutions literature. Acemoglu, Johnson and Robinson (2002) establish that pre-industrial population density negatively predicts current income (reversal of fortune). This paper reframes that as partly attributable to trade institutions, showing that Bourbon-era reforms interacted with pre-existing geography to shape the reversal. Fourth, the literature on 18th-century Spanish empire reforms. Valencia (2019), Alvarez-Villa and Guardado (2020), Arteaga (2022), and Chiovelli et al. (2024) examine Bourbon administrative and ecclesiastical reforms. This paper is distinct in focusing on commercial policy and in constructing time-varying bilateral transportation time matrices rather than relying on cross-sectional variation.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central policy implication is that trade liberalization that reduces access costs to long-distance markets can reshape the spatial distribution of economic activity within a country, particularly benefiting peripheral regions that were previously excluded from international trade networks. However, this finding comes with important scope conditions. First, the magnitude of the effect is larger in locations with small pre-existing internal markets. Regions with larger pre-existing urban agglomerations are relatively insulated from reallocation because their size makes them less dependent on external trading opportunities. Policy interventions that reduce international trade costs may therefore have limited spatial rebalancing effects in already-urbanized contexts. Second, the adaptation is not a rapid reallocation: the estimated 2 percent population gain per day-reduction in transportation time reflects cumulative adjustment over 50-year periods. Third, the historical context involves extractive colonial institutions. The paper notes that lower transportation costs influenced spatial development within countries even under extractive institutions, suggesting the result does not require inclusive institutions - but the magnitude and form of adjustment may differ under different institutional regimes. Fourth, migration plays an important amplifying role: in the model, restricting trans-Atlantic migration halves to nearly eliminates the population effect. In modern contexts where immigration is restricted, the spatial reallocation effect of trade liberalization may be substantially smaller. Fifth, the author cautions that the reform involved abrupt, large changes in trade costs, which may produce different adjustment dynamics than gradual reductions.&lt;/p&gt;
&lt;h3 id="q7-how-is-the-transportation-network-constructed-and-validated"&gt;Q7. How is the transportation network constructed and validated?&lt;/h3&gt;
&lt;p&gt;Maritime transportation times are estimated by regressing daily sailing speed (in knots) from 188,687 logbook entries (after removing implausibly fast observations above 10 knots, anchored ships, steamships, and coastal entries) on wind speed and the cosine of the angle between direction of travel and wind direction. The model is estimated on a training sample (179,255 entries) and validated on a holdout sample (9,432 entries), yielding a mean squared error of 2.16. Fitted sailing speeds are then extrapolated to a 0.16 x 0.16 degree global grid using modern wind data from NOAA&amp;rsquo;s Global Forecasting System (2011-2017), assuming wind patterns are sufficiently stable (the correlation between historical logbook wind speed and modern wind speed is 0.24; for wind direction, 0.33). The Dijkstra algorithm finds time-minimizing routes through this grid. Land transportation is modeled using a Tobler-style hiking function adjusted for slope, elevation, and landcover, based on the Weiss et al. (2018) parameterization, applied to a 0.16-degree land grid with postal route locations from Stangl (2019b) treated as roads. Validation compares maritime times to seadistances.org sailing times across 21 ports (strong positive correlation), and land times to the Human Mobility Index and Google Maps driving times (again strongly correlated). The transportation time to Europe from city i in period t is defined as the minimum over the set of ports open to direct trade at time t of the sum of the inland travel time to the nearest open port plus the maritime travel time from that port to Cadiz.&lt;/p&gt;
&lt;h3 id="q8-how-are-the-spatial-models-parameters-identified-and-what-are-the-key-parameter-values"&gt;Q8. How are the spatial model&amp;rsquo;s parameters identified and what are the key parameter values?&lt;/h3&gt;
&lt;p&gt;The model has six parameters (sigma = elasticity of substitution, theta = Frechet shape parameter for migration, mu = expenditure share on traded goods, b = preference shifter for transatlantic goods, a1 = static agglomeration externality, a2 = dynamic/historical agglomeration externality), two vectors of location fundamentals (A and H), and time-varying trade and migration cost matrices (T and M). Sigma is set to 5 following Simonovska and Waugh (2014). Theta is set to 3.18 following Bryan and Morten (2019). Mu = 0.5 is the midrange estimate of the land income share for colonial Mexico and Peru from Arroyo Abad and van Zanden (2016). a1 = 0.055 is taken from the mid-range of estimates in Combes and Gobillon (2015). The trade cost elasticity with respect to transportation time (kappa) is estimated from a port-level gravity model of Spanish imports from Spanish America (1797-1820) using PPML with viceroyalty fixed effects, yielding a transportation time elasticity of trade flows of -2.23, which gives kappa = 0.56. The preference shifter b = 0.45 is chosen to match the observed Spanish import share from the Americas in 1750 (approximately 25 percent per Prados de la Escosura and Casares 1983). The migration cost elasticity lambda is estimated similarly from migration gravity, yielding -lambda*theta = -1.16, so lambda = 0.363. The dynamic agglomeration parameter a2 = 0.063 is identified by estimating the structural version of the reduced-form city-size equation (regressing log population on log price index, log real income, lagged log population, and location controls), where the coefficient on lagged population identifies a2 via the model&amp;rsquo;s equilibrium conditions. Location fundamentals A and H are recovered by inverting the model to exactly match the observed population distribution and nominal wages in 1750.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-paper-contribute-to-understanding-of-the-reversal-of-fortune-in-the-americas"&gt;Q9. What does the paper contribute to understanding of the &amp;lsquo;reversal of fortune&amp;rsquo; in the Americas?&lt;/h3&gt;
&lt;p&gt;Acemoglu, Johnson and Robinson (2002) established that areas with higher pre-industrial (circa 1500) population density tend to have lower income today, interpreting this as evidence that Spanish colonization was most extractive in densely populated areas (which later fell behind) and that sparser-populated frontier areas had better institutions (property rights) that supported later development. This paper complements that institutional story by showing that trade institutions also matter for explaining the reversal. The Bourbon reform - driven by dynastic change from Habsburg to Bourbon rule and by European interstate competition - specifically opened direct trade access to peripheral areas that had been systematically excluded under the Habsburg mercantilist system. The paper&amp;rsquo;s persistence results (lower elasticity of contemporary to pre-colonial population density in areas more exposed to the reform) suggest that the trade reform contributed to the subsequent relative rise of peripheral regions. The finding thus supports the view that the reversal of fortune is partly rooted in institutional change (trade liberalization) interacting with pre-existing geography, rather than in population-density-determined institutions alone. The scope condition is important: the core-versus-fringe heterogeneity shows the reform&amp;rsquo;s spatial effects were largest precisely in the sparsely populated periphery - consistent with the Acemoglu et al. mechanism but augmenting it with a trade-access channel.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-limitations-of-the-analysis-acknowledged-by-the-author"&gt;Q10. What are the limitations of the analysis acknowledged by the author?&lt;/h3&gt;
&lt;p&gt;The author acknowledges four main limitations. First, the reform involved sizeable and abrupt changes in trade costs. More gradual liberalizations might produce different adjustment dynamics, potentially slower convergence or different spatial sorting. Second, the absence of individual-level migration data prevents a more direct examination of whether the city-population effects operate primarily through trans-Atlantic immigration, intracolonial migration, or natural population growth. The model-based inference that trans-Atlantic migration matters substantially is indirect. Third, path dependence likely plays a more important role in industrialized contexts with stronger agglomeration economies (larger a2 than estimated here). The pre-industrial colonial setting, with relatively modest agglomeration forces and thin labor markets, may not generalize to modern industrialized spatial economies. Fourth, the study&amp;rsquo;s focus on within-country (within-empire) variation means it cannot directly address the effect of trade liberalization on aggregate national income, only on the spatial distribution of activity within the empire.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-significance-of-the-finding-that-the-reform-primarily-affected-city-size-rather-than-frontier-settlement"&gt;Q11. What is the significance of the finding that the reform primarily affected city size rather than frontier settlement?&lt;/h3&gt;
&lt;p&gt;The settlement-level analysis uses a balanced panel of 53,581 grid-cell-decade observations for 1710-1810, with an indicator for whether a cell contains any settlement. The baseline result is that a ten-day increase in transportation time to Europe reduces the probability of a cell containing a settlement by one percentage point, against a sample mean of 11 percent, and this effect is small relative to the urban population effects. Event-study plots for settlement formation show no significant pre-trends and only modest post-reform effects. This implies that the reform&amp;rsquo;s primary spatial impact was to concentrate more people in existing urban centers rather than to push economic activity into entirely new locations. This is consistent with the model, in which cities have pre-existing productivity advantages (embedded in A and H) that make them focal points for agglomeration. It also implies the reform did not create entirely new urban systems in frontier areas but rather amplified existing ones, which is important for interpreting the persistence results: even in the fringe, the settlements that grew were already established before 1765.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Transportation time to Europe&lt;/strong&gt;: The time-minimizing route from a given location in Spanish America to Cadiz (the dominant European trading port), computed by combining maritime sailing speed estimates (from logbooks, conditional on wind speed and direction) and land travel speed estimates (based on slope, elevation, landcover, and road location) via the Dijkstra algorithm; time-varying because the set of ports permitted to trade directly with Europe changes as the reform proceeds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Comercio Libre (free trade reform)&lt;/strong&gt;: The staggered series of Spanish royal decrees between 1765 and the early 19th century that progressively lifted the mercantilist restriction confining direct transatlantic trade to four American ports and a single Spanish port, ultimately opening more than 45 American ports to direct trade with Europe; motivated by European interstate competition rather than by the commercial potential of specific American locations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic agglomeration externality (a2)&lt;/strong&gt;: In the Allen-Donaldson (2022) framework as applied here, the component of city-level total factor productivity that depends on the city&amp;rsquo;s own population in the previous period rather than the current period; it encodes the idea that historically larger cities are persistently more productive through channels such as durable local infrastructure, accumulated local knowledge, or input-sharing networks. Estimated at a2 = 0.063 in this setting, smaller than values found in Allen and Donaldson (2022).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;First-nature fundamentals&lt;/strong&gt;: Time-invariant geographic endowments that determine a location&amp;rsquo;s intrinsic productivity and land availability independent of the scale of economic activity, captured in the model by the vectors A (productivity) and H (arable land); these are recovered by inverting the spatial model to match observed 1750 population and wages and are correlated with caloric potential, elevation, terrain ruggedness, and proximity to rivers and coasts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second-nature fundamentals&lt;/strong&gt;: The agglomeration forces that arise from the scale of economic activity already present at a location, including static (current population) and dynamic (lagged population) agglomeration economies; in this paper, the term is used to explain why larger pre-reform cities in the core are insulated from the trade reform&amp;rsquo;s spatial reallocation effects - their scale generates internal-market advantages that reduce reliance on long-distance external trade.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Internal market size (market insulation)&lt;/strong&gt;: The degree to which a city&amp;rsquo;s price index for traded varieties is determined by local production rather than external trade costs; in the model, cities with larger local productivity (higher Ait) have a less sensitive price index to changes in the trade cost with Europe because local goods compete with imported varieties, dampening the welfare and migration effects of trade liberalization; this is the central mechanism explaining the core-fringe heterogeneity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Persistence elasticity&lt;/strong&gt;: The coefficient relating contemporary (year 2000) population density or size to pre-reform (1500 or 1750) population density or size in a cross-sectional regression, interpreted as a measure of how much historical settlement patterns predict current ones; found to be 0.866 for cities with below-median changes in transportation time (little treated) and 0.369 for cities with above-median changes (strongly treated), documenting that the reform attenuated the persistence of pre-reform settlement patterns.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Migration-facilitation channel&lt;/strong&gt;: The mechanism by which lower transportation times to Europe reduce not only trade costs but also migration frictions - through lowering the direct cost of travel and through improving information flows about opportunities in American cities - thereby amplifying city population growth beyond the pure trade-cost effect; quantified in the model by comparing counterfactuals that allow migration frictions to decline with those that hold them fixed at 1760 levels.&lt;/p&gt;</description></item><item><title>Macroeconomic Effects of 'Free' Secondary Schooling in the Developing World</title><link>https://macropaperwarehouse.com/papers/macroeconomic-effects-of-free-secondary-schooling-in-the-developing-world/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/macroeconomic-effects-of-free-secondary-schooling-in-the-developing-world/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether publicly funded (&amp;ldquo;free&amp;rdquo;) secondary schooling in developing countries raises GDP per capita. The question is policy-relevant because many low-income countries — including Ghana, Kenya, Tanzania, Uganda, and others listed in the paper&amp;rsquo;s appendix — have recently adopted or are considering such policies, motivated by the combination of low secondary enrollment (roughly one-third of secondary-school-age children enrolled in the poorest countries, versus near-universal enrollment in rich countries) and evidence that credit constraints keep talented students out of school.&lt;/p&gt;
&lt;p&gt;The analysis is built around an overlapping-generations (OLG) model with heterogeneous households and credit constraints, estimated to match experimental evidence from a randomized controlled trial (RCT) in Ghana (Duflo, Dupas, and Kremer, 2021). The RCT randomly offered full four-year scholarships covering 100 percent of tuition and fees to approximately two thousand poor but high-ability students who had passed the Basic Education Certificate Examination (BECE) but had not enrolled in Senior High School (SHS). Scholarship winners were 27 percentage points more likely to complete secondary school than the control group, scored 0.16 standard deviations (equivalent to 7.6 percent wage gains in the model) higher on math and literacy tests, and experienced a 10.6 percent decline in fertility after 12 years.&lt;/p&gt;
&lt;p&gt;The model departs from standard human capital OLG models in three ways. First, it incorporates an explicit opportunity cost of schooling: teenagers who attend SHS forgo labor income during ages 15–19, which is economically significant given that secondary-school-age individuals are near their prime working years in developing countries. Second, the model includes a merit-based entrance exam (the BECE), so that removing the exam requirement as part of free schooling causes negative selection — the new marginal students induced to attend have lower average ability than those already attending. Third, the model features education-dependent fertility: more-educated households have fewer children (estimated fertility of 2.07 per less-educated family vs 1.19 per more-educated family, in line with Ghanaian Demographic and Health Survey data). The model also incorporates imperfect substitutability between skilled and unskilled labor (elasticity of substitution set to 4, following long-run cross-country estimates), savings wedges that match low liquid asset holdings, and Ghana&amp;rsquo;s actual progressive income tax schedule.&lt;/p&gt;
&lt;p&gt;The model is estimated using the Simulated Method of Moments (SMM) targeting ten moments — five non-experimental (aggregate population growth rate of 2.2 percent per year, aggregate SHS completion rate, SHS completion in the top and bottom test-score quartiles of the control group, and variance of the permanent component of log wages) and five experimental or quasi-experimental (RCT treatment effects on human capital, fertility, overall SHS completion, the Q4 vs Q1 difference in SHS completion, and the intergenerational schooling correlation from administrative data).&lt;/p&gt;
&lt;p&gt;The central quantitative finding is that nationwide free secondary schooling — eliminating both fees and the entrance-exam requirement — raises secondary school completion by about 12 percentage points (from 30 percent to 42 percent of the population) but reduces GDP per capita by approximately 1 percent in the long run. The 95 percent confidence interval for the GDP effect excludes any positive value (lower bound -4.2 percent, upper bound -0.7 percent), so the model can statistically reject any positive GDP impact. The direct fiscal cost of the policy is 1.4 percent of GDP, implying a total cost (direct cost plus lost GDP) of approximately 2.4 percent of GDP. Taxes per capita increase by 1.4 percent. Adult earnings rise by about 1.2 percent, but this is more than offset by a 7.5 percent decline in child earnings (the opportunity cost of schooling for newly enrolled students). The skilled-to-unskilled wage ratio falls by about 10 percent, reflecting general-equilibrium wage compression from the expanded supply of secondary graduates.&lt;/p&gt;
&lt;p&gt;Three counterfactual experiments decompose the negative GDP result. (i) Eliminating the opportunity cost of schooling reverses the GDP effect from -1.0 percent to +2.9 percent, a swing of nearly 4 percentage points — the dominant channel. (ii) Holding the ability distribution of new secondary attendees to match the experimental sample (removing negative selection) moves GDP from -1.0 percent to essentially 0, accounting for about 1 percentage point of the gap. (iii) Holding fertility constant for new secondary attendees moves GDP from -1.0 percent to +1.2 percent, contributing about 2.2 percentage points. When all three channels are shut down simultaneously, GDP rises by 6.9 percent — close to the naive back-of-the-envelope projection of 6 percent based on the RCT&amp;rsquo;s test-score estimates.&lt;/p&gt;
&lt;p&gt;As a policy comparison, an economy-wide improvement in schooling quality that raises test scores by 0.1 standard deviations (a conservative estimate consistent with randomized teacher-incentive interventions in India and Kenya) raises GDP per capita by 2.7 percent and increases SHS completion by 13.8 percentage points — more than free schooling and at lower fiscal cost (the policy pays for itself in equilibrium). Improving schooling quality avoids the negative selection and opportunity-cost channels because it raises human capital for both new and inframarginal students.&lt;/p&gt;
&lt;p&gt;On welfare and distribution, the policy is predominantly redistributive. The bottom 25 percent of parents gain welfare equivalent to a 7.3 percent increase in lifetime consumption, while the top 25 percent lose 4.2 percent. For children, the bottom 25 percent gain 23 percent in consumption-equivalent welfare, while the top 75 percent lose about 5.3 percent. These distributional predictions are validated against a new nationally representative survey of 3,500 Ghanaian households (conducted by the authors in August–September 2022): households with at most a JHS education were 3.1 percentage points more likely to support the policy than average, while those with SHS education or more were 5.2 percentage points less likely — remarkably close to the model&amp;rsquo;s predicted values of 2.6 and 5.9 percentage points, respectively. The authors conclude that free secondary schooling in developing countries is primarily a redistributive policy and not an efficient path to economic growth at current levels of schooling quality.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses a two-step strategy. First, it estimates the OLG model using SMM, with the experimental moments from Duflo, Dupas, and Kremer&amp;rsquo;s (2021) RCT serving as the key identifying variation. The RCT randomly assigned scholarships to poor but high-ability students in Ghana who had passed the BECE but had not enrolled in SHS, making the treatment effect on schooling completion, test scores, and fertility credibly causal in partial equilibrium. Second, the estimated model is used to compute general-equilibrium counterfactuals for a nationwide policy. The main threats to validity are: (a) external validity of the RCT sample to the general population — the sample is explicitly &amp;lsquo;smart kids from poor families,&amp;rsquo; which the authors account for through the negative-selection counterfactual; (b) the model misses on the intergenerational schooling correlation (model: 0.32 vs data: 0.45) and on the treatment effect on SHS completion (model: 21.3 pp vs data: 27 pp), though the authors show in Appendix C that forcing the model to match these moments does not reverse the negative GDP conclusion (a 40 percent higher schooling cost parameter yields a -0.8 percent GDP result vs -1.0 percent baseline; a 15 percent higher ability-persistence parameter yields -2.0 percent); (c) abstracting from human capital externalities (Lucas 1988 type spillovers) and crime reduction effects of education — the authors note these omissions but argue the low estimated effects of the policy make them unlikely to matter quantitatively; and (d) partial equilibrium of the RCT itself — the authors assume no general-equilibrium effects of the experiment since it covered only 2,064 students.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the three main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The three channels are (i) opportunity cost — attendees ages 15–19 forgo labor income; (ii) negative selection — removing the BECE requirement means new marginal students have lower average ability than current attendees; (iii) differential fertility — newly educated households reduce fertility, shifting the long-run population distribution toward less-educated (higher-fertility) households, diluting the share of educated workers over time. The paper isolates each channel through sequential counterfactual experiments: (i) is isolated by eliminating the option for ages-15–19 children to work (forcing the choice between schooling and idleness), which raises the GDP effect from -1.0 to +2.9 percent; (ii) is isolated by artificially boosting the ability of new secondary attendees to match the experimental sample&amp;rsquo;s ability distribution, which moves GDP from -1.0 to approximately 0; (iii) is isolated by setting new attendees&amp;rsquo; fertility to the uneducated-household level, which moves GDP from -1.0 to +1.2 percent. The magnitudes reveal that the opportunity cost channel is the largest (approximately 4 pp swing), followed by the fertility channel (approximately 2.2 pp), and then the selection channel (approximately 1 pp).&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Several dimensions of heterogeneity are documented. In the experimental sample, the treatment effect on SHS completion is not particularly skewed toward high-ability students: the difference in treatment effects between the top and bottom test-score quartiles is only 4 percentage points in the data (and 3 in the model), implying broadly similar gains across the ability distribution within the selected sample. In the estimated model&amp;rsquo;s misallocation analysis, the attendance probability plot (Figure 3) shows that the highest-ability children are fairly likely to attend SHS even when born to low-ability parents — suggesting relatively low misallocation in the estimated model compared to the stylized high-misallocation case. On welfare, the paper documents large heterogeneity by income quartile: the bottom 25 percent of parents gain 7.3 percent in consumption-equivalent welfare while the top 25 percent lose 4.2 percent; for children the bottom 25 percent gain 23 percent while the top 75 percent lose about 5.3 percent. Welfare also differs across generations: gains for grandchildren who always exist are smaller (9 percent) than for children (12 percent), reflecting the compounding fertility effect. The survey confirms these patterns across urban/rural, male/female, and across the Volta (42.3 percent average support for free SHS) and Ashanti (78.2 percent average support) regions of Ghana.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The authors report three robustness checks in Appendix C. First, they increase the schooling cost parameter ΨS by 40 percent to force the model to match the (currently undershot) treatment effect on SHS completion; the free schooling policy then produces a -0.8 percent GDP result (vs -1.0 percent baseline) and a 14 percent increase in attendance (vs 12 percent baseline) — the conclusion is unchanged. Second, they increase the ability-persistence parameter ρ by 15 percent to match the intergenerational schooling correlation; the result is a -2.0 percent GDP decline and a 4 percent attendance increase — the GDP decline is larger, so if anything the baseline is too generous to free schooling. Third, they experiment with lower values of the elasticity of substitution between skilled and unskilled labor (down to 1.4 from the baseline value of 4) and report no substantive change in conclusions. The authors also use bootstrapped 95 percent confidence intervals for all aggregate predictions, which is unusual in general-equilibrium counterfactual exercises in macroeconomics.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does the paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The paper is most closely related to Abbott, Gallipoli, Meghir, and Violante (2019) and Daruich (2020), both of which study public education expansions in the United States and find largely positive effects on GDP and welfare. The authors argue the contrast with their pessimistic findings reflects lower school quality in developing countries — in a rich-country setting, opportunity costs are lower relative to the returns to schooling. Hendricks and Schoellman (2014) find similar negative selection of college students in the US as enrollment expands, lending support to the selection channel. Khanna (2023) documents substantial declines in the relative wages of skilled workers after an education expansion in India, consistent with the model&amp;rsquo;s 10 percent skilled-to-unskilled wage compression, though Khanna&amp;rsquo;s short-run effects are larger due to lower short-run elasticity of substitution. In terms of methodology, the paper follows Daruich (2020) in using RCT evidence to discipline an OLG model, and is the first paper to do so for the macroeconomic effects of education policy in the developing world. The paper also builds on the macro-development literature emphasizing school quality (Hanushek and Woessmann, 2007; Schoellman, 2012) over average years of schooling as the proximate cause of low human capital in poor countries.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central policy implication is that free secondary schooling in developing countries, at current low levels of schooling quality, is primarily redistributive rather than growth-enhancing. Countries considering free schooling should expect secondary enrollment to rise substantially (by around 12 percentage points in the baseline) but GDP per capita to fall or stay flat. The alternative of improving schooling quality — modeled as a 0.1 standard deviation increase in test scores, using teacher incentives or additional teachers at a cost of approximately US$5.78 per student per year (based on Mbiti et al. 2019 in Tanzania) — raises GDP by 2.7 percent and schooling enrollment by even more (13.8 percentage points), while paying for itself in equilibrium. A key scope condition: the negative GDP finding is driven by the combination of high opportunity costs of schooling (secondary-school-age workers have economically significant labor income in developing countries), negative selection from removing merit requirements, and low schooling quality that limits the human capital return per year of schooling. In rich countries where these conditions do not hold, the same policy has been found to be beneficial. The paper also shows (Table 6) that maintaining the entrance-exam requirement alongside free schooling substantially mitigates the GDP decline (-0.3 percent vs -1.0 percent), and that keeping both the test and a positive fee results in approximately zero GDP change — suggesting that the test-requirement component of the policy design is important.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-paper-find-about-misallocation-in-the-estimated-model"&gt;Q7. What does the paper find about misallocation in the estimated model?&lt;/h3&gt;
&lt;p&gt;The estimated model exhibits relatively low misallocation. The misallocation concept refers to situations where high-ability children of poor parents are kept out of secondary school by borrowing constraints even though the net-present-value of additional schooling exceeds the cost. The paper shows (Figure 2) that economies can have similar aggregate secondary enrollment rates of around 30 percent but very different degrees of misallocation — one where enrollment is low because returns are low (low-misallocation case), and one where enrollment is low because high-ability children are credit-constrained (high-misallocation case). The estimated model falls closer to the low-misallocation case (Figure 3), with the highest-ability children fairly likely to attend SHS even if born to low-ability parents. This finding is consistent with the modest increase in SHS completion induced by free schooling (12 percentage points) relative to the experimental treatment effect on the selected sample (27 percentage points): most high-ability children are already attending, so there is limited room for a free schooling policy to reduce misallocation.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-welfare-analysis-reveal-about-the-puzzle-of-large-welfare-gains-alongside-a-gdp-decline"&gt;Q8. What does the welfare analysis reveal about the puzzle of large welfare gains alongside a GDP decline?&lt;/h3&gt;
&lt;p&gt;The paper documents an apparent puzzle: the free schooling policy reduces long-run GDP per capita by 1 percent but produces large positive welfare gains for parents (average 3.9 percent in consumption-equivalent welfare) and even larger gains for children (average 12.4 percent). The resolution is that (a) welfare gains for parents come entirely from redistribution — the very poor gain 7.3 percent while the rich lose 4.2 percent, and the progressive tax schedule is the mechanism; (b) the welfare gains for the children&amp;rsquo;s generation partially reflect large gains to the small number of previously misallocated children who now attend secondary school (the bottom 25 percent of children gain 23 percent, primarily through income gains for those who previously could not afford school); and (c) these gains erode across generations — grandchildren who always exist gain less (9 percent vs 12 percent for children), because the grandchildren who would only have existed without the free schooling policy (i.e., the &amp;lsquo;unborn&amp;rsquo; due to reduced fertility among educated households) would have experienced disproportionately large gains (almost 17 percent). The composition of the population thus shifts toward those experiencing smaller gains, compounding over generations and producing the long-run GDP decline.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-role-of-the-entrance-exam-design-in-free-schooling-policy-outcomes"&gt;Q9. What is the role of the entrance exam design in free schooling policy outcomes?&lt;/h3&gt;
&lt;p&gt;The paper shows that how access is structured matters as much as whether schooling is free. In the main analysis, free schooling eliminates both fees and the BECE entrance requirement, consistent with Ghana&amp;rsquo;s 2017 policy. In alternative simulations (Table 6), free schooling that maintains the existing entrance requirement (a &amp;lsquo;relaxed test&amp;rsquo; policy) produces a GDP decline of only -0.3 percent instead of -1.0 percent. Free schooling that keeps the test at full stringency (so fewer new students gain access) produces essentially no change in GDP (-0.0 percent), but also a much smaller increase in secondary attendance (3.0 pp vs 11.8 pp). Eliminating only the test requirement while keeping a positive fee produces a -0.4 percent GDP decline. These results confirm that the negative selection channel is a quantitatively important driver of the adverse GDP effect and is specifically activated by the removal of the merit requirement.&lt;/p&gt;
&lt;h3 id="q10-how-is-the-model-estimated-and-what-moments-does-each-parameter-primarily-identify"&gt;Q10. How is the model estimated and what moments does each parameter primarily identify?&lt;/h3&gt;
&lt;p&gt;The model is estimated by SMM minimizing the sum of squared differences between model moments and their data counterparts, using a vector of 10 parameters (fertility parameters νJ and νS; schooling efficiency ηS; goods cost of schooling ΨS; intergenerational altruism b; exam score noise σε; Gumbel taste-shock scale θ; savings wedge χ; ability persistence ρ; ability shock standard deviation συ). Six parameters are chosen directly from the literature or normalization (A, α, β, r*, λ, σζ). Ten moments are targeted: population growth rate (primarily identifies νJ, νS), aggregate SHS completion rate and quartile completion rates (identify ηS, b, ΨS, χ), variance of the permanent component of wages (identifies συ, ρ), and five experimental moments from the Duflo et al. RCT (treatment effects on human capital, fertility, SHS completion, the Q4–Q1 completion difference, and the intergenerational schooling correlation). Confidence intervals are bootstrapped by re-sampling the five experimental moments 100 times, treating the non-experimental moments as fixed. The Jacobian matrix (Appendix Table C.1) and sensitivity matrix (Appendix Table C.2) are computed following Kaboski and Townsend (2011) and Andrews, Gentzkow, and Shapiro (2017) to document identification.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-survey-design-details-and-how-well-does-it-validate-the-model"&gt;Q11. What are the survey design details and how well does it validate the model?&lt;/h3&gt;
&lt;p&gt;The authors conducted a new nationally representative household survey in Ghana in August–September 2022, covering 3,500 households selected via two-stage cluster sampling from seven regions accounting for about 61 percent of the Ghanaian population. Respondents were asked whether eight categories of government expenditure should be abolished, cut substantially, cut somewhat, maintained, or expanded. For free SHS, respondents with at most a JHS education were 3.1 percentage points more likely to support the policy than average; those with SHS education or more were 5.2 percentage points less likely. These empirical patterns align closely with the model&amp;rsquo;s predicted values of 2.6 and 5.9 percentage points respectively. The pattern is robust across urban/rural subsamples, male/female subsamples, and across the Volta and Ashanti regions (which differ substantially in overall support levels — 42.3 percent vs 78.2 percent — but maintain the same qualitative pattern of lower-educated households being more supportive). The one discrepancy is that the model over-predicts the support of JHS-educated households who have children enrolled in SHS.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Opportunity cost of schooling&lt;/strong&gt;: In this paper&amp;rsquo;s model, the foregone labor income of teenagers aged 15–19 who attend secondary school rather than work. This cost persists even when the school fee is eliminated by government policy and is identified as the single largest channel explaining why free secondary schooling reduces rather than raises GDP per capita in developing countries, contributing approximately 4 percentage points to the adverse GDP effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Negative selection of new students&lt;/strong&gt;: The reduction in average ability of the marginal students who enter secondary school once both fees and the merit-based entrance exam are eliminated. The existing pool of secondary attendees was positively selected by the entrance exam, so broadening access induces a lower-ability pool of new entrants, reducing the average human capital gain per new graduate. The paper estimates this channel accounts for approximately 1 percentage point of the adverse GDP gap relative to the back-of-the-envelope projection.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Differential fertility by education&lt;/strong&gt;: The model feature by which secondary-educated households have significantly fewer children (parameter νS = 0.19 implying 2.4 children per family) than non-secondary-educated households (νJ = 1.07 implying 4.1 children per family). When free schooling induces more households to obtain secondary education, aggregate fertility falls, and crucially the share of high-ability households in the long-run population declines because those households now have fewer children, reducing the long-run supply of educated workers and contributing approximately 2.2 percentage points to the adverse GDP gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Misallocation of talent&lt;/strong&gt;: In this paper&amp;rsquo;s sense: the situation in which high-ability children of poor parents are prevented by borrowing constraints from attending secondary school even though the net-present-value of additional schooling exceeds the combined goods and opportunity costs. The paper finds that the estimated model of Ghana corresponds more closely to a low-misallocation economy (Figure 3), meaning the highest-ability children attend SHS at fairly high rates regardless of parental income, so the scope for free schooling to reduce misallocation is limited.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Balanced growth path&lt;/strong&gt;: In this paper: a recursive competitive equilibrium in which aggregate population grows at a constant rate while the relative distribution of households across individual states (ability, education, assets) is stationary, and household policy functions are independent of the aggregate population level. All policy counterfactuals are conducted by introducing a policy into the balanced growth path and computing transition dynamics to the new balanced growth path.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Schooling quality (ηS)&lt;/strong&gt;: The efficiency parameter governing how much human capital a student of given ability acquires from a year of secondary schooling, defined in the production function h(z,S) = z · ηS. In the estimated model, ηS = 5.66, implying an annual return to education of 7.9 percent for the experimental sample. The paper shows that a policy raising ηS (schooling quality) by enough to increase average test scores by 0.1 standard deviations raises GDP by 2.7 percent and expands SHS enrollment by 13.8 percentage points, outperforming free schooling on both counts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Savings wedge (χ)&lt;/strong&gt;: A wedge between the international market rate of return on capital (r*) and the return available to households in the model (r = r* - χ), calibrated to match the low savings rates observed in low-income economies. In the estimated model χ = 0.09, implying households earn approximately 2 percent per year on savings. Together with the borrowing constraint (no borrowing against children&amp;rsquo;s future income), this ensures that poor parents cannot save their way out of the constraint preventing them from sending high-ability children to school.&lt;/p&gt;</description></item><item><title>Macroeconomic Effects of Public R&amp;D</title><link>https://macropaperwarehouse.com/papers/macroeconomic-effects-of-public-rd/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/macroeconomic-effects-of-public-rd/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper estimates the dynamic macroeconomic effects of US government R&amp;amp;D investment using a Structural Vector Autoregressive (SVAR) framework, with an extension to a Rational Expectations SVAR (RE-SVAR) that explicitly captures private-sector anticipation of public spending decisions. The central questions are: (1) what is the fiscal multiplier of public R&amp;amp;D spending on GDP and private R&amp;amp;D investment, and how does it compare to other government spending categories; (2) does public R&amp;amp;D crowd in or crowd out private R&amp;amp;D; and (3) how much does the private sector&amp;rsquo;s anticipation of future public R&amp;amp;D commitments amplify these effects?&lt;/p&gt;
&lt;p&gt;The dataset covers 1947Q1–2017Q3 and is drawn from the US Bureau of Economic Analysis, deflated to 2009 prices and expressed in per-capita terms. The five-variable system includes government R&amp;amp;D investment (GI), government residual spending (GG), net taxes (T), private R&amp;amp;D investment (GR), and GDP (Y), all modelled in log-levels to preserve cointegrating relationships. The lag length is set to six quarters (chosen by Hannan-Quinn criterion, consistent with the R&amp;amp;D-to-productivity lag literature). Identification rests on three mild contemporaneous restrictions: (i) government R&amp;amp;D decisions are independent of current-quarter GDP, consistent with their long-term, mission-oriented character; (ii) R&amp;amp;D spending can influence all other government expenditures in the same quarter but not vice versa; (iii) taxes affect government spending contemporaneously but not the reverse. An alternative identification (SVAR model B) reverses the within-quarter tax-spending causality and produces very similar results. The RE-SVAR extends the system by including the expected next-period public R&amp;amp;D shock, identified by assuming perfect foresight of one-quarter-ahead government R&amp;amp;D innovations and an additional restriction that public R&amp;amp;D does not respond to lagged GDP or private R&amp;amp;D.&lt;/p&gt;
&lt;p&gt;Main quantitative findings from the leading estimation (RE-SVAR model A, full sample):&lt;/p&gt;
&lt;p&gt;GDP fiscal multiplier — anticipated shock: within the quarter of implementation (one quarter after the announcement), one dollar of public R&amp;amp;D spending raises GDP by approximately 52 dollars (pure multiplier at t = 0 is 51.59; see Table 2). The multiplier peaks immediately and then declines to roughly 22–24 dollars over a six-year horizon. Critically, this GDP increase is permanent across all SVAR and RE-SVAR specifications, whereas generic government spending produces only a temporary rise.&lt;/p&gt;
&lt;p&gt;GDP fiscal multiplier — unanticipated shock: setting aside the anticipation effect, the impact-period multiplier falls to approximately 13–14 dollars (13 dollars in the scenario with no anticipation), which is still substantially larger than the peak multiplier of roughly 0.73–0.76 dollars for residual government spending (Table 1, SVAR model A).&lt;/p&gt;
&lt;p&gt;Expectations channel: at t = 0, before the actual spending increase occurs at t = 1, the news alone raises GDP by 16.48 dollars. The total peak GDP effect (55.75 dollars) is nearly double the counterfactual effect without the anticipation component (31.64 dollars). The coefficient on expected next-period public R&amp;amp;D in the private R&amp;amp;D equation is 0.58 (p-value 0.035), confirming a statistically significant anticipation channel for private R&amp;amp;D.&lt;/p&gt;
&lt;p&gt;Crowding-in of private R&amp;amp;D: public R&amp;amp;D crowds in private R&amp;amp;D at all horizons. The public-to-private R&amp;amp;D multiplier peaks at 1.81 in the quarter following the news shock (t = 0), and stabilizes at 0.75 after six years — an elasticity of 0.72, close to Moretti et al.&amp;rsquo;s (2021) estimate of 0.52 from production-function methods. At t = 0, private R&amp;amp;D rises by 0.52 in response to the announcement alone.&lt;/p&gt;
&lt;p&gt;Persistence of public spending: a one-dollar public R&amp;amp;D shock keeps GI above 2 dollars six years later, whereas residual government spending returns to baseline within four years. Cumulative total government spending over six years following a one-dollar R&amp;amp;D shock is 220 dollars, versus only 22 dollars for a generic spending increase.&lt;/p&gt;
&lt;p&gt;Output elasticity at longer horizons: the GDP multiplier expressed in elasticity terms is 0.34 one year after the anticipated shock, stabilizing between 0.23 and 0.25 over three to six years. The corresponding range for private R&amp;amp;D (GR shock) is 0.18 to 0.16, broadly consistent with cross-country evidence from Coe-Helpman (1995) and Guellec-van Pottelsberghe (2004).&lt;/p&gt;
&lt;p&gt;The paper argues that the large short-run multipliers reflect three mechanisms that can materialize quickly: (1) process-innovation cost reductions; (2) early entry of private co-investors seeking first-mover advantage; (3) embodiment of new knowledge in physical capital. At longer horizons, supply-side productivity gains and knowledge spillovers dominate. The policy conclusion is that public R&amp;amp;D is unusually effective both as a demand-side stimulus and as a long-run growth instrument, provided government credibly announces and maintains multi-year funding commitments that stabilize private-sector expectations.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The baseline SVAR identification (model A) imposes three contemporaneous exclusion restrictions: government R&amp;amp;D decisions are exogenous to same-quarter GDP and to other fiscal variables (because R&amp;amp;D budgets reflect long-term strategic priorities, not countercyclical reactions); GI can influence GG contemporaneously but not vice versa; and taxes affect spending in the same quarter but not the reverse. A key threat is non-fundamentalness: because public R&amp;amp;D programs are announced well in advance, what appears to the econometrician as a surprise shock is actually largely anticipated by the private sector, biasing the SVAR impulse responses. The paper addresses this by extending the SVAR to a Rational Expectations SVAR (RE-SVAR) that adds the expected next-period GI shock to the information set of private agents, identified by the additional assumption that GI does not respond to lagged GDP or private R&amp;amp;D. A secondary threat is the direction of same-period causality between taxes and spending; an alternative model (SVAR model B) reverses this and finds only minor quantitative differences. The Lucas Critique applies to the counterfactual simulation of an unanticipated shock since the model was estimated under a perfect-foresight assumption.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-re-svar-separate-the-anticipation-effect-from-the-effect-of-the-actual-spending-increase"&gt;Q2. How does the RE-SVAR separate the anticipation effect from the effect of the actual spending increase?&lt;/h3&gt;
&lt;p&gt;The RE-SVAR model includes E[GI_{t+1} | Omega_t] — the expectation of next-period public R&amp;amp;D — as a forward-looking right-hand-side variable in the private R&amp;amp;D and GDP equations. Under the perfect-foresight assumption, this expectation equals the realized next-period structural shock. The IRF for an anticipated GI shock therefore starts at t = 0 when the news arrives and the actual spending rise occurs at t = 1. By comparing (i) the full anticipated IRF (news at t = 0 + realization at t = 1) to (ii) a modified version where the news term is removed from the information set (unanticipated shock), the paper isolates the incremental contribution of expectations. At t = 0 the news alone raises GDP by 16.48 and private R&amp;amp;D by 0.52; the total peak GDP effect with anticipation is 55.75, versus 31.64 without it — a difference of roughly 24 dollars at the one-year horizon.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-main-mechanisms-proposed-to-explain-the-unusually-large-short-run-fiscal-multiplier"&gt;Q3. What are the main mechanisms proposed to explain the unusually large short-run fiscal multiplier?&lt;/h3&gt;
&lt;p&gt;Three channels are proposed for the large immediate GDP response. First, process innovation can reduce production costs without long lags from the start of R&amp;amp;D investment. Second, anticipatory entry of private co-investors seeking first-mover advantages intensifies investment at the very beginning of a research program, even before results are commercialized. Third, innovation embodied in new physical capital means R&amp;amp;D expenditure is accompanied by complementary investment in physical equipment, amplifying the aggregate demand stimulus. At longer horizons, supply-side productivity gains from knowledge spillovers across firms and sectors become the dominant channel. The paper also notes that public R&amp;amp;D programs are frequently accompanied by large-scale complementary government procurement (e.g., defense agency procurements), further magnifying the total mobilization of public resources.&lt;/p&gt;
&lt;h3 id="q4-what-do-the-multipliers-for-residual-government-spending-gg-look-like-and-how-do-they-compare-to-public-rd"&gt;Q4. What do the multipliers for residual government spending (GG) look like, and how do they compare to public R&amp;amp;D?&lt;/h3&gt;
&lt;p&gt;From SVAR model A (Table 1), one dollar of residual government spending raises GDP by 0.73 at t = 0 (also its peak), declining to around 0.45 after six years. The peak private R&amp;amp;D multiplier of GG spending is 0.08 (after six years), rising very slowly from near zero. Compared to the GDP multiplier of public R&amp;amp;D (13.68 at t = 0, peak 16.18), the residual spending multiplier is roughly 20 times smaller. Moreover, the GDP increase from GG spending is temporary, reverting to baseline within four years, while the GDP increase from GI spending is permanent. These contrasts hold across both SVAR models A and B and across the RE-SVAR estimations.&lt;/p&gt;
&lt;h3 id="q5-what-evidence-is-there-for-the-crowding-in-of-private-rd-by-public-rd"&gt;Q5. What evidence is there for the crowding-in of private R&amp;amp;D by public R&amp;amp;D?&lt;/h3&gt;
&lt;p&gt;The paper finds strong, statistically significant crowding-in across all specifications. In the SVAR model A (Table 1), the multiplier of GI on private R&amp;amp;D (GR) reaches its peak of 0.76 after two quarters and remains at 0.41 after six years. In the RE-SVAR model A (Table 2), the anticipated public R&amp;amp;D shock raises private R&amp;amp;D by 1.81 dollars per dollar of public R&amp;amp;D at t = 0, declining to 0.75 after six years, translating to an elasticity of 0.72. Even in the alternative identification (RE-SVAR model B), the result persists, though the peak private R&amp;amp;D multiplier from anticipated GI spending is lower (0.40 after four quarters). The response of private R&amp;amp;D to both its own shock and to public R&amp;amp;D shocks is permanent across all RE-SVAR estimations, supporting the conclusion that public R&amp;amp;D accelerates the total national innovation effort rather than displacing it.&lt;/p&gt;
&lt;h3 id="q6-what-mechanisms-explain-the-crowding-in-of-private-rd"&gt;Q6. What mechanisms explain the crowding-in of private R&amp;amp;D?&lt;/h3&gt;
&lt;p&gt;The paper identifies five complementary channels: (1) Public funding covers large fixed costs (laboratories, human capital), making private research projects profitable that would not otherwise be undertaken. (2) Public R&amp;amp;D removes credit constraints faced by private innovators. (3) Anticipated technological spillovers signal profitable investment opportunities to private firms. (4) The government funding decision itself conveys a signal about the long-run profitability and viability of a research area. (5) The public-private partnership alleviates asymmetric information and the high riskiness that typically deters private R&amp;amp;D. Additionally, transparency in public procurement and entry requirements into publicly funded programs may signal quality, further encouraging private investment.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-conducted-and-what-do-they-show"&gt;Q7. What robustness checks are conducted, and what do they show?&lt;/h3&gt;
&lt;p&gt;Three robustness checks are applied to both the SVAR and RE-SVAR estimations: (i) alternative identification (SVAR model B / RE-SVAR model B) where the contemporaneous causal direction between taxes and government spending is reversed; (ii) a shorter sample excluding the period from the 2008 financial crisis onward (1947Q1–2007Q4); (iii) a longer lag length of eight quarters. For check (i), results are very similar: the GDP multiplier for GI is slightly smaller at short horizons (10.02 vs 13.68 at t = 0 in the SVAR, and 31.19 vs 51.59 at t = 0 in the anticipated RE-SVAR) but converges to similar long-horizon values. For check (ii), the impact of GI on GDP at t = 0 is 15.5 (vs 13.54), with similar hump shape; GI&amp;rsquo;s impact on GR is slightly lower. For the RE-SVAR robustness checks, the paper reports that the shape, timing, and order of magnitude remain stable, as does the finding that the anticipated GI multiplier considerably exceeds the unanticipated one. The general conclusion is no qualitative variation and only minor quantitative differences.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-re-svars-handling-of-the-non-fundamentalness-problem-and-how-is-it-justified-specifically-for-public-rd"&gt;Q8. What is the RE-SVAR&amp;rsquo;s handling of the non-fundamentalness problem and how is it justified specifically for public R&amp;amp;D?&lt;/h3&gt;
&lt;p&gt;Non-fundamentalness arises when the VAR&amp;rsquo;s implied information set is smaller than that of private agents — i.e., what the econometrician calls a surprise is actually anticipated by the economy, so estimated structural shocks are combinations of current and future structural innovations and the fundamental VAR representation is not identified. The paper argues this problem is particularly severe for public R&amp;amp;D because: (1) R&amp;amp;D budgets are part of long-term plans with detailed technical reports and high-profile public announcements (as documented with historical episodes in Section 2); (2) established procurement links between government agencies and private firms provide early information flows. The RE-SVAR addresses this by explicitly adding E[GI_{t+1} | Omega_t] to the system (Blanchard-Perotti approach applied to a non-causal VAR) and assuming perfect foresight of next-period GI innovations. External forecast measures are unavailable for government R&amp;amp;D spending, making this the only viable route. Perfect foresight is defended as particularly appropriate given the highly public, plan-driven nature of government R&amp;amp;D decisions.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q9. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The closest precursors are Deleidi and Mazzucato (2021) and Antolin-Diaz and Surico (2022). Deleidi and Mazzucato use a recursively identified SVAR where defense R&amp;amp;D spending is ordered first and find a first-quarter GDP multiplier of 24 dollars. This paper differs by: (a) using total government R&amp;amp;D (defense + non-defense) rather than only defense R&amp;amp;D; (b) providing a more general and explicitly motivated identification that goes beyond simple recursive ordering; (c) developing the RE-SVAR extension to capture the anticipation channel, which raises the estimated multiplier substantially above 24 dollars. Antolin-Diaz and Surico (2022) study military spending news with a 125-year VAR (60 lags, Bayesian shrinkage) and find a long-run defense spending GDP multiplier of 2.08 and argue that public R&amp;amp;D specifically drives long-run productivity. The present paper uses a shorter but richer five-variable quarterly system with explicit crowding-in measurement. On the crowding-in question, the paper contrasts with earlier work (Goolsbee 1998, Wallsten 2000) finding crowding-out due to inelastic supply of scientists, and aligns with more recent evidence (Becker 2015, Moretti et al. 2021) showing crowding-in once a broader set of mechanisms is accounted for.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Three core policy implications are identified. First, public R&amp;amp;D is a highly effective instrument for stimulating long-run technological innovation and economic growth: the permanent GDP response and the strong private R&amp;amp;D crowding-in indicate that public investment substantially elevates the country&amp;rsquo;s aggregate innovation capacity. Second, fiscal multipliers are class-specific: the multiplier for public R&amp;amp;D dramatically exceeds that for generic government spending, implying that the composition of government expenditure matters greatly for both short-run stabilization and long-run growth. The absence of crowding-out and the large short-run multipliers suggest substantial untapped productive capacity due to market failures in R&amp;amp;D. Third, the anticipation channel is quantitatively important: ignoring private-sector foresight understates the true multiplier, and this implies that the credibility and advance communication of government R&amp;amp;D commitments are themselves policy instruments — long-term, publicly announced programs that stabilize expectations can effectively mobilize private co-investment that would not occur under uncertain or ad hoc spending. Scope conditions: results are estimated on US data 1947Q1–2017Q3, a country with large and heterogeneous federal R&amp;amp;D programs; extrapolation to countries with different institutional settings, R&amp;amp;D compositions, or capital market structures requires caution. The model uses a 1.5-year lag structure that may not fully capture very long-run R&amp;amp;D-to-productivity channels estimated at 5–20 years in micro studies.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-pure-fiscal-multiplier-and-why-does-the-paper-use-it-instead-of-the-standard-multiplier"&gt;Q11. What is the &amp;lsquo;pure fiscal multiplier&amp;rsquo; and why does the paper use it instead of the standard multiplier?&lt;/h3&gt;
&lt;p&gt;Standard fiscal multipliers are calculated by dividing the cumulative IRF of GDP to a unit shock in a given spending category by the cumulative IRF of total government spending to the same shock. The problem is that total spending includes other categories that dynamically respond to the initial shock (e.g., GI shocks cause GG to rise significantly via cross-equation dynamics), so the denominator conflates the effect of GI with the effect of induced GG changes, making multipliers across spending categories incomparable. The paper therefore uses &amp;lsquo;pure multipliers&amp;rsquo; (following Perotti 2004): the counterfactual total government spending is calculated from a version of the SVAR where the dynamics of GG are switched off (all coefficients in the GG equation are set to zero), so the denominator captures only the direct mechanical effect of the GI shock on aggregate spending without the induced cross-spending effects. This allows clean apples-to-apples comparison of one average dollar spent across different categories.&lt;/p&gt;
&lt;h3 id="q12-what-do-long-run-gdp-elasticities-imply-about-the-social-return-to-rd"&gt;Q12. What do long-run GDP elasticities imply about the social return to R&amp;amp;D?&lt;/h3&gt;
&lt;p&gt;Expressed in elasticity terms, the GDP multiplier from an anticipated GI shock is 0.34 one year after implementation and stabilizes at 0.23–0.25 over three to six years. For private R&amp;amp;D (GR shock), the corresponding elasticity is 0.18 after one year, stabilizing at 0.15–0.16. These are broadly consistent with existing cross-country production function estimates: Coe and Helpman (1995) obtain 0.22 for G7 economies; Guellec and van Pottelsberghe (2004) find 0.13 for private and 0.17 for public R&amp;amp;D spending; Ornaghi (2006) finds 0.24 for Spanish firms including spillovers. The paper notes that Jones and Summers (2020) calculate that the social return to innovation can easily generate a GDP effect of 20 dollars per dollar of R&amp;amp;D once the full set of spillovers is captured at the aggregate level, which is consistent with the dollar multipliers obtained here at longer horizons.&lt;/p&gt;
&lt;h3 id="q13-how-does-private-rd-gr-compare-to-public-rd-gi-as-a-gdp-stimulus"&gt;Q13. How does private R&amp;amp;D (GR) compare to public R&amp;amp;D (GI) as a GDP stimulus?&lt;/h3&gt;
&lt;p&gt;In the leading RE-SVAR model A, a unit shock to private R&amp;amp;D raises GDP by 27.65 at t = 0 and reaches a peak of 39.62 after one year, before stabilizing at around 24 dollars after six years. This is slightly below the public R&amp;amp;D effect (peak 55.75 at t = 0, declining to ~38 dollars and eventually ~22 after six years). The short-run superiority of public R&amp;amp;D over private R&amp;amp;D is attributed to: (1) breadth of goals — public programs simultaneously mobilize a wider set of industries; (2) longer planning horizon — reducing uncertainty and encouraging private co-investment; (3) the expectations channel available to public but not private R&amp;amp;D; (4) entry requirements and transparency signaling research quality; (5) government agencies as both funder and user, accelerating knowledge transfer. However, the superiority of public over private R&amp;amp;D is not confirmed in all specifications of the robustness analysis.&lt;/p&gt;
&lt;h3 id="q14-what-historical-evidence-does-the-paper-marshal-to-motivate-the-anticipation-mechanism"&gt;Q14. What historical evidence does the paper marshal to motivate the anticipation mechanism?&lt;/h3&gt;
&lt;p&gt;Section 2 documents several large defense and non-defense R&amp;amp;D programs where public announcements substantially pre-dated actual spending: the Sputnik response (DARPA and NASA created in 1958 following October 1957 Sputnik launch; spending projections published in Business Week months in advance); Nixon&amp;rsquo;s Strategic Nuclear Doctrine (January–February 1974 announcements of record defense budget of 92.6 billion, with Congress extending Pentagon research commitments in June 1975); Reagan&amp;rsquo;s Strategic Defense Initiative (publicly announced March 23, 1983; CBO published detailed multi-year cost projections by May 1984); Kennedy&amp;rsquo;s Moon Mission (announced May 25, 1961; NYT reported cost projections the following day; estimates revised multiple times through 1969); Nixon&amp;rsquo;s War on Cancer (December 1970 Senate report and May 1971 Nixon speech; National Cancer Act passed December 23, 1971 with pre-specified multi-year budget); Human Genome Initiative (DOE announcement March 1986; Department of Health endorsement April 1987; project ran 1990–2013); Obama&amp;rsquo;s Climate Action Plan (energy transition plans mooted from 2009; America COMPETES Acts 2007, 2010, 2014). These examples document both the forward-looking nature of R&amp;amp;D budgeting and the detailed public information available to private agents ahead of actual spending.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Rational Expectations SVAR (RE-SVAR)&lt;/strong&gt;: An extension of the standard SVAR framework that adds a forward-looking expectational variable — specifically the expected next-period public R&amp;amp;D structural shock E[GI_{t+1} | Omega_t] — to the system, allowing the model to capture the influence of private-sector anticipation on current economic outcomes rather than treating all fiscal shocks as surprises.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-fundamentalness&lt;/strong&gt;: A condition arising when the VAR&amp;rsquo;s implied information set is a strict subset of the actual information set of private agents, causing the reduced-form VAR residuals to be non-invertible linear combinations of current and future structural innovations. For public R&amp;amp;D, this means that what the econometrician identifies as a surprise shock to GI is in fact largely anticipated by the private sector, biasing estimated impulse responses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pure fiscal multiplier&lt;/strong&gt;: A class-specific fiscal multiplier calculated by isolating the GDP response to one dollar spent in a given category of government spending while holding other spending categories constant (switching off their dynamics). Contrasts with the standard multiplier, which conflates the direct effect of the shock with induced changes in other spending categories triggered by dynamic cross-equation correlations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mission-oriented spending&lt;/strong&gt;: Government R&amp;amp;D investment directed at achieving long-term strategic national goals (e.g., space exploration, defense superiority, cancer research, climate transition). Defined by three features that distinguish it from generic government expenditure: (i) long-term policy motivation independent of short-run macroeconomic conditions; (ii) advance public announcements that create private-sector expectations; (iii) potential for permanent productivity-level effects through knowledge spillovers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crowding-in&lt;/strong&gt;: In this paper, the phenomenon whereby an exogenous increase in public R&amp;amp;D investment triggers a statistically significant and persistent increase in private R&amp;amp;D investment — the opposite of the crowding-out (substitution) effect posited when an inelastic supply of scientists and engineers constrains total R&amp;amp;D activity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fiscal foresight&lt;/strong&gt;: The ability of private economic agents to predict future government spending decisions ahead of their actual implementation, arising from legislative lags, public announcements, procurement contracts, and established information channels between policy makers and private co-investors. Fiscal foresight makes standard SVAR fiscal shocks non-fundamental and amplifies the macroeconomic impact of spending by triggering anticipatory private responses before the actual dollar is spent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Anticipation channel (expectations effect)&lt;/strong&gt;: The component of the macroeconomic response to public R&amp;amp;D spending that is activated at the time of the public announcement rather than at the time of actual spending. In the RE-SVAR model, this channel accounts for the extra GDP boost of approximately 21 dollars at t = 1 and a peak of 24 dollars after one year, relative to the counterfactual scenario of an unanticipated shock.&lt;/p&gt;</description></item><item><title>Manipulation of information in times of crisis: evidence from Covid excess mortality</title><link>https://macropaperwarehouse.com/papers/manipulation-of-information-in-times-of-crisis-evidence-from-covid-excess-mortality/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/manipulation-of-information-in-times-of-crisis-evidence-from-covid-excess-mortality/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Karlinsky and Shayo ask which governments manipulate public information, in which direction, and by how much — questions that are normally intractable because the ground truth is unobservable. The Covid-19 pandemic supplies an unusual opportunity: all countries faced a broadly similar crisis simultaneously, and all-cause mortality — collected by national statistical offices as a routine bureaucratic function independently of Covid — provides a manipulation-resistant benchmark against which officially-reported Covid deaths can be evaluated.&lt;/p&gt;
&lt;p&gt;The authors hand-collect all-cause mortality data for 134 countries and territories from national statistical offices, population registries, health ministries, and, in some cases, right-to-information requests facilitated by local journalists. Data span 2015–2021 at weekly, monthly, or annual frequency. Their sample covers 93 percent of countries with at least 75 percent Death Registration Completeness. They compute, for each country, a Misreporting Rate (MRR) defined as estimated Covid deaths minus officially reported Covid deaths, normalised by expected total deaths derived from pre-pandemic trends. Estimated Covid deaths equal excess mortality — itself estimated from a country-specific model with weekly/monthly fixed effects and an annual trend (R² = 0.997 in pre-pandemic prediction) — minus adjustments for excess deaths attributable to conflicts, natural disasters, and other identifiable non-Covid causes. Those adjustments are small: the mean total adjustment across the sample is 0.04 percent of expected deaths.&lt;/p&gt;
&lt;p&gt;Six main findings emerge. First, between 45 and 55 percent of the 134 countries misreported Covid deaths. Second, the direction of manipulation is overwhelmingly one-sided: of 131 countries with sufficient data to estimate confidence intervals, 59 reported accurately, 62 significantly underreported, and only 10 overreported. The theoretical prediction that governments might exaggerate a crisis — to rally populations, legitimise repressive measures, or attract foreign aid — finds no empirical support. Third, the magnitude of underreporting is large: the sample reported 5.08 million Covid deaths in 2020–2021 while estimated actual Covid deaths were 12.47 million, nearly 2.5 times the official figure; the implied global MRR is 12.8 percent. Among the 62 underreporting countries, the average MRR is 14.5 percent of expected total deaths and the median is 12 percent. Individual-country MRRs range from above 37 percent (Bolivia, Nicaragua) downward, with Russia at 24 percent. Fourth, state capacity in counting and registering deaths explains some but far from most cross-country variation; the R² of the best capacity-only regression is 0.115. Chile and Russia have virtually identical Death Registration Completeness and Percent Well-Certified Death Registrations, yet Chile accurately reported while Russia&amp;rsquo;s MRR is 24 percent. Fifth, the extent of underreporting is strongly associated with constraints on governmental power. In individual regressions conditioning on capacity, each of three institutional constraint measures — Clean Elections, Executive Constraints, and Freedom of the Press — is associated with a 0.4–0.5 standard deviation lower MRR per one standard deviation stronger constraint. In a joint model including all 12 factors from four domains (macroeconomic incentives, culture, audience sophistication, institutions), institutional constraints are the strongest predictor (partial R² ≈ 0.11), followed by audience sophistication (partial R² ≈ 0.04–0.06). Macroeconomic incentives — tourism reliance, unemployment, foreign direct investment — are not jointly significant. Cultural factors (trust, individualism, religiosity) lose significance once other factors are controlled. The full model explains more than 50 percent of MRR variation. Sixth, countries with a communist legacy (defined as having had a communist or socialist regime for at least 10 years, covering 34 countries) show significantly higher misreporting even holding current institutional and cultural conditions constant. Countries that held elections during 2020–2021 also show significantly higher misreporting.&lt;/p&gt;
&lt;p&gt;The results are robust to alternative expected-mortality models, alternative MRR normalisations, the inclusion of Bangladesh, China, and Indonesia (treated separately due to data quality concerns), year-by-year (2020 vs. 2021) splits, controls for age structure and GDP per capita, and alternative manipulation measures (underdispersion, Benford&amp;rsquo;s law deviations). The evidence that manipulation cannot be attributed to varying standards for false-positive attribution of cause of death is direct: four pre-pandemic measures of a country&amp;rsquo;s tendency to use unspecified cause-of-death categories are uncorrelated with MRR and individually account for less than 1 percent of its variation.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s contribution to the economics of information manipulation is methodological as well as empirical: it provides a comparable, country-level measure of governmental misinformation based on actual observable actions regarding a policy issue of central importance, covering a large and diverse cross-section of countries.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The strategy compares officially-reported Covid deaths (the variable that attracted political attention and over which governments had strong incentives and ability to intervene) with estimated Covid deaths derived from excess all-cause mortality (a statistic collected routinely by national bureaucracies under very different incentive structures, harder to manipulate, and less visible publicly during the pandemic). The identifying assumption is that all-cause mortality data are not themselves systematically manipulated in response to Covid. The authors defend this on four grounds: (1) all-cause mortality has long been collected independently of Covid; (2) ascertaining that someone died is far easier than attributing a cause of death; (3) Covid figures attracted vastly more public attention, making their manipulation more urgent; (4) when governments appear to have discovered the evidential value of excess mortality, their response has been to delay publication of all-cause data rather than to alter it (Belarus is cited as an example). The main remaining threat is that the adjustment for non-Covid excess deaths (conflicts, disasters, traffic accidents, suicides, homicides) is imperfect in countries with poor data on those causes. The authors note this caveat but show mean adjustments are tiny (0.04% of expected deaths) and the largest individual adjustments (Armenia 6.1%, Azerbaijan 3.2%) are driven by the Nagorno-Karabakh war and are handled explicitly.&lt;/p&gt;
&lt;h3 id="q2-how-is-excess-mortality-estimated-and-how-sensitive-are-the-results-to-modelling-choices"&gt;Q2. How is excess mortality estimated, and how sensitive are the results to modelling choices?&lt;/h3&gt;
&lt;p&gt;Country-specific models are estimated using 2015–2019 all-cause mortality data, including country-specific weekly or monthly fixed effects and a country-specific annual trend to capture seasonality and long-run factors (population ageing, improvements in health care, etc.). The model achieves R² = 0.997 in predicting pre-pandemic mortality. The authors report in Supplementary Material B that alternative expected-mortality approaches from the literature yield very similar results, as do alternative normalisations of the MRR. Sensitivity to model choice is low because the discrepancies between excess and reported deaths in weak-institution countries are so large that they persist across methodological variants.&lt;/p&gt;
&lt;h3 id="q3-how-do-the-authors-distinguish-intentional-manipulation-from-limited-state-capacity"&gt;Q3. How do the authors distinguish intentional manipulation from limited state capacity?&lt;/h3&gt;
&lt;p&gt;They use two pre-pandemic, capacity-specific measures: (1) Death Registration Completeness (DRC) — the share of deaths captured by the vital registration system — and (2) Percent of Well-Certified Death Registrations (PWC) — the share with proper cause-of-death attribution. Both are computed before the pandemic so they are not contaminated by Covid-era behaviour. Regressions confirm that capacity predicts MRR negatively (R² up to 0.115), but the residual variation remains large. The clearest illustration is Chile vs. Russia: both have complete DRC and near-identical high PWC, yet Chile reports accurately and Russia has an MRR of 24 percent. All subsequent analysis of correlates conditions on these capacity measures.&lt;/p&gt;
&lt;h3 id="q4-how-do-the-authors-rule-out-the-possibility-that-differences-in-false-positive-aversion-rather-than-manipulation-explain-mrr-variation"&gt;Q4. How do the authors rule out the possibility that differences in false-positive aversion (rather than manipulation) explain MRR variation?&lt;/h3&gt;
&lt;p&gt;They construct four pre-pandemic measures from WHO Mortality Database ICD-10 cause-of-death data: (1) number of ICD codes reported; (2) share of specific-viral deaths among all viral deaths; (3) share of specific-infection deaths among all infection deaths; (4) share of specific-respiratory deaths among all respiratory deaths. A country more averse to false positives would report less specific causes. None of the four measures is significantly associated with MRR, and none accounts for more than 1 percent of its variation. This rules out differences in diagnostic/reporting standards as a driver of the observed discrepancies.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-direction-of-manipulation-and-what-does-this-imply-for-theories-of-governmental-information-behaviour"&gt;Q5. What is the direction of manipulation and what does this imply for theories of governmental information behaviour?&lt;/h3&gt;
&lt;p&gt;Of 131 countries with estimable confidence intervals, 62 significantly underreported and only 10 overreported. The four main theoretical channels for overreporting — rally-around-the-flag effects, legitimising repression, attracting foreign aid, and inducing flight-to-safety compliance — find no empirical support. The authors argue that the rally-around-the-flag mechanism requires an outgroup-related threat (Covid, unlike a foreign military attack, was not easily framed this way), that Covid mortality does not signal repressive capacity, and that international economic actors appear sufficiently sophisticated to be sceptical of inflated figures. The pattern is consistent instead with governments downplaying to project competence, reduce accountability, and justify inadequate responses.&lt;/p&gt;
&lt;h3 id="q6-what-factors-are-most-strongly-associated-with-misreporting-and-how-are-they-ranked"&gt;Q6. What factors are most strongly associated with misreporting, and how are they ranked?&lt;/h3&gt;
&lt;p&gt;In joint regressions with all 12 factors from four domains, after conditioning on capacity: (1) Institutional constraints (Clean Elections, Executive Constraints, Freedom of the Press) have the highest partial R² (approximately 0.11 for Executive Constraints alone) and are jointly significant at p &amp;lt; 0.001; each standard deviation of stronger institutional constraint is associated with roughly 0.4–0.5 standard deviations lower MRR. (2) Audience Sophistication (tertiary education, HDI Education Index, internet access) is the second strongest domain (partial R² in the range of 0.04–0.06 per variable; jointly significant at p &amp;lt; 0.05). (3) Cultural factors (trust, individualism, religiosity) are individually significant in bivariate regressions but lose significance when institutional and other factors are controlled. (4) Macroeconomic incentives (tourism, unemployment, net FDI) are not jointly significant in any specification. Specification-curve analysis across all combinations of controls confirms that Executive Constraints is the single most robust predictor, retaining sign, magnitude, and significance across all models. The full model (Table 4, column 1) has R² exceeding 0.50.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-communist-legacy-finding-and-how-is-it-interpreted"&gt;Q7. What is the communist legacy finding and how is it interpreted?&lt;/h3&gt;
&lt;p&gt;Countries defined as having had a communist or socialist regime for at least 10 years (34 countries) show significantly higher MRRs even after conditioning on contemporary institutional constraints, audience sophistication, culture, and capacity. The coefficient is statistically significant at p &amp;lt; 0.05 or better in the main and most robustness specifications. The authors point to Harrison (2017) on the pervasiveness of information manipulation in communist states as a historical precedent, and interpret the finding as a persistent legacy operating through channels not fully captured by current measures. This suggests that historical exposure to a political culture of systematic information manipulation may have durable effects on bureaucratic behaviour or political norms that current V-Dem indices do not fully absorb.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-elections-finding"&gt;Q8. What is the elections finding?&lt;/h3&gt;
&lt;p&gt;Countries holding national parliamentary or presidential elections during 2020–2021 (76 of 134 countries) show significantly higher misreporting, consistent with electoral incentive theories of information manipulation. This finding is robust to including controls for GDP per capita, population age structure, and other domains, and is stable across the 2020-only and 2021-only sub-samples.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-performed"&gt;Q9. What robustness checks are performed?&lt;/h3&gt;
&lt;p&gt;The authors conduct: (1) specification-curve analysis across all combinations of covariates; (2) a joint model with all 12 individual factors; (3) principal component analysis within each domain to recover common variation and reduce dependence on specific measurement choices; (4) alternative expected-mortality models (Supplementary Material B.1); (5) alternative MRR normalisations (Supplementary Material B.2); (6) separate year-by-year analysis for 2020 and 2021; (7) inclusion of Bangladesh, China, and Indonesia as robustness cases despite lower data reliability; (8) addition of GDP per capita to check whether the institution-misreporting link is proxying for development; (9) analysis using underdispersion (Kobak 2022) and Benford&amp;rsquo;s law deviations as alternative manipulation measures; (10) exploration of colonial legacy as an additional historical variable (no significant effect found). The primacy of institutional constraints is robust across all of these.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-authors-treat-china-bangladesh-and-indonesia"&gt;Q10. How do the authors treat China, Bangladesh, and Indonesia?&lt;/h3&gt;
&lt;p&gt;These three large countries are excluded from the main analysis because their all-cause mortality data come from surveys (Bangladesh, China) rather than vital registration systems, or are very incomplete (Indonesia), making excess mortality estimation unreliable. They are included in a robustness regression (Table 4, column 6) and results are described as qualitatively similar. The authors flag that China&amp;rsquo;s data may itself be informative as a potential indicator of data suppression.&lt;/p&gt;
&lt;h3 id="q11-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q11. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;The paper is closest in spirit to Olken (2007), who uses the gap between reported and actual infrastructure spending to measure corruption, and Martinez (2022), who compares GDP growth to night-time-light-implied growth and finds autocracies overstate growth by more than a third. The authors extend this approach to a different domain (health/mortality) with broader country coverage. Prior Covid-specific work documented anomalies — underdispersion (Kobak 2022) and Benford&amp;rsquo;s law deviations (Kapoor et al. 2020; Kilani 2021) — and noted that autocratic regimes reported lower-than-expected deaths (Annaka 2021; Cassan and Van Steenvoort 2021), but these studies relied on regime type as the sole or primary explanatory variable and did not systematically rank competing factors. Neumayer and Plümper (2022) and Wigley (2024) used the authors&amp;rsquo; own World Mortality Dataset to test data manipulation. This paper is distinctive in that it: (a) provides what the authors describe as the most systematic estimates to date of Covid mortality and misreporting; (b) examines a broad range of factors across four domains without a priori privileging any; (c) directly tests and rejects capacity and false-positive aversion as alternative explanations; and (d) identifies communist legacy and elections as additional significant correlates.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Three implications are highlighted. First, unconstrained regimes appear to manipulate not only economic statistics but also health information during the most salient public policy event of the era; travel restrictions and multilateral actions during the pandemic relied on reported Covid figures, so manipulation had direct international externalities. This raises broader questions about the credibility of official data from such governments across domains — foreign aid targeting, climate action, vaccination campaigns. Second, the MRR provides a comparable cross-country measure of institutional quality grounded in actual governmental behaviour, potentially useful as an input to studies of institutions, conflict, electoral outcomes, and economic performance. Third, some countries that score respectably on conventional executive constraint indices — Albania, El Salvador, India, Serbia — show high MRRs, suggesting these rates may be leading indicators of democratic erosion not yet captured by standard measures. The scope condition the authors flag is external validity: if pandemic mortality is an extreme case with unique incentive structures (tourism, investment, aid eligibility), then findings about determinants of manipulation may not generalise beyond crisis settings. The authors argue against this interpretation on the grounds that macroeconomic factors — which would be pandemic-specific — are not significant, while institutional constraints — which reflect general governmental behaviour — are.&lt;/p&gt;
&lt;h3 id="q13-what-limitations-do-the-authors-acknowledge"&gt;Q13. What limitations do the authors acknowledge?&lt;/h3&gt;
&lt;p&gt;First, the analysis is explicitly descriptive rather than causal; factors are correlates, not proven determinants. Second, the MRR may understate true manipulation if all-cause mortality data are themselves selectively withheld or manipulated; the authors argue this is probably modest but acknowledge it cannot be fully ruled out. Third, important large countries — Pakistan, Nigeria, Ethiopia, Venezuela — cannot be scored because sufficient all-cause mortality data are not publicly available; the authors note this absence may itself be informative but cannot be quantified. Fourth, data on other causes of excess deaths (traffic accidents, suicides, homicides) are patchy in many countries, though the scale of these adjustments is very small. Fifth, some capacity controls (PWC) use data from as early as 2003, introducing measurement error. The paper does not claim to fully separate the channels through which institutions reduce manipulation (electoral accountability, press scrutiny, judicial oversight, professional agency independence), treating them as joint constraints rather than separately identified mechanisms.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Misreporting Rate (MRR)&lt;/strong&gt;: The paper&amp;rsquo;s central measure, defined as (estimated Covid deaths minus officially reported Covid deaths) divided by expected total deaths for the country in the same period based on pre-pandemic trends. A positive MRR indicates underreporting; a negative MRR indicates overreporting. Normalising by expected total deaths rather than by reported Covid deaths accounts for differences in population size, age structure, and baseline mortality across countries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Excess mortality&lt;/strong&gt;: The number of deaths above and beyond what would have been expected in the absence of the pandemic, estimated from country-specific models with weekly or monthly fixed effects and an annual trend fitted to 2015–2019 data. Used as the primary building block for estimated Covid deaths after subtracting excess deaths due to identified non-Covid causes (conflict, natural disasters, traffic accidents, homicides, suicides).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Death Registration Completeness (DRC)&lt;/strong&gt;: In this paper&amp;rsquo;s usage, the share of all deaths in a country captured by its vital registration system each year, measured using pre-pandemic data. Treated as the most basic indicator of a country&amp;rsquo;s capacity to count deaths. Used as a control to separate capacity constraints from intentional manipulation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Percent of Well-Certified Death Registrations (PWC)&lt;/strong&gt;: The share of death certificates in a country that carry a properly specified cause of death, measured using pre-pandemic data. Used alongside DRC as a second capacity control capturing not just whether deaths are registered but whether causes are correctly attributed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Informational Autocrat&lt;/strong&gt;: Following Guriev and Treisman (2022), the paper uses this concept to describe executives in countries where formal and informal checks and balances are weak, who systematically manipulate public information to project competence and reduce accountability. The paper&amp;rsquo;s empirical results are interpreted as evidence that such executives behave as informational autocrats not only in economic statistics but also in health data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;False-positive aversion&lt;/strong&gt;: The tendency of some countries to apply a higher evidentiary bar before attributing a death to a specific cause — such as Covid — rather than leaving the cause unspecified, independently of capacity or intention to deceive. The paper operationalises this using pre-pandemic ICD-10 data on specificity of reported causes of death and shows it is uncorrelated with MRR, ruling it out as a driver of observed discrepancies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Communist legacy&lt;/strong&gt;: The paper&amp;rsquo;s binary indicator for countries that had a communist or socialist regime for at least 10 consecutive years (34 countries). The variable captures historical exposure to a political culture of systematic information manipulation and is found to be a significant positive predictor of MRR even after conditioning on current institutional constraints, consistent with persistent norms or bureaucratic practices.&lt;/p&gt;</description></item><item><title>Medical innovation and health disparities</title><link>https://macropaperwarehouse.com/papers/medical-innovation-and-health-disparities/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/medical-innovation-and-health-disparities/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks why medical innovation can widen health disparities even when it unambiguously improves health for everyone who takes it. The authors argue that the standard access-versus-preferences dichotomy is a false one: disadvantaged patients can rationally forgo effective medications because treatment side effects interfere with work, and the income cost of not working is particularly severe for low-education workers who hold physically demanding, inflexible jobs. Health-maximizing and welfare-maximizing behavior are therefore not the same thing, and the gap between the two is systematically larger for lower-education individuals.&lt;/p&gt;
&lt;p&gt;The empirical setting is the introduction of Highly Active Antiretroviral Therapy (HAART) for HIV in the mid-1990s. HAART was substantially more effective than prior mono- and combo-therapy at preventing AIDS progression and death, but it produced harsh physical side effects (fatigue, diarrhea, headache, fever). Data come from the Multi-Center AIDS Cohort Study (MACS), a semi-annual panel of men who have sex with men in Baltimore, Chicago, Pittsburgh, and Los Angeles, covering 1991–2003. After sample restrictions, the analysis uses 11,290 person-visit observations for 1,201 HIV-positive individuals aged 30–64, approximately 63% of whom hold a college degree or more. The study dichotomizes education into less-than-college versus college-or-more and tracks treatment choices, labor supply, immune-system health (CD4 count, with AIDS threshold at 250), physical ailments, income, insurance, and out-of-pocket medical expenditures.&lt;/p&gt;
&lt;p&gt;The structural model is a lifecycle discrete-choice dynamic programming framework in which forward-looking individuals simultaneously choose treatment (no treatment, monotherapy, combotherapy, and post-1995 HAART) and full-time work or non-work each half-year period to maximize expected lifetime utility. Health and survival evolve stochastically as functions of prior health, treatment, and age. Utility is a function of consumption (income minus out-of-pocket expenses), ailments, and labor supply, with utility parameters allowed to differ by education. The model is estimated via maximum likelihood using nested backwards induction; the quasi-experimental introduction of HAART as an unanticipated shock helps identify utility parameters.&lt;/p&gt;
&lt;p&gt;Key quantitative results: (1) HAART drastically reduced mortality for both groups—six-month mortality fell from 9% to 2% for less-educated men and from 6% to 1% for college graduates—and raised the probability of maintaining a high CD4 count from 62% to 78% (less-educated) and 68% to 83% (college+). (2) Despite equivalent access (both groups face roughly 91-95% insurance coverage and similarly low out-of-pocket costs), lower-educated men adopted HAART at a lower rate (58% of post-HAART visits versus 66% for college graduates) and approximately five months later. (3) The structural utility parameters confirm that while the direct disutility of ailments is not significantly different across education groups, the disutility of working while experiencing ailments is substantially larger in magnitude for less-educated men (estimated parameter -2.73) than for college graduates (-1.97). (4) Measured as expected lifetime utility, HAART&amp;rsquo;s introduction increased value for low-CD4 men by 236.1% (less-educated) versus 176.6% (college+), but in absolute utility units the gains were larger for college graduates—establishing that HAART increased welfare inequality. (5) Decompositions show the largest single driver of the education gap in HAART value is the differential survival process; income differences also matter but financial access variables (insurance, out-of-pocket costs) explain little. (6) A simulated six-month HAART mandate improves health—by 1.7 percentage points more for less-educated men—but reduces expected lifetime value by 2.8% for the less-educated versus 1.4% for college graduates, and reduces employment by 4.1% versus 1.6%, as mandated HAART forces men into ailment-producing treatment whose side effects they cannot manage alongside work. (7) A counterfactual $10,000-per-six-months non-labor income subsidy (similar to COVID-19 transfer policies) reduces work by 31–49% for less-educated men and by 25–39% for college graduates, while inducing an 81.2% increase in HAART take-up among less-educated men in good health who were not previously on treatment (from 5% to 9% baseline probability), and a 44.5% increase for similar college graduates (8% to 11%). For men with AIDS-level CD4 counts not on treatment, the policy raises the probability of being healthy next period by 12.6% for less-educated men and 5.3% for college graduates.&lt;/p&gt;
&lt;p&gt;The central mechanism is a wedge between health and welfare that is steeper for disadvantaged workers: occupational conditions make it harder to work while experiencing side effects, so the opportunity cost of HAART compliance is higher. This means effective medical innovation—precisely by creating more severe side effects than older regimens—can widen welfare inequality even as it compresses mortality gaps. Clinical trials that randomize assignment to treatment and measure health outcomes will register the innovation as a success while masking the distributional welfare costs. Policy interventions that reduce the cost of not working (income transfers, labor market restructuring) can simultaneously increase HAART take-up and improve health, with effects concentrated among the disadvantaged.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-main-identification-strategy-and-what-are-the-key-threats-to-identification"&gt;Q1. What is the main identification strategy and what are the key threats to identification?&lt;/h3&gt;
&lt;p&gt;The model is estimated by maximum likelihood using nested backwards induction over observable state variables. A key identifying variation is the quasi-experimental, unanticipated introduction of HAART in 1995, which shifts the choice set mid-panel and allows the authors to trace behavioral responses to an exogenous change in treatment efficacy and side-effect profiles. Disutility of ailments and work parameters are identified by conditional choice probabilities given state variables (health, ailment status, prior treatment) and by comparing behavior before and after HAART availability. The authors follow Magnac and Thesmar (2002) to establish that under the distributional assumptions (Type I EV shocks, fixed discount factor β=0.95) and the normalization imposed, the likelihood has a unique maximum. The main threats are: (a) the assumption that individuals were surprised by HAART (no forward-looking anticipation), which simplifies the model but is explicitly noted—Hamilton et al. (2021) show that incorporating individual expectations substantially complicates the framework; (b) the exclusion of unobserved heterogeneity in the utility function, though specifications including it produce very small probabilities of a second type (below 5%); (c) the absence of borrowing and saving, which could allow more educated individuals to smooth consumption across treatment cycles—the authors note this would bias downward the disutility of working with ailments for higher-educated individuals, meaning the estimated cross-education difference in that parameter is a lower bound; (d) the sample is restricted to white men in four cities, limiting external validity; and (e) the education dichotomy collapses heterogeneity within education groups.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-through-which-education-moderates-the-health-welfare-tradeoff-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms through which education moderates the health-welfare tradeoff, and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The paper identifies two nested channels. First, the estimated structural utility parameter for working while experiencing ailments is larger in magnitude for less-educated men (θ = -2.73) than for college graduates (θ = -1.97), indicating greater disutility from combining work and side effects. The paper argues this reflects occupational sorting: lower-education men are significantly more likely to hold manual occupations (occupation score 5.12 versus 4.49 for college graduates, where higher scores indicate more manual tasks per Autor et al. 2003), making physical side effects especially incompatible with job performance. Second, lower-educated men have lower incomes ($15,373 versus $22,290 per half-year for less-educated versus college-educated, pre-HAART), so the income cost of not working is larger in relative terms, creating stronger incentives to maintain employment even at the cost of forgoing treatment. The authors decompose the relative contribution of these mechanisms in the non-labor income subsidy simulation: when they give lower-educated men the income process of higher-educated men (Appendix Figure A1), the gap in behavioral response narrows but does not close; when they give lower-educated men the disutility parameters of higher-educated men (Figure A2), similarly the gap narrows but remains. Both mechanisms are jointly operative.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-in-haart-take-up-and-welfare-value-is-documented"&gt;Q3. What heterogeneity in HAART take-up and welfare value is documented?&lt;/h3&gt;
&lt;p&gt;Education is the primary heterogeneity dimension examined. Post-HAART, lower-educated men used HAART in 58% of observations versus 66% for college graduates, were slower to start (5 months later on average), and less likely to ever use it (67% versus 81%). Health status interacts with education: low-CD4 men gain more in percentage terms from HAART because they are more in need of its health-improving effects (236.1% gain for less-educated low-CD4 versus 176.6% for college-educated low-CD4; 85.7% versus 76.3% for high-CD4 men, with college graduates gaining more in absolute utility units throughout). The welfare cost of a treatment mandate is higher for less-educated men (2.8% lifetime value decline versus 1.4%), and the employment reduction induced by the mandate is also larger for them (4.1% versus 1.6%). In the income subsidy simulation, low-CD4 men not on any medication show the largest health response. The paper does not examine race/ethnicity heterogeneity, having excluded non-white individuals from the analysis due to sampling methodology concerns.&lt;/p&gt;
&lt;h3 id="q4-what-does-the-value-decomposition-reveal-about-why-haart-benefited-more-educated-men-more"&gt;Q4. What does the value decomposition reveal about why HAART benefited more-educated men more?&lt;/h3&gt;
&lt;p&gt;Table A17 sequentially replaces the processes and parameters of lower-educated agents with those of higher-educated agents. Giving lower-educated men the income process of college graduates narrows but does not close the gap—income is not the primary driver. Replacing the insurance and medical expenditure processes slightly reduces value for less-educated men relative to giving them only the income process, because more-educated individuals actually have somewhat higher out-of-pocket costs. Changing the health and ailments processes has modest positive effects. The largest single contributor to closing the education gap is the survival process: less-educated men face much higher baseline mortality, which depresses the expected present value of all future flows including the gains from HAART. This suggests that policies targeting survival differentials (e.g., access to other health services) could partially close the HAART welfare gap. Finally, replacing the utility parameters mechanically closes the remaining gap, but preferences are less amenable to direct policy intervention than the survival process.&lt;/p&gt;
&lt;h3 id="q5-what-do-the-treatment-mandate-simulations-show-and-why-do-they-matter-for-evaluating-clinical-trials"&gt;Q5. What do the treatment mandate simulations show, and why do they matter for evaluating clinical trials?&lt;/h3&gt;
&lt;p&gt;A six-month HAART mandate mimics randomized assignment to treatment in a clinical trial. It improves health—the probability of high CD4 rises by 1.7 percentage points more for less-educated men than baseline (reflecting a larger baseline gap in HAART use)—which would appear a policy success from a health-only perspective. However, expected lifetime utility falls by 2.8% for less-educated men and 1.4% for college graduates, because mandated HAART forces individuals into ailment-inducing treatment they would not have chosen, inhibiting labor supply. Employment falls by 4.1% for less-educated men versus 1.6% for college graduates. Appendix analyses removing the ailment-producing properties of treatment largely eliminate both the welfare cost and the employment effect, confirming that ailments are the mediating channel. This shows that clinical trials—which typically report health endpoints and do not measure welfare or distributional consequences—can mask the costs that effective but side-effect-heavy treatments impose, and that those costs fall disproportionately on less-advantaged patients.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-non-labor-income-subsidy-simulation-show-and-which-groups-respond-most"&gt;Q6. What does the non-labor income subsidy simulation show, and which groups respond most?&lt;/h3&gt;
&lt;p&gt;A permanent $10,000-per-six-months increase in non-employment income (approximately 50% of median income, calibrated to COVID-era transfer policies) induces labor force exit across all groups but concentrates its health-promoting effects among disadvantaged men who were not already on HAART. Among relatively healthy (high-CD4) less-educated men not using any medication, HAART take-up rises by 81.2% (from 5% to 9%); the corresponding figure for college graduates is 44.5% (from 8% to 11%). Among men with AIDS-level (low) CD4 not on treatment, the probability of being healthy next period increases by 12.6% for less-educated men and 5.3% for college graduates. Men already on HAART—who are unlikely to change treatment regardless—show little response. The policy has small but positive health externalities beyond the immediate recipients, since people on antiretrovirals have lower viral loads and lower transmission risk. Decomposition simulations (Appendix Figures A1–A2) show that both the income-level channel and the disutility-of-work-with-ailments channel independently contribute to the larger lower-education response, with neither alone sufficient to fully explain the differential.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The paper is most closely related to Papageorge (2016, Quantitative Economics), which uses the same MACS data and setting to link non-uptake of HAART to labor supply and side effects. The key difference is scope: Papageorge (2016) focuses on individual-level mechanisms; the present paper&amp;rsquo;s goal is to characterize distributional differences in the health-welfare tradeoff across education groups and to show that innovation can exacerbate existing inequality. Chan, Hamilton, and Papageorge (2016, Review of Economic Studies) also use the MACS setting to study the value of medical innovation, and Hamilton, Hincapié, Miller, and Papageorge (2021, International Economic Review) examine the diffusion of HAART. Relative to the sociological fundamental cause theory literature (Link and Phelan 1995; Phelan et al. 2010), which documents that medical innovations tend to widen health disparities, the present paper provides a structural quantification of the specific mechanisms and their relative magnitude. Relative to papers attributing health disparities primarily to access barriers (insurance, cost), the paper provides evidence that for this sample—where insurance coverage exceeds 91% even for less-educated men and HIV drugs are inexpensive—access explains little of the educational disparity in HAART use or health outcomes.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The core implication is that policies reducing the cost of not working—income transfers, disability benefits, worker protections—can raise HAART adoption and improve health among disadvantaged patients, precisely the group for whom standard health-access policies have limited traction. The non-labor income subsidy simulation suggests that the health improvements are modest in absolute magnitude (a 0.2% rise in probability of being healthy next period for the best-responding group among high-CD4 non-HAART users, and 13% for low-CD4 non-HAART users), but there are unmodeled positive externalities through reduced transmission risk that would multiply the social return. Scope conditions: (1) The sample is white men who have sex with men in four U.S. cities during 1991–2003, enrolled in a prospective cohort study; generalizability to other populations (women, racial minorities, other diseases) is uncertain. (2) The income subsidy that triggers HAART take-up must be large enough to induce labor force exit; a $10,000 per-six-months transfer is needed to generate the simulated behavioral response, larger for higher-income workers. (3) The paper explicitly notes that drug costs and insurance are not binding constraints in this sample, and the policy conclusions may differ in settings with weaker drug coverage. (4) Mental health is excluded from the model; the paper shows depression variables have smaller effects on treatment choice than the physical mechanisms included, but mental health could independently affect some populations&amp;rsquo; response. The paper&amp;rsquo;s conclusions extend to other conditions where effective treatment has disabling side effects and disadvantaged patients hold inflexible physical jobs—the authors invoke COVID-19 as a contemporary analog.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-conducted"&gt;Q9. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The authors report several robustness exercises. Treatment transition results are shown to be robust to defining the HAART introduction period as survey visit 23 or 25 rather than 24. Ailment specifications are noted to be robust to varying the type or frequency of ailments counted (citing Papageorge 2016 for this). Specifications including unobserved heterogeneity in the utility function produce very small second-type probabilities (below 5%), arguing against its inclusion. The treatment mandate simulations are run under three alternative shock-assignment methods (2 draws, 8 draws, and the preferred 2-draw approach), with results consistent across methods on the main welfare-versus-health asymmetry. Appendix Tables A19 and A20 remove ailments from all medications and from HAART only, respectively, confirming that the welfare cost of mandates is driven by treatment-induced ailments. Appendix Figures A1 and A2 mechanically decompose the education-differential response to the income subsidy by replacing income processes and disutility parameters separately, confirming that both channels are active. The model fit (Table A9) shows overall employment (66% model, 66% data) and HAART use (33% model, 36% data) closely matching, though the model slightly over-predicts medication use among low-CD4 individuals.&lt;/p&gt;
&lt;h3 id="q10-why-does-the-paper-focus-on-white-men-only-and-what-does-this-imply-for-interpretation"&gt;Q10. Why does the paper focus on white men only, and what does this imply for interpretation?&lt;/h3&gt;
&lt;p&gt;The authors drop 1,098 observations from 390 non-white individuals because of concerns about the sampling methodology used to recruit the refresher sample for those individuals—specifically, non-white participants entered the panel via a different selection process that could confound estimates. The paper does not investigate racial disparities in HAART take-up, which are also well-documented in the literature. This is a significant limitation because HIV/AIDS has disproportionately affected Black men in the United States, and the mechanisms the paper identifies—occupational sorting, income constraints, disutility of working with ailments—may operate differently or more intensely along racial lines. The authors acknowledge this limitation and note that the structural framework could in principle be applied to other groups if appropriate data were available.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Health-welfare tradeoff&lt;/strong&gt;: In this paper, the wedge between the action that maximizes health (taking effective medication despite side effects) and the action that maximizes lifetime utility (avoiding medication to remain employed and maintain income). The tradeoff is not a bias or error but a rational response to economic constraints, and it is wider for less-educated individuals whose occupational conditions make working with side effects especially costly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HAART (Highly Active Antiretroviral Therapy)&lt;/strong&gt;: A combination antiretroviral HIV treatment introduced in the mid-1990s, far more effective than prior mono- or combo-therapy at improving CD4 count and preventing AIDS-level immune decline and death. In this paper&amp;rsquo;s model, HAART serves as the innovation whose adoption the authors study: it is more efficacious but produces harsher side effects than earlier treatments, and its introduction is treated as an unanticipated aggregate shock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Disutility of working with ailments&lt;/strong&gt;: A structural utility parameter (θ_2,f=0) capturing how much worse-off an agent feels from working while experiencing physical ailments (fatigue, diarrhea, headache, fever). Estimated at -2.73 for less-educated men and -1.97 for college graduates, this parameter is the primary driver of the differential health-welfare tradeoff across education groups and explains why side-effect-bearing treatments like HAART are disproportionately avoided by lower-education workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Treatment mandate simulation&lt;/strong&gt;: A counterfactual in which all agents are assigned to HAART for six months (eliminating choice among other treatment options), used to mimic randomized assignment in a clinical trial. The simulation is designed specifically to illustrate that health improvements observable in a clinical trial coexist with welfare reductions and employment disruptions that would not be captured in standard trial endpoints.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fundamental cause theory&lt;/strong&gt;: A sociological framework (Link and Phelan 1995) arguing that socioeconomic status is a &amp;lsquo;fundamental cause&amp;rsquo; of health disparities that persists despite or is even amplified by medical innovation, because more advantaged individuals are better positioned to adopt and benefit from new treatments. The paper provides structural economic microfoundations for this theory by quantifying the mechanisms through which HAART&amp;rsquo;s introduction widened the welfare gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-labor income subsidy&lt;/strong&gt;: A counterfactual policy simulation in which non-employment income is raised by $10,000 per six months (approximately 50% of the median person&amp;rsquo;s income), modeled after COVID-19 transfer policies. In the paper&amp;rsquo;s model this policy reduces employment but increases HAART take-up and health improvements particularly for less-educated HIV-positive men who were previously forgoing treatment to maintain income from work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Source text origin&lt;/strong&gt;: Not a paper-specific concept but denoted here: the full working paper text was obtained from the NBER Working Paper (No. 28864), not from abstract-only, satisfying the GUARD requirement.&lt;/p&gt;</description></item><item><title>On the elasticity of substitution between labor and ICT and IP capital and traditional capital</title><link>https://macropaperwarehouse.com/papers/on-the-elasticity-of-substitution-between-labor-and-ict-and-ip-capital-and-traditional-capital/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/on-the-elasticity-of-substitution-between-labor-and-ict-and-ip-capital-and-traditional-capital/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper estimates the elasticities of substitution between labor, information and communication technology (ICT) and intellectual property (IP) capital, and traditional capital using a nested constant elasticity of substitution (CES) production function. The motivation is twofold: standard macroeconomic models aggregate all capital into a single input and thus miss potentially distinct substitution relationships, and competing estimates of the labor-capital elasticity of substitution diverge sharply — with some finding gross substitutability (Karabarbounis and Neiman 2013) and others gross complementarity (Glover and Short 2020) — leaving unexplained the observed decline in labor income share across advanced economies.&lt;/p&gt;
&lt;p&gt;The data come from the 2023 release of the EU KLEMS database for nine Euro Area economies (Austria, Belgium, Finland, France, Germany, Italy, Netherlands, Portugal, and Spain) over 1996-2020 (with Germany ending in 2019 and Portugal starting in 2001). The nesting structure places an ICT-IP capital aggregate (itself a CES nest of ICT equipment and IP capital, which includes software, databases, patents, and R&amp;amp;D capital) together with labor in an inner nest, and that combined aggregate is then nested with traditional capital in an outer nest. The rationale for grouping ICT and IP capital is their joint and complementary use — computers and software — and the observation that roughly 25% of granted patents in the sample period are ICT-related. Estimation follows the normalized CES methodology of Grandville (1989), Klump, McAdam, and Willman (2007), and Leon-Ledesma, McAdam, and Willman (2010), which jointly estimates the logged and normalized production function together with its first-order conditions using feasible generalized nonlinear least squares, weighting by country-year employment shares and correcting for heteroscedasticity and serial correlation. This approach is preferred because normalization anchors the point elasticity at sample averages and Monte Carlo evidence shows it outperforms first-order-condition-only or translog alternatives, especially when identifying factor-augmenting technological change alongside substitution elasticities.&lt;/p&gt;
&lt;p&gt;The main results (Table 4, column 1) are as follows. The elasticity of substitution between labor and traditional capital (ε1) is estimated at 0.745 (standard error 0.009), statistically significantly below 1, implying gross complementarity. The elasticity between labor and the ICT-IP aggregate (ε2) is 1.187 (0.010), significantly above 1, implying gross substitutability. The elasticity between ICT and IP capital themselves (ε3) is 0.961 (0.003), significantly below 1, implying gross complementarity within the ICT-IP nest. The ICT capital-augmenting technological change parameter (γ_ICT) is estimated at 0.725, several orders of magnitude larger than the labor-augmenting parameter (γ_L = 0.003), consistent with rapid technological progress in ICT. The IP capital-augmenting parameter (γ_IP) is negative (−0.111), and the traditional capital-augmenting parameter (γ_TK) is negative but statistically insignificant (−0.002). For the US, ε2 is substantially larger at 1.712 (0.133), with ε1 = 0.724 (0.024) and ε3 = 0.922 (0.017).&lt;/p&gt;
&lt;p&gt;A counterfactual accounting exercise (fixing ICT and IP technological progress indexes and capital stocks at their 1996 levels) finds that absent these developments, labor income share would have slightly increased in European countries rather than declining, and would have declined by about 75% less in the US over the sample period. ICT accumulation and technological progress is the dominant driver of the fall: absent ICT changes alone, labor share would have risen significantly in Europe.&lt;/p&gt;
&lt;p&gt;The paper also derives the implied aggregate labor-capital elasticity (εL,K) using Hicks&amp;rsquo;s formula applied to the nested production function. The imputed εL,K for European countries ranges from approximately 1.36 to 1.43 over 1996-2020, rising through 1996-2008 and declining afterward. The US imputed values are substantially higher, ranging from approximately 2.14 to 2.37. By contrast, when the author directly estimates a two-input CES function combining labor with aggregate capital, the estimated elasticity is significantly below 1 (approximately 0.988 for European countries in the constant-CES specification), far below the imputed values. This divergence demonstrates that production function specification is consequential for identifying the labor-capital elasticity, and that models treating all capital as a single input can generate downward-biased estimates of this parameter.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The author jointly estimates a normalized CES production function and first-order conditions (capital return equations and the wage equation) using feasible generalized nonlinear least squares with multiple starting points, selecting results by log likelihood, AIC, BIC, and R-squared. Normalization anchors the elasticity as a point elasticity at geometric sample averages, which is theoretically motivated and improves finite-sample identification. Main threats include: (1) endogeneity of factor inputs — the system of equations is estimated jointly but without instrumental variables, relying on non-arbitrage conditions to close the model; (2) negative estimates for γ_IP and γ_TK, which the author acknowledges may capture markups or capital underutilization rather than true technical change (Jiang and Leon-Ledesma 2018 show that omitting markups can bias the sign of capital-augmenting technology); (3) the US results are sensitive to initial values for the estimation algorithm, possibly because of the small sample size (24 observations); and (4) the counterfactual exercise abstracts from equilibrium effects and free-factor supply adjustments.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-distinguishing-the-three-capital-types-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms distinguishing the three capital types, and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;ICT capital (computers, communication devices, peripherals) and IP capital (software, databases, patents, R&amp;amp;D capital) are grouped in an inner nest on the grounds of their complementary joint use. Traditional capital (machinery, transport, construction and structures) forms the outer nest. This nesting allows the elasticity of substitution between labor and the ICT-IP aggregate (ε2 &amp;gt; 1, gross substitute) to differ from the elasticity between labor and traditional capital (ε1 &amp;lt; 1, gross complement), which the paper argues is consistent with the automation literature&amp;rsquo;s emphasis on ICT displacing routine tasks. The elasticity of substitution within the ICT-IP nest (ε3 &amp;lt; 1) reflects gross complementarity between ICT equipment and IP assets (one needs software to use computers). The empirical distinction comes from the separate first-order conditions for each capital type, which link each capital&amp;rsquo;s income share to its stock and price, allowing the three elasticities to be separately identified.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented-across-countries-or-time"&gt;Q3. What heterogeneity is documented across countries or time?&lt;/h3&gt;
&lt;p&gt;The main estimates pool 9 European countries weighted by employment shares; the author does not report country-by-country elasticity estimates but does report country-level descriptive statistics (Table I in the Data Appendix). Time-series heterogeneity is addressed through the imputed aggregate elasticity εL,K, which rises from approximately 1.367 in 1996 to a peak around 1.388-1.426 near 2008 (varying across the sensitivity columns of Table 6) and then declines to approximately 1.369-1.411 by 2020. The US elasticities are systematically higher than the European ones (εL,K ranging approximately 2.14-2.37 for the US vs. 1.36-1.43 for Europe; ε2 = 1.712 for the US vs. 1.187 for Europe). The time-varying aggregate capital specification in Table 7 shows the estimated ε1 for European countries follows an inverted-U shape over the sample period, while the US estimate shows the contrary pattern (though the latter is imprecise due to the small sample).&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper estimates two alternative CES nesting structures (equations 20 and 21, reported in columns 2 and 3 of Table 4) to assess sensitivity to the nesting assumption. In specification (20), labor and traditional capital are nested first and then combined with the ICT-IP aggregate, so the elasticity between labor and ICT-IP equals that between traditional capital and ICT-IP. In specification (21), the different capital types are nested first and then combined with labor. Both alternatives confirm that ICT and IP capital are gross substitutes for labor. The paper also estimates a two-input labor-aggregate capital function in three variants: constant CES, elasticity as a linear function of compensation shares and relative prices, and elasticity as a quadratic polynomial of time (Table 7). Results using US data from the EU KLEMS database are reported separately (column 4 of Table 4 and columns 8-9 of Table 6). The imputed εL,K is further verified using data counterparts of the compensation shares rather than model-predicted shares (column 7 of Table 6), yielding essentially identical results with higher variability.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Relative to Karabarbounis and Neiman (2013), this paper agrees that labor and aggregate capital are gross substitutes (imputed εL,K &amp;gt; 1) and that capital deepening drives the labor share decline, but attributes the mechanism specifically to ICT and IP capital accumulation rather than the fall in all capital prices. It contrasts with Glover and Short (2020), whose below-1 estimates the paper reconciles by showing that treating all capital as a single input biases the aggregate elasticity downward. Relative to Eden and Gaggl (2018, 2019), who use US data and find ICT (including software) substitutes for labor in first-order-condition-only estimates, this paper adds normalization and biased technical change parameters and uses European panel data, and also separates ICT equipment from IP/software. Relative to Koh, Santaeulalia-Llopis, and Zheng (2020), who perform an accounting exercise attributing the labor share decline to IP capital capitalization, this paper provides structural estimates of substitution elasticities and corroborates the IP capital importance. Relative to Aum and Shin (2024), who use Korean firm-level data and find software substitutes for labor while ICT equipment complements it, this paper uses a different nesting (ICT and IP grouped together) and European aggregate data, and finds the combined ICT-IP aggregate is a gross substitute for labor — consistent with Aum and Shin&amp;rsquo;s software result driving the within-nest finding. The normalization approach distinguishes the paper from Antras (2004) and earlier aggregate studies that estimate only first-order conditions (which can produce upward-biased elasticity estimates when biased technical change is omitted).&lt;/p&gt;
&lt;h3 id="q6-what-does-the-paper-find-about-the-source-of-the-labor-share-decline-and-what-are-the-scope-conditions-on-this-result"&gt;Q6. What does the paper find about the source of the labor share decline, and what are the scope conditions on this result?&lt;/h3&gt;
&lt;p&gt;The counterfactual exercise (Section 4.2, Panel B of Table 3) finds that absent ICT and IP capital technological progress and accumulation, labor income share would have slightly increased in European countries over 1996-2020 rather than falling. Absent ICT changes alone, labor share would have risen significantly in Europe. The ICT-driven decline is the dominant contributor. By contrast, absent IP capital trends, labor share would have fallen substantially more (suggesting IP capital compensation growth, when attributed to capital rather than labor, partially offsets the ICT effect on labor&amp;rsquo;s share but its own share rise is the proximate driver of labor share decline). For the US, absent ICT and IP developments, labor share decline would have been about 75% smaller. Scope conditions: this is a static accounting exercise holding free factors at initial values and abstracting from general equilibrium effects. The results apply to total industrial value added (not individual sectors) and to the nine Euro Area countries in the sample. The exercise assumes the estimated production function parameters are the correct structural parameters, and thus inherits any limitations of the identification strategy.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-implication-for-the-measured-aggregate-labor-capital-elasticity-and-why-does-it-differ-from-standard-estimates"&gt;Q7. What is the implication for the measured aggregate labor-capital elasticity, and why does it differ from standard estimates?&lt;/h3&gt;
&lt;p&gt;When the paper estimates a two-input (labor, aggregate capital) CES function directly, the estimated aggregate elasticity is significantly below 1 and close to estimates from Herrendorf, Herrington, and Valentinyi (2015). When it instead imputes the aggregate elasticity from the nested-CES parameter estimates using Hicks&amp;rsquo;s formula, the imputed values exceed 1 and are much larger. The paper shows analytically that εL,K &amp;gt; ε2 when the relative capital cost of ICT compared to traditional capital (pKICT&lt;em&gt;KICT / pTK&lt;/em&gt;TK) takes sufficiently low values, which is the case in the data. This divergence arises because the single-input capital specification conflates the high substitutability of labor with ICT-IP capital and the low substitutability with traditional capital, yielding a biased estimate that depends on the capital composition. The paper concludes that production function specification is consequential for identifying the aggregate labor-capital substitution elasticity.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-key-data-features-that-drive-the-results"&gt;Q8. What are the key data features that drive the results?&lt;/h3&gt;
&lt;p&gt;ICT investment prices fell at an average annual rate of -4.6% relative to value added prices over the sample, while IP and traditional capital investment prices changed by -0.3% and +0.1% per year, respectively. Real ICT capital stocks grew at 4.9% per year, versus 3.4% for IP capital and 1.6% for traditional capital. ICT and IP capital depreciate rapidly (20.1% and 24.1% per year) compared to traditional capital (3.6%). These patterns imply computed rates of return on ICT capital that were very high at the start of the sample (131% in 1996, largely reflecting the fall in ICT prices that year) and fell sharply to 24% by 2020. The average share of labor and ICT-IP compensation in value added is approximately 71%, with labor making up about 92% of that combined share. The ICT share within the ICT-IP nest is about 21%, meaning IP capital compensation is substantially larger than ICT capital compensation.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Allen-Uzawa elasticity of substitution&lt;/strong&gt;: A point elasticity measuring the percentage change in the ratio of two inputs in response to a percentage change in their price ratio, holding output and other input prices constant. In this paper, it is estimated as a structural parameter of the nested CES production function, normalized at sample geometric averages; values above 1 imply gross substitutability and values below 1 imply gross complementarity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Normalized CES production function&lt;/strong&gt;: A CES specification that is indexed to sample averages of output and inputs so that the elasticity of substitution is defined as a point elasticity at those averages. This normalization, following Grandville (1989) and Leon-Ledesma et al. (2010), facilitates identification of both elasticity parameters and factor-augmenting technological change parameters, avoiding the conflation that arises in unnormalized specifications.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gross substitutes / gross complements&lt;/strong&gt;: Two inputs are gross substitutes (elasticity of substitution &amp;gt; 1) if a fall in the relative price of one leads to a rise in the share of cost devoted to it, reducing the other input&amp;rsquo;s cost share. They are gross complements (elasticity &amp;lt; 1) if a fall in relative price instead reduces cost share. In this paper, labor and ICT-IP capital are gross substitutes; labor and traditional capital and ICT with IP capital are gross complements.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Traditional capital (TK)&lt;/strong&gt;: In this paper&amp;rsquo;s taxonomy, all non-ICT, non-IP capital: machinery, transport equipment, construction, and structures. It is the residual capital category and is defined as a gross complement of labor in the estimated nested CES structure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intellectual property (IP) capital&lt;/strong&gt;: Capital comprising software, databases, patents (including R&amp;amp;D capital), and other forms of intellectual property as measured in the EU KLEMS database. IP capital is grouped with ICT equipment in an inner CES nest on the grounds of complementary use. Its compensation share rise is the proximate accounting factor in the labor share decline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Factor-augmenting technological change&lt;/strong&gt;: Hicks-neutral or biased technical progress that enters multiplicatively with a specific factor input in the production function (e.g., γ_ICT for ICT capital), scaling the effective quantity of that input. In this paper, the ICT-augmenting parameter is estimated to be very large and positive (0.725), reflecting rapid ICT productivity growth, while IP- and traditional-capital-augmenting parameters are negative, which the author suggests may partly reflect markups or underutilization rather than pure technology.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Imputed aggregate labor-capital elasticity&lt;/strong&gt;: The elasticity of substitution between labor and total capital derived analytically from the nested CES parameters using Hicks&amp;rsquo;s formula, rather than estimated directly from a two-input specification. In this paper, the imputed value exceeds 1 for Europe (~1.36-1.43) and is substantially higher for the US (~2.14-2.37), contrasting with directly estimated values that are below 1, illustrating the sensitivity of this parameter to production function specification.&lt;/p&gt;</description></item><item><title>Optimal Combination of Patent Instruments in a Cumulative-Innovation Growth Model</title><link>https://macropaperwarehouse.com/papers/optimal-combination-of-patent-instruments-in-a-cumulative-innovation-growth-model/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/optimal-combination-of-patent-instruments-in-a-cumulative-innovation-growth-model/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper develops a tractable general equilibrium model of endogenous growth driven by cumulative innovation, and uses it to characterize optimal patent policy — both for patent breadth (via a &amp;ldquo;non-infringing inventive step&amp;rdquo; requirement) and patent length — with a focus on their welfare implications and optimal combination.&lt;/p&gt;
&lt;p&gt;The central motivation is that cumulative innovation creates positive knowledge spillovers: each new idea strictly builds on the best existing technology, and the disclosure that patenting requires diffuses knowledge to future innovators. Because private firms do not internalize these spillovers, the decentralized equilibrium features strictly lower R&amp;amp;D investment than the social optimum. The key wedge is an intertemporal spillover effect: firms discount future profits at a rate that includes the hazard of being superseded (rho + lambda&lt;em&gt;v&lt;/em&gt;L), while the social planner uses only the pure time preference rate (rho). Appropriability and business-stealing externalities exactly offset each other, so the intertemporal spillover is the sole source of under-investment.&lt;/p&gt;
&lt;p&gt;The model has a continuum of differentiated varieties, a single labor input, a Poisson idea arrival process (rate lambda per R&amp;amp;D worker), and productivity improvements drawn i.i.d. from a standardized Pareto distribution with shape parameter theta &amp;gt; 1. The Pareto structure yields the key tractability: the log of the k-th best productivity level is Gamma-distributed with mean k/theta, which allows closed-form welfare expressions. In steady state, all outcomes depend on just three deep parameters: the discount rate rho, the Pareto shape theta, and the innovative capacity lambda*L.&lt;/p&gt;
&lt;p&gt;The patent breadth instrument is formalized as a &amp;ldquo;non-infringing inventive step&amp;rdquo; (NIS) requirement B &amp;gt;= 1: a new idea must deliver a productivity at least B times the current patent-holder&amp;rsquo;s productivity to qualify for a patent. Raising B creates two opposing forces. The &amp;ldquo;profit effect&amp;rdquo; extends incumbent monopoly duration by reducing the hazard rate of supersession (from lambda&lt;em&gt;v&lt;/em&gt;L to lambda&lt;em&gt;v&lt;/em&gt;L&lt;em&gt;B^{-theta}), raising innovation incentives. The &amp;ldquo;hurdle effect&amp;rdquo; raises the bar an idea must clear to be patentable, reducing the expected return to R&amp;amp;D. These forces generate a non-monotonic (inverted-U) relationship between R&amp;amp;D effort and B (Proposition 2): there is a unique B_v that maximizes the innovation rate, with dv/dB &amp;gt; 0 for B &amp;lt; B_v and dv/dB &amp;lt; 0 for B_v &amp;lt; B &amp;lt; B_0 (the upper bound beyond which no R&amp;amp;D occurs). Explicitly, B_v = [lambda&lt;/em&gt;L / (rho*(theta-1))]^{1/theta}. Proposition 3 further establishes that in economies whose innovative capacity falls just below the threshold for positive growth at B=1, a well-chosen NIS can shift the economy from a zero-growth to a positive-growth steady state.&lt;/p&gt;
&lt;p&gt;The welfare-maximizing breadth B_w is shown to be unique, binding (B_w &amp;gt; 1), and strictly below B_v (Proposition 4 and 5). The welfare optimum trades off the dynamic gain from greater innovation against the static consumer surplus loss from higher markup power. Because the dynamic gain is still positive when B &amp;lt; B_v (R&amp;amp;D is still rising) but the static loss grows continuously in B, the welfare maximum necessarily occurs in the region where research is still increasing — i.e., B_w &amp;lt; B_v.&lt;/p&gt;
&lt;p&gt;Numerically, at baseline parameters (rho = 0.07, theta = 4, lambda&lt;em&gt;L = 1), B_w = 1.14 and the equilibrium R&amp;amp;D share is v(B_w) = 0.22, implying an asymptotic maximum real wage growth rate of 4.8%. The optimal breadth is most sensitive to theta (Pareto tail thickness) and less sensitive to rho and lambda&lt;/em&gt;L.&lt;/p&gt;
&lt;p&gt;When patent length (Omega) is added as a second instrument, the model yields a sharp result: the welfare-maximizing policy sets Omega → infinity together with B = B_w (Proposition 6). Unlike patent breadth, patent length has no hurdle effect — a longer patent duration raises R&amp;amp;D monotonically (dv/dOmega &amp;gt; 0, Lemma 2). With no diminishing returns to innovation effort in this model (the Poisson arrival rate is proportional to vL), the marginal dynamic gain from extending Omega always strictly outweighs the marginal static loss, so infinite patent length is always superior to any finite length. With Omega = 20 years (the TRIPS standard), the baseline calibration implies B_w = 1.13 and v(B_w) = 0.21 — only slightly below the infinite-length benchmark — suggesting the qualitative infinite-length result has limited quantitative bite for realistic patent durations.&lt;/p&gt;
&lt;p&gt;Proposition 7 shows that patent breadth and patent length are policy complements: when patent length is exogenously constrained to a finite value, the welfare-maximizing breadth increases in Omega (dB_w/dOmega &amp;gt; 0). Intuitively, a shorter patent duration weakens innovation incentives, so the optimal NIS compensates by providing stronger breadth protection.&lt;/p&gt;
&lt;p&gt;The paper provides a unified rationalization of several empirical puzzles: the weak or negative relationship between patent strength and innovation rates (Sakakibara-Branstetter 2001 on Japan; Bessen-Maskin 2009 on US software) is consistent with B being set above B_v, where the hurdle effect dominates; the causal evidence in Galasso-Schankerman (2014) that patents impede cumulative knowledge accumulation is consistent with the hurdle effect operating at the margin.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-is-this-a-theoretical-or-empirical-paper"&gt;Q1. What is the identification strategy, and is this a theoretical or empirical paper?&lt;/h3&gt;
&lt;p&gt;This is a purely theoretical paper. There is no empirical identification strategy. The core contribution is an analytically tractable general equilibrium model in which the key results (Propositions 1–7) are derived from first-order conditions, comparative statics, and the application of the intermediate value theorem. The Pareto-improvement distribution is the key parametric assumption that enables closed-form expressions for welfare and the growth rate.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-key-model-departure-from-kortum-1997-and-eaton-kortum-2001"&gt;Q2. What is the key model departure from Kortum (1997) and Eaton-Kortum (2001)?&lt;/h3&gt;
&lt;p&gt;Kortum (1997) and Eaton-Kortum (2001) model ideas as drawn from a stationary distribution over productivity levels — new ideas may or may not surpass the existing frontier, and as ideas accumulate it becomes progressively less likely that a new draw beats the current best. This generates growth only if the workforce grows. Chor and Lai instead model productivity improvements (ratios Z_{k+1}/Z_k) as i.i.d. Pareto draws, so each new idea strictly improves on the frontier regardless of how many ideas have arrived. This cumulative structure generates endogenous growth with a constant workforce and introduces knowledge spillovers that are absent in Kortum (1997).&lt;/p&gt;
&lt;h3 id="q3-what-exactly-is-the-non-infringing-inventive-step-nis-and-how-does-it-differ-from-other-breadth-concepts-in-the-literature"&gt;Q3. What exactly is the &amp;rsquo;non-infringing inventive step&amp;rsquo; (NIS) and how does it differ from other breadth concepts in the literature?&lt;/h3&gt;
&lt;p&gt;The NIS requirement B stipulates that a new idea must achieve a productivity at least B times the productivity of the current best patent (i.e., Z_new &amp;gt;= B * Z_current) to be patentable and non-infringing (what the paper calls &amp;rsquo;leading breadth&amp;rsquo;). The paper notes this is distinct from — though related to — patentability requirements studied by O&amp;rsquo;Donoghue (1998), which focused on the minimum improvement to qualify for a new patent but not necessarily on infringement. It also differs from the Gilbert-Shapiro (1990) and Klemperer (1990) breadth concepts, which focus on horizontal product differentiation (consumer willingness to substitute away from a patent) rather than vertical quality improvements. In the paper&amp;rsquo;s model, both patentability and non-infringement requirements are captured by a single parameter B, with the simplifying assumption that meeting the B hurdle is both necessary and sufficient for non-infringement.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-three-externalities-in-the-model-and-which-one-drives-the-market-planner-wedge"&gt;Q4. What are the three externalities in the model, and which one drives the market-planner wedge?&lt;/h3&gt;
&lt;p&gt;Three externalities are present: (1) The intertemporal spillover effect — firms do not internalize that their innovation raises the knowledge base for future innovators. (2) The appropriability effect — firms capture only private profits, not the full consumer surplus gain from each innovation. (3) The business-stealing effect — each innovator imposes a negative externality on the incumbent patent-holder by eroding their profits. Effects (2) and (3) exactly offset each other in the Pareto specification, so only the intertemporal spillover effect remains. This is verified formally: the market equilibrium condition features a discount rate of rho + lambda&lt;em&gt;v&lt;/em&gt;L (including the creative destruction hazard), whereas the social planner&amp;rsquo;s problem involves only rho. The wedge between v_eqm and v_SP stems entirely from this higher effective discount rate in decentralized equilibrium.&lt;/p&gt;
&lt;h3 id="q5-why-is-the-welfare-maximizing-patent-breadth-strictly-less-than-the-innovation-rate-maximizing-breadth"&gt;Q5. Why is the welfare-maximizing patent breadth strictly less than the innovation-rate-maximizing breadth?&lt;/h3&gt;
&lt;p&gt;At B_v, research effort is at its maximum, but this is achieved by granting patent-holders maximum protection, imposing the largest static consumer surplus loss. For B between B_w and B_v, increasing B further raises the static loss but no longer raises the innovation rate significantly enough to compensate; in fact for B &amp;gt; B_v, research effort falls while the static loss remains. The welfare optimum trades off the dynamic benefit (higher innovation) against the static cost (monopoly pricing). Because welfare must also account for the static loss at each period, and this loss is already large at B_v, the welfare optimum is achieved at a lower level of protection. Formally, dU_0/dB &amp;lt; 0 for all B in [B_v, B_0), and the unique welfare maximum lies strictly in [1, B_v).&lt;/p&gt;
&lt;h3 id="q6-why-is-the-optimal-patent-length-infinite"&gt;Q6. Why is the optimal patent length infinite?&lt;/h3&gt;
&lt;p&gt;Unlike patent breadth, patent length has only a profit effect and no hurdle effect — a longer patent strictly raises R&amp;amp;D effort (Lemma 2). Moreover, the model has no diminishing returns to innovation effort: the Poisson arrival rate of ideas is simply proportional to the total number of R&amp;amp;D workers at each date (lambda&lt;em&gt;v&lt;/em&gt;L), so each additional unit of research labor generates the same expected innovation flow regardless of how much research has already been done. This means the marginal dynamic gain from raising Omega (via increased innovation) is approximately constant, while the marginal static loss (additional consumer surplus ceded per period) is also roughly constant. The dynamic gain always strictly exceeds the static loss as long as the economy can sustain positive R&amp;amp;D (Lemma 1 condition holds), so Omega → infinity is always welfare-improving. This result breaks down if one introduces diminishing returns to R&amp;amp;D (e.g., a fishing-out effect or a congestion externality in research).&lt;/p&gt;
&lt;h3 id="q7-are-patent-breadth-and-patent-length-policy-substitutes-or-complements"&gt;Q7. Are patent breadth and patent length policy substitutes or complements?&lt;/h3&gt;
&lt;p&gt;They are policy complements (Proposition 7): when patent length is shorter (e.g., exogenously constrained by TRIPS or ethical considerations), the welfare-maximizing breadth B_w is lower; conversely, a longer patent length calls for a higher optimal breadth. This is because a longer patent length increases the dynamic gain from research, which raises the marginal value of also increasing breadth (since breadth further amplifies the monopoly profit effect). Formally, d^2U^l_0/(dB d Omega) &amp;gt; 0 at B_w, implying dB_w/d Omega &amp;gt; 0 by the implicit function theorem.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-quantitative-calibration-and-what-are-the-key-numerical-results"&gt;Q8. What is the quantitative calibration, and what are the key numerical results?&lt;/h3&gt;
&lt;p&gt;The calibration is illustrative rather than structural. Baseline: rho = 0.07 (matching real stock market returns as in Kortum 1997), theta = 4 (implying expected profits = 25% of per-variety expenditure, since 1/(1+theta) = 0.20 &amp;hellip; actually 1/(1+4) = 0.20, with the text stating 1/(1+theta) = 0.25 implying theta=3; the paper states theta=4 gives 1/(1+theta) = 0.20 — there is a slight inconsistency in the text&amp;rsquo;s wording, but the stated result is 25% of expenditures per variety), lambda*L = 1 (one expected new idea per variety per year). These yield: B_w = 1.14 (infinite patent length), v(B_w) = 0.22 (22% of labor in R&amp;amp;D), and an asymptotic maximum real wage growth rate of 4.8%. The optimal breadth B_w is most sensitive to theta: lowering theta (fatter tail, larger average improvements) raises B_w substantially. Under a finite patent length of Omega = 20, the results change minimally: B_w = 1.13, v(B_w) = 0.21.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-model-handle-the-possibility-that-economies-with-low-innovative-capacity-might-not-innovate-at-all-without-policy"&gt;Q9. How does the model handle the possibility that economies with low innovative capacity might not innovate at all without policy?&lt;/h3&gt;
&lt;p&gt;When lambda&lt;em&gt;L &amp;lt; rho&lt;/em&gt;theta, the economy has no R&amp;amp;D in the decentralized equilibrium at B = 1 (v(1) &amp;lt; 0 per equation 22). However, Proposition 3 shows that if lambda&lt;em&gt;L falls in the intermediate range (rho&lt;/em&gt;(theta-1)&lt;em&gt;(theta^2/(theta^2-1))^theta &amp;lt; lambda&lt;/em&gt;L &amp;lt; rho*theta), there exists a range of binding NIS values B &amp;gt; 1 that can shift the economy from zero to positive growth. Setting B = B_v achieves this transition. This is because the profit effect of introducing a binding NIS can more than offset the hurdle effect in this regime, making it profitable for some workers to engage in R&amp;amp;D.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-key-welfare-improving-scope-conditions-for-the-nis-policy"&gt;Q10. What are the key welfare-improving scope conditions for the NIS policy?&lt;/h3&gt;
&lt;p&gt;The welfare gain from a binding NIS requires Assumption 1: lambda&lt;em&gt;L &amp;gt; rho&lt;/em&gt;theta. This ensures the economy already features positive R&amp;amp;D at B = 1, and that the innovative capacity is large enough so the dynamic gains from raising B above 1 exceed the static consumer surplus losses. Without this condition, the NIS may either fail to generate R&amp;amp;D (if lambda*L is very low) or may tip the economy into R&amp;amp;D via Proposition 3&amp;rsquo;s mechanism, but welfare-optimality of the NIS still requires the economy be in a regime where the profit effect dominates for small B. Additionally, the NIS must remain below B_v to generate any dynamic gain.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-model-relate-to-japans-narrow-patent-breadth-policy-from-1960-1993"&gt;Q11. How does the model relate to Japan&amp;rsquo;s narrow patent breadth policy from 1960-1993?&lt;/h3&gt;
&lt;p&gt;The paper cites Ordover (1991) and Maskus-McDaniel (1999) to note that Japan deliberately adopted narrow patent breadth to encourage more incremental innovation and technology catch-up. In the model&amp;rsquo;s terms, Japan was setting B close to 1 (or even at 1) to lower the hurdle for new patents, maximizing the number of patentable ideas. This is consistent with a strategy of maximizing the innovation rate (operating near B_v or even below it), potentially at the cost of some dynamic welfare optimization. The Apple v. Samsung example illustrates that the US tends toward broader patent breadth (higher B) than Japan, consistent with the model&amp;rsquo;s international variation in NIS standards.&lt;/p&gt;
&lt;h3 id="q12-how-does-the-paper-handle-the-price-markup-and-profit-structure-under-the-nis"&gt;Q12. How does the paper handle the price markup and profit structure under the NIS?&lt;/h3&gt;
&lt;p&gt;Under Bertrand competition with limit pricing, the incumbent with the best patentable technology sets price equal to the marginal cost of the second-best technology (the previous patent-holder). The price markup m = Z_k/Z_{k-1} is drawn from a Pareto distribution with shape theta and lower bound 1 (no NIS) or B (with NIS). Flow profits are therefore: Pi = B(1+theta)^{-theta} / [B(1+theta) - theta] &amp;hellip; more precisely from equation (19): Pi = [B(1+theta) - theta] * (B(1+theta))^{-1}. As B rises, Pi increases (higher average markups from higher minimum improvement), which is the profit effect. The expected log productivity of the k-th patentable idea is E[ln Z~_k] = k/theta + k*ln(B), confirming that higher B raises not just the probability threshold but also the expected productivity of successful innovations.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-limitations-and-potential-extensions-noted-by-the-authors"&gt;Q13. What are the limitations and potential extensions noted by the authors?&lt;/h3&gt;
&lt;p&gt;The authors acknowledge several limitations and propose extensions: (1) The model assumes fully cumulative innovation — each idea strictly builds on the frontier. Generalizing to partial cumulativeness (where some ideas are non-cumulative or only partially built on existing knowledge) is flagged as a natural extension. (2) The analysis is confined to a single-country setting. A multi-country extension would allow study of cross-border patent policy spillovers and optimal international IPR harmonization (e.g., under TRIPS). (3) The model does not allow directed research — firms cannot target specific varieties. Relaxing this could introduce additional policy margins. (4) The model abstracts from imitation threats, which Gallini (1992) shows can make broader patent protection optimal.&lt;/p&gt;
&lt;h3 id="q14-how-does-the-paper-compare-to-odonoghue-1998-and-odonoghue-zweimüller-2004"&gt;Q14. How does the paper compare to O&amp;rsquo;Donoghue (1998) and O&amp;rsquo;Donoghue-Zweimüller (2004)?&lt;/h3&gt;
&lt;p&gt;O&amp;rsquo;Donoghue (1998) shows a patentability requirement can raise social welfare in a partial equilibrium setting, and Hunt (2004) finds an inverted-U relationship between innovation rate and requirement strength — both echo Chor-Lai&amp;rsquo;s findings. O&amp;rsquo;Donoghue-Zweimüller (2004) embed patentability in a quality-ladder endogenous growth model but focus more on innovation effects than welfare. The contribution of Chor-Lai relative to these papers is: (i) a fully general equilibrium treatment with explicit welfare analysis; (ii) derivation of both the welfare-maximizing breadth and the innovation-maximizing breadth and proof that Bw &amp;lt; Bv; (iii) extension to jointly optimal patent breadth and length, showing infinite patent length is optimal; and (iv) the Pareto-Gamma tractability that yields closed-form expressions and enables clean comparative statics on three deep parameters.&lt;/p&gt;
&lt;h3 id="q15-what-robustness-checks-does-the-paper-provide"&gt;Q15. What robustness checks does the paper provide?&lt;/h3&gt;
&lt;p&gt;The paper notes in the main text that results are robust to removing the scale effect (the feature that the innovation rate increases in L). An online appendix (referenced but not included in this draft) proves that the main qualitative results — inverted-U in innovation vs. B, unique welfare-maximizing B_w &amp;lt; B_v, and infinite optimal patent length — survive in a model variant without the scale effect. The numerical sensitivity analysis in Section 3.4 also demonstrates robustness of the qualitative findings across wide ranges of rho (0.02 to 0.12) and theta (2 to 6) and lambda*L.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Non-Infringing Inventive Step (NIS) requirement&lt;/strong&gt;: A patent policy parameter B &amp;gt;= 1 stipulating that a new idea must achieve a productivity at least B times that of the current best patent to qualify for a patent and be deemed non-infringing. In the paper&amp;rsquo;s usage, this simultaneously captures both the patentability requirement and the leading breadth (protection of incumbents against near-imitation), and is used interchangeably with &amp;lsquo;patent breadth.&amp;rsquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cumulative innovation&lt;/strong&gt;: An innovation process in which each new idea strictly improves upon the existing technological frontier. Formally, the productivity improvement Z_{k+1}/Z_k is drawn i.i.d. from a Pareto distribution with support [1, infinity), so each arriving idea always delivers a strictly positive productivity gain over the current best technology. This contrasts with non-cumulative models (e.g., Kortum 1997) where draws are from a stationary distribution and may fall below the frontier.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Profit effect (of patent breadth)&lt;/strong&gt;: The mechanism by which a higher NIS requirement B reduces the hazard rate that an incumbent patent-holder is superseded (from lambda&lt;em&gt;v&lt;/em&gt;L to lambda&lt;em&gt;v&lt;/em&gt;L*B^{-theta}), thereby extending the expected duration of monopoly power and raising the value of each patent. This increases R&amp;amp;D incentives by raising expected profits from successful innovation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hurdle effect (of patent breadth)&lt;/strong&gt;: The mechanism by which a higher NIS requirement B reduces the probability that any given arriving idea is patentable (probability B^{-theta}), thereby lowering the expected return to engaging in R&amp;amp;D. This discourages research effort and is the force that eventually dominates when B becomes sufficiently large, causing the innovation rate to fall.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Innovative capacity&lt;/strong&gt;: The product lambda&lt;em&gt;L, where lambda is the per-worker Poisson arrival rate of ideas and L is the total labor endowment. All steady-state outcomes in the model depend on lambda and L only through this product, not their individual values. It is the key parameter determining whether positive R&amp;amp;D equilibrium exists (requires lambda&lt;/em&gt;L &amp;gt; rho*theta) and the magnitude of welfare gains from patent policy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intertemporal spillover externality&lt;/strong&gt;: The sole market failure driving under-investment in R&amp;amp;D in this model&amp;rsquo;s Pareto specification. Because the knowledge embodied in each marketed innovation diffuses freely and becomes the base for subsequent cumulative improvements, private innovators do not internalize the benefit their R&amp;amp;D confers on future innovators. This causes firms to use an effective discount rate of rho + lambda&lt;em&gt;v&lt;/em&gt;L (including the creative destruction hazard) rather than rho alone, leading to strictly less R&amp;amp;D than the social optimum. Appropriability and business-stealing externalities exactly cancel in this model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy complementarity (breadth and length)&lt;/strong&gt;: The property that the welfare-maximizing patent breadth B_w is increasing in patent length Omega: dB_w/d Omega &amp;gt; 0. When the patent authority is constrained to set a shorter patent length, the optimal breadth should also be narrower, and vice versa. This arises because a longer patent length raises the marginal dynamic benefit of providing stronger breadth protection.&lt;/p&gt;</description></item><item><title>Optimal Fiscal Policy in a Climate-Economy Model with Heterogeneous Households</title><link>https://macropaperwarehouse.com/papers/optimal-fiscal-policy-in-a-climate-economy-model-with-heterogeneous-households/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/optimal-fiscal-policy-in-a-climate-economy-model-with-heterogeneous-households/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether inequality and redistributive taxation should make climate policy more or less ambitious, and how optimal carbon taxes interact with optimal income taxes when households differ in productivity, wealth, and energy demand. The motivation is twofold: equity considerations belong at the center of normative climate analysis, and the distributional consequences of environmental policies are increasingly recognized as critical for their political feasibility — as illustrated by the Yellow Vests episode in France. The paper extends Barrage (2020)&amp;rsquo;s representative-agent dynamic climate-Ramsey model to a heterogeneous-agent setting, using the Werning (2007) technique to characterize the Ramsey optimum in terms of aggregate variables. The government maximizes utilitarian social welfare choosing linear taxes on labor income, capital income, energy, and pollution plus a uniform lump-sum transfer. The climate module is calibrated to DICE 2016 (Nordhaus, 2017). Household heterogeneity is calibrated to US data: ten productivity groups from SCF 2013 hourly wages ranging from $6.44 (bottom decile) to $101.35 (top decile), yielding a model consumption Gini of 0.33, very close to the empirical value of 0.32 (Heathcote et al., 2010). Tax rates are set at effective US rates from Trabandt and Uhlig (2012): capital income tax of 41.1% and labor income tax of 25.5%. The model period is five years beginning in 2015, and the discount factor follows DICE at beta = 1/(1.015) per year, with inverse IES sigma = 1.45. The main quantitative exercise compares optimal policy to a climate-skeptic planner who sets carbon taxes to zero. Key findings: (i) Tax distortions have a negligible effect on the optimal carbon tax in the heterogeneous-agent setting. The second-best carbon tax is initially only 0.5% below the social cost of carbon (SCC) and subsequently fluctuates within about 0.2% above or below it — in sharp contrast to Barrage (2020), who finds tax distortions reduce optimal carbon taxes by 8% in the representative-agent setting. The key mechanism is that, with heterogeneous agents, the government optimally levies distortionary taxes for redistributive purposes (not merely to finance public spending), so the marginal cost of public funds (MCF) averages to 1 over time and its temporal deviations are quantitatively trivial. (ii) Income inequality only slightly reduces the optimal carbon tax: residual consumption inequality after optimal income-tax redistribution lowers the SCC by 3.9% in the baseline. The mechanism is that inequality raises the average marginal utility of consumption (because the marginal utility function is convex), increasing the opportunity cost of abatement; this effect dominates when IES &amp;lt; 1 (sigma &amp;gt; 1 in the calibration). (iii) The optimal carbon tax path starts at $21.7/tCO2 in 2020 and reaches $229.2/tCO2 one century later — levels consistent with Barrage (2020) and Nordhaus (2017/2018) but insufficient to achieve the Paris +2°C target under baseline damages. (iv) Comparing optimal policy to the climate-skeptic baseline, the additional carbon tax revenue is split nearly equally: the present value of labor taxes falls by 0.7% of GDP, while transfers rise by 0.8% of GDP. This violates the weak double-dividend hypothesis, which prescribes using carbon tax revenue exclusively to cut distortionary taxes. (v) The optimal policy has progressive welfare effects in the 21st century, because increased tax progressivity benefits lower-income households. The average discounted welfare gain is 5.8% of consumption under baseline damages. In the long run, gains become regressive because richer households (with IES &amp;lt; 1) are willing to pay proportionally more in consumption to avoid temperature increases. By contrast, a representative-agent double-dividend policy — using all carbon revenue to cut labor taxes — is regressive from the outset, with low-income households bearing a net cost even in the short run. The 3.9% inequality effect on the SCC is robust to changes in fiscal pressure and damage calibration but is sensitive to sigma: with sigma = 2, inequality reduces optimal carbon taxes by 16.2% rather than 3.9%. Extensions with wealth heterogeneity, heterogeneous energy demand (calibrated to CEX), and heterogeneous environmental damage sensitivity confirm that the MCF remains negligible and the inequality effect on carbon taxes remains small in quantitative terms.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-theoretical-result-on-the-optimal-carbon-tax-and-why-does-it-differ-from-barrage-2020"&gt;Q1. What is the core theoretical result on the optimal carbon tax and why does it differ from Barrage (2020)?&lt;/h3&gt;
&lt;p&gt;The optimal carbon tax is approximately Pigouvian — set equal to the social cost of carbon — because the MCF averages to 1 over time with balanced-growth preferences when households are heterogeneous and the government can optimize a uniform lump-sum transfer. In Barrage (2020)&amp;rsquo;s representative-agent model, the government cannot choose the level of lump-sum taxes or transfers because there is no redistribution motive, so distortionary taxes are the only way to finance public spending and the MCF exceeds 1, reducing optimal carbon taxes by 8%. With heterogeneous agents, the government optimally provides lump-sum transfers for redistribution, so the constraint on transfers is barely binding and the MCF is close to 1 even when the ability to adjust transfers is removed.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-mechanism-by-which-inequality-affects-the-optimal-carbon-tax-and-what-is-the-sign"&gt;Q2. What is the mechanism by which inequality affects the optimal carbon tax, and what is the sign?&lt;/h3&gt;
&lt;p&gt;Inequality reduces the optimal carbon tax when IES &amp;lt; 1 (sigma &amp;gt; 1). The mechanism operates through the Pigouvian tax formula: pollution abatement reduces aggregate consumption, and the welfare cost of this reduction depends on the social marginal utility of consumption (Vc,t). With inequality, Vc,t is affected by two opposing forces. First, the average marginal utility of consumption is higher because of Jensen&amp;rsquo;s inequality (convex marginal utility function), increasing the opportunity cost of abatement and pushing the pollution tax down. Second, additional consumption goes disproportionately to richer households with lower marginal utilities, reducing Vc,t and pushing the tax up. When IES &amp;lt; 1, the first (higher average marginal utility) effect dominates, so inequality unambiguously reduces the SCC and hence the optimal pollution tax. When IES = 1, the two effects exactly cancel and inequality has no effect.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-mcf-and-why-does-it-average-to-1-in-the-heterogeneous-agent-setting"&gt;Q3. What is the MCF and why does it average to 1 in the heterogeneous-agent setting?&lt;/h3&gt;
&lt;p&gt;The MCF is defined as the ratio of the public (planner&amp;rsquo;s Lagrange multiplier on the resource constraint) to the private (aggregate welfare-weighted) marginal utility of consumption. It measures the social cost of transferring resources from the private to the public sector. The MCF averages to 1 because the first-order condition for the uniform lump-sum transfer implies that the sum of the Lagrange multipliers on agents&amp;rsquo; implementability constraints is zero. With balanced-growth preferences, this implies the welfare-weighted average MCF equals 1 from period 0. The temporal covariance between type-specific shadow costs (theta_i) and the type-specific implementability term (I_{c,i,t}) averages to zero over time, so while the MCF can deviate temporarily from 1, it is 1 on average.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-double-dividend-hypothesis-and-how-does-the-papers-optimal-policy-relate-to-it"&gt;Q4. What is the double-dividend hypothesis and how does the paper&amp;rsquo;s optimal policy relate to it?&lt;/h3&gt;
&lt;p&gt;The weak double-dividend hypothesis holds that it is optimal to use carbon tax revenue exclusively to reduce distortionary taxes, yielding both environmental and efficiency dividends. The paper shows this does not hold with heterogeneous agents: at the optimum, the welfare gain from a marginal reduction in tax distortions equals the welfare loss from increased inequality, so the government splits carbon revenue between cutting distortionary taxes and increasing redistribution. In the baseline quantification, the split is roughly equal: present-value labor taxes fall by 0.7% of GDP and lump-sum transfers rise by 0.8% of GDP. By contrast, following the double-dividend prescription — using all carbon revenue to reduce labor taxes without raising transfers — generates a strongly regressive policy in which low-income households bear net welfare costs even in the short run.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-calibration-strategy-and-how-does-the-model-match-us-inequality-data"&gt;Q5. What is the calibration strategy and how does the model match US inequality data?&lt;/h3&gt;
&lt;p&gt;The economic side is calibrated to the US, while the climate side uses DICE 2016. The discount factor follows DICE (beta = 1/(1.015) per year), and sigma = 1.45 (IES = 1/1.45). Household productivity is calibrated using SCF 2013 hourly wage deciles, yielding ten equal-sized groups with hourly wages from $6.44 (bottom) to $101.35 (top), normalized so that the productivity-weighted average is 1. Although productivity inequality is directly targeted rather than moments of the consumption distribution, the model correctly predicts the consumption Gini of 0.33, close to the empirical 0.32 (Heathcote et al., 2010). Capital and labor income tax rates are from Trabandt and Uhlig (2012): 41.1% and 25.5% respectively. Government debt-to-GDP is approximately 111% (average 2011-2015, IMF). The Frisch elasticity of labor supply is targeted at 0.75 (Chetty et al., 2011). Production in both sectors is Cobb-Douglas with energy share nu = 0.04 from Golosov et al. (2014).&lt;/p&gt;
&lt;h3 id="q6-what-happens-to-optimal-income-taxes-in-the-model"&gt;Q6. What happens to optimal income taxes in the model?&lt;/h3&gt;
&lt;p&gt;The optimal labor income tax roughly doubles from its calibrated level of 25% to about 50% in the first period and stabilizes there. Revenue from these taxes is rebated via the uniform lump-sum transfer, achieving most of the desired redistribution. Because optimal labor income taxes are approximately constant over time, the associated intertemporal distortions are small, and the optimal capital income tax converges to zero quickly after the second period. The mechanism is that, with access to lump-sum transfers, the only reason to tax capital income is to mitigate intertemporal distortions created by labor income taxation; when labor taxes are roughly constant, this motive is weak.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-sensitivity-analysis-reveal-about-the-robustness-of-the-39-inequality-effect"&gt;Q7. What does the sensitivity analysis reveal about the robustness of the 3.9% inequality effect?&lt;/h3&gt;
&lt;p&gt;The effect of inequality on optimal carbon taxes is robust along several dimensions but sensitive to sigma. Under the high-damage scenario (cubic rather than quadratic damage function, yielding an SCC about four times larger), the inequality effect falls to 2.6% rather than 3.9%, because higher carbon taxes reduce warming and thus the share of utility (rather than production) damages. The effect is roughly proportional to the degree of productivity inequality: half the inequality implies about half the effect on the carbon tax. The effect changes more than proportionally with sigma: with sigma = 2 (IES = 0.5), inequality reduces carbon taxes by 16.2%, versus 3.9% with the DICE value of sigma = 1.45. With sigma = 1, the effect is exactly zero. Government expenditure levels and fiscal pressure have negligible effects on the results. The share of damages entering utility directly matters: if only 10% of damages affect utility directly (versus the baseline 26%), the inequality effect falls to 1.8%; if 40% affect utility directly, it rises to 5.2%.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-role-of-initial-wealth-inequality"&gt;Q8. What is the role of initial wealth inequality?&lt;/h3&gt;
&lt;p&gt;Initial wealth inequality (studied in Section 6.1) creates an additional motive for deviating from Pigouvian taxation in period 0 only. Because the planner cannot use the period-0 capital tax to expropriate initial wealth (it is fixed at 41.1%), higher damages would reduce interest rates and thereby partially mitigate wealth inequality (a subtle indirect redistribution mechanism), calling for lower pollution taxes in period 0. Quantitatively, this produces a significant reduction in the initial-period optimal carbon tax. However, from period 1 onward, the optimal tax rules are unaffected by initial wealth heterogeneity, and the effects of MCF and income inequality remain very similar to the baseline. Welfare gains from carbon taxation in the wealth-heterogeneity extension are U-shaped with income but strictly increasing in initial wealth.&lt;/p&gt;
&lt;h3 id="q9-how-does-energy-demand-heterogeneity-stone-geary-extension-affect-the-results"&gt;Q9. How does energy-demand heterogeneity (Stone-Geary extension) affect the results?&lt;/h3&gt;
&lt;p&gt;The extension introduces a second dirty consumption good with Stone-Geary preferences, calibrated using CEX data to match the average energy expenditure share of 10.8% and the observed distribution of energy budget shares across and within income groups. Target emissions share from household energy consumption is 30%. The optimal pollution tax formula remains a modified Pigouvian rule (the MCF structure is unchanged), and the MCF effect remains negligible. The inequality effect on carbon taxes stays near 3.9%, rising marginally to 4.1% with identical energy necessity and 4.1% with heterogeneous energy necessity. Theoretically, the optimal excise tax on the energy good is zero when energy preferences are homogeneous; with heterogeneous necessity levels calibrated to the US, the optimal energy excise tax is quantitatively tiny: about -0.4% of energy prices (a small subsidy). The negative sign arises because within-income-group heterogeneity in energy needs means that energy-intensive households (who are valued more by the planner on average) can be partially targeted via a subsidy. Under the double-dividend scenario with energy inequality, regressive effects are magnified: the poorest, most energy-intensive households actually lose in welfare terms even accounting for long-run climate mitigation benefits.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-paper-establish-theoretically-about-heterogeneous-environmental-damages"&gt;Q10. What does the paper establish theoretically about heterogeneous environmental damages?&lt;/h3&gt;
&lt;p&gt;Proposition 6 (Section 6.3) shows that with additively separable environmental utility and a utilitarian planner, heterogeneous marginal utility damages from pollution have no effect on the optimal pollution tax: they enter the welfare criterion symmetrically and cancel in the aggregate. The pollution tax increases relative to the utilitarian benchmark only if the planner&amp;rsquo;s welfare weights are positively correlated with marginal utility damages — that is, if the planner cares relatively more about the households that are more exposed. A Rawlsian planner would set a higher pollution tax if and only if the least-well-off household is also more sensitive to environmental degradation.&lt;/p&gt;
&lt;h3 id="q11-what-are-third-best-policy-results-when-either-income-tax-is-fixed"&gt;Q11. What are third-best policy results when either income tax is fixed?&lt;/h3&gt;
&lt;p&gt;The paper analyzes policies where either the labor or capital income tax is fixed at its current calibrated level (studied in Appendix E, with results referenced in the main text). These constraints introduce an additional fiscal interaction effect on the optimal carbon tax — the carbon tax is pushed below its second-best Pigouvian level when the fixed tax is set at a sub-optimally low level, and above it when the fixed tax is sub-optimally high. The roles of the MCF and income inequality remain similar to the second-best baseline under these third-best constraints.&lt;/p&gt;
&lt;h3 id="q12-how-does-the-paper-relate-to-and-differ-from-the-double-dividend-and-pollution-taxation-literatures"&gt;Q12. How does the paper relate to and differ from the double-dividend and pollution taxation literatures?&lt;/h3&gt;
&lt;p&gt;The paper builds on three earlier pillars. First, Pigou (1920) established first-best Pigouvian taxation. Second, a large literature (Sandmo, 1975; Bovenberg and de Mooij, 1994; Bovenberg and Goulder, 1996) showed that in representative-agent second-best settings the MCF exceeds 1 and optimal pollution taxes fall below the Pigouvian level. Barrage (2020) is the closest dynamic general-equilibrium predecessor, finding the 8% reduction from tax distortions. Third, Jacobs and de Mooij (2015) and Jacobs and van der Ploeg (2019) showed in static models with heterogeneous agents and a uniform lump-sum transfer that the MCF equals 1. This paper extends this insight to a fully dynamic climate-economy framework with general equilibrium and a rich model of household heterogeneity. The key innovation relative to Barrage (2020) is agent heterogeneity, which both provides microfoundations for distortionary taxation and significantly changes the quantitative implications for optimal carbon taxes. Relative to Jacobs and de Mooij (2015), the contribution is the dynamic setting, the linkage to the DICE climate module, and the full quantitative characterization including distributional welfare analysis and multiple sources of heterogeneity.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q13. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary policy implication is that a carbon tax should be set approximately equal to the SCC (Pigouvian level) and the associated revenue should be split roughly equally between increasing lump-sum transfers and reducing distortionary labor taxes — rather than following the double-dividend prescription of using all revenue to reduce distortionary taxes. This combination is both more efficient (the MCF argument) and more equitable (progressive in the short run). The scope conditions are: (a) the result applies under a utilitarian welfare criterion with linear income taxes and a uniform lump-sum transfer; (b) it requires that the government can optimize the level of lump-sum transfers for redistribution; (c) the approximately Pigouvian result is quantitatively robust to alternative damage functions, fiscal pressure, and energy demand heterogeneity, but the degree to which inequality lowers the carbon tax depends sensitively on the IES/inequality aversion parameter sigma; (d) the calibration is designed to capture US conditions assuming that the US internalizes the full global impact of its emissions (strategic considerations are abstracted away); (e) heterogeneous environmental damage sensitivity does not affect the utilitarian optimum, but would increase the optimal carbon tax under a more inequality-averse social planner.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Marginal Cost of Public Funds (MCF)&lt;/strong&gt;: The ratio of the public (planner&amp;rsquo;s shadow price on the resource constraint) to the private (aggregate welfare-weighted) marginal utility of consumption. In this paper, it captures the divergence between second-best and first-best pollution taxes due to fiscal distortions. With heterogeneous agents and an optimized uniform lump-sum transfer, the MCF averages to 1 over time under balanced-growth preferences, implying that tax distortions do not systematically push the carbon tax below the Pigouvian level — unlike in the representative-agent setting where the MCF exceeds 1.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pigouvian tax (second-best)&lt;/strong&gt;: In this paper&amp;rsquo;s context, the Pigouvian tax refers to the pollution tax equal to the social cost of pollution (the discounted present value of marginal production and utility damages), evaluated at the second-best allocation rather than the first-best. When the MCF equals 1 (as it approximately does in the heterogeneous-agent setting), the second-best optimal pollution tax is equal to this second-best Pigouvian level, which may itself differ from the first-best Pigouvian level due to residual consumption inequality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Cost of Carbon (SCC)&lt;/strong&gt;: The present discounted value of marginal climate damages (both production and utility losses) from emitting one additional ton of CO2, converted into consumption units using the social marginal utility of consumption. In the paper, the SCC corresponds to the case where the MCF is set to 1 in every period, and it is affected by consumption inequality through its effect on the social marginal utility of consumption. With sigma &amp;gt; 1, residual inequality raises the opportunity cost of abatement, reducing the SCC by 3.9% in the baseline calibration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Double-dividend hypothesis (weak)&lt;/strong&gt;: The claim that it is optimal to use the entire proceeds of a carbon tax to reduce existing distortionary taxes, yielding both an environmental dividend (less pollution) and an efficiency dividend (lower tax distortions). The paper shows this does not hold with heterogeneous agents: because distortionary taxes serve a redistributive purpose, reducing them at the margin has a welfare cost (increased inequality), so the planner optimally splits revenue between tax reduction and increased transfers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ramsey problem (climate-economy)&lt;/strong&gt;: The government&amp;rsquo;s optimization problem in this paper: maximizing utilitarian social welfare over an infinite horizon by choosing paths for linear taxes on labor income, capital income, energy, and pollution, plus a uniform lump-sum transfer, subject to households&amp;rsquo; optimality conditions (implementability constraints), resource constraints, climate dynamics from DICE, and abatement technology constraints. The approach extends Werning (2007) to a dynamic climate-economy context.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implementability condition&lt;/strong&gt;: The constraint in the Ramsey problem that captures each household&amp;rsquo;s lifetime budget constraint in terms of aggregate variables and market weights. It requires that the present value of a household&amp;rsquo;s consumption minus labor income equals its initial assets plus its share of the present value of lump-sum transfers, evaluated using the social marginal utilities implied by the planner&amp;rsquo;s choice of taxes. The shadow cost of this constraint for each household type (theta_i) determines the MCF through its covariance with a fiscal externality term.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Residual inequality&lt;/strong&gt;: The level of inequality that remains after the planner has optimally set all income taxes and the lump-sum transfer — i.e., the inequality that cannot be eliminated because individualized lump-sum transfers are not feasible and only linear instruments are available. In the paper, it is this residual inequality (not total inequality) that affects the optimal carbon tax: the carbon tax responds to the inequality that income-tax policy cannot address, not to the underlying productivity or wealth dispersion per se.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Balanced-growth preferences&lt;/strong&gt;: A preference specification of the form u(c, h, Z) = [c(1 - varsigma*h)^gamma]^(1-sigma)/(1-sigma) + u_hat(Z), with 1/sigma the intertemporal elasticity of substitution. This specification ensures that the economy admits a balanced growth path and plays a key role in the paper&amp;rsquo;s theoretical results: under balanced-growth preferences, the welfare-weighted average MCF equals 1 from period 0, and when IES = 1 (sigma = 1) the MCF is exactly 1 in every period.&lt;/p&gt;</description></item><item><title>Population and Welfare: Measuring Growth when Life is Worth Living</title><link>https://macropaperwarehouse.com/papers/population-and-welfare-measuring-growth-when-life-is-worth-living/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/population-and-welfare-measuring-growth-when-life-is-worth-living/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;The paper asks how much economic progress looks different when one applies a total utilitarian welfare criterion — counting every person&amp;rsquo;s flow utility — rather than the standard per-capita consumption measure. The motivation is both philosophical and practical: philosophers have long debated whether the number of people matters for social welfare alongside average living standards, yet growth economists have almost exclusively used the per-capita approach. The authors do not adjudicate the debate; they quantify its stakes across a broad cross-country sample.&lt;/p&gt;
&lt;p&gt;The framework is parsimonious. Under total utilitarianism, flow social welfare is W = N·u(c). Consumption-equivalent (CE) welfare growth is gλ = v(c)·gN + gc, where gN is population growth, gc is per-capita consumption growth, and v(c) = u(c)/[u&amp;rsquo;(c)·c] is the value of a year of life in units of per-capita consumption. Diminishing marginal utility guarantees v(c) &amp;gt; 1: each percentage point of population growth is worth more than a percentage point of per-capita consumption growth. The baseline utility function is u(c) = ū + log(c). The key parameter ū is calibrated to the U.S. Environmental Protection Agency&amp;rsquo;s Value of Statistical Life (VSL) of $7.4 million (2006 prices): dividing by remaining life expectancy (~40 years) and U.S. per-capita consumption of $38,000 gives v(c_US,2006) ≈ 4.87. Normalizing U.S. 2006 consumption to 1 sets ū = 4.87. Under log utility, v(c) = ū + log(c) rises with living standards: it averaged roughly 2 in 1820 for the U.S. and nearly 5 by 2019, and ranges from about 2 (Ethiopia) to 5 (U.S.) across countries in 2019. The world-sample average of v(c) over 1960–2019 is 2.7.&lt;/p&gt;
&lt;p&gt;Applying the formula to Penn World Table 10.0 data for 101 countries over 1960–2019 yields the following main findings. CE welfare growth averages 6.2% per year versus 2.1% per year for per-capita consumption growth; at 2.1% growth per-capita consumption doubles every 33 years, but under the CE measure social welfare doubles every 12 years. Population growth (averaging 1.8% per year) accounts for 66% of CE welfare growth unweighted across countries, and 51% weighting by country population (which gives China a large weight). For the United States specifically, CE welfare growth averages 6.5% per year versus 2.2% for per-capita consumption growth. Country rankings shift dramatically. Mexico rises from the 35th to the 88th percentile (CE welfare growth: 8.6% per year; population contribution: 79%). South Africa and Kenya similarly move up sharply. Germany falls to the 11th percentile, Japan to the 32nd, and China to the 44th — all below the United States. The cross-country correlation between CE welfare growth and per-capita consumption growth is 0.51; with population growth, 0.29. Over the very long run (1500–2018, Maddison data), per-capita consumption rose 20-fold (0.6% per year) while CE welfare rose 3,700-fold (1.6% per year) due to population growing at 0.5% per year scaled by v(c).&lt;/p&gt;
&lt;p&gt;Robustness checks confirm the core result. Halving the baseline VSL (setting ū = 2.4) still leaves population contributing 38% of CE welfare growth on average. Incorporating within-country consumption inequality under a log-normal distribution lowers CE welfare growth by an average of just 10 basis points (from 6.1% to 6.0% for 1980–2007). Attributing migrants to source rather than destination countries produces a correlation of 0.92 between adjusted and baseline CE welfare growth rates. Decomposing population growth, roughly three-quarters of actual population growth in a 24-country subsample reflected increases in the number of lives lived (i.e., births), not longevity extension — so the welfare contribution of births exceeds that of rising longevity.&lt;/p&gt;
&lt;p&gt;An extended model adds leisure, parental altruism toward children&amp;rsquo;s consumption and human capital, and endogenous fertility. Using time-use data from six countries (U.S. 2003–2019; Netherlands 1975–2006; Japan 1991–2016; South Korea 1999–2019; Mexico 2006–2019; South Africa 2000–2010), the extension modestly reduces the population share of CE welfare growth in most countries. The main reason is that parental altruism &amp;ldquo;double-counts&amp;rdquo; children&amp;rsquo;s consumption in the social welfare function, making consumption growth relatively more valuable and thus scaling down the weight on population growth. Rising quality of children (human capital) roughly offsets falling fertility in most countries, leaving net CE welfare growth little changed. Mexico is the sharpest exception: under extended preferences, CE welfare growth falls from 6.5% to 3.3% because of sharply declining leisure and little offsetting gain in children&amp;rsquo;s quality. Japan and South Korea also see smaller population shares under the extended model. The qualitative conclusion — that population growth is a major contributor to CE welfare growth — survives across all specifications.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-identification-strategy-and-what-does-it-rely-on"&gt;Q1. What is the paper&amp;rsquo;s identification strategy and what does it rely on?&lt;/h3&gt;
&lt;p&gt;This is a welfare accounting exercise rather than a causal identification exercise. There is no identification problem in the traditional econometric sense: the authors are computing a welfare index given a social welfare function and observed data on population and consumption. The two key inputs are (1) data on population and consumption per capita from the Penn World Table 10.0 for 101 countries over 1960–2019, and (2) a calibrated value of the parameter ū, which is the value of a year of life measured in units of per-capita consumption. The calibration of ū is anchored to external VSL estimates (EPA&amp;rsquo;s $7.4 million in 2006 prices), divided by life expectancy and per-capita consumption. The paper is explicit that it cannot make causal policy recommendations because it says nothing about the production side of the economy or externalities (pollution, ideas, human capital spillovers).&lt;/p&gt;
&lt;h3 id="q2-what-is-vc-and-why-does-it-matter-so-much-for-the-results"&gt;Q2. What is v(c) and why does it matter so much for the results?&lt;/h3&gt;
&lt;p&gt;v(c) = u(c)/[u&amp;rsquo;(c)·c] is the value of a year of life measured in consumption-equivalent units — specifically, how many years&amp;rsquo; worth of per-capita consumption an individual would require as compensation for losing one year of life. Under log utility u(c) = ū + log(c), v(c) = ū + log(c), so it rises with the log of consumption. The key implication is that each percentage point of population growth is worth v(c) percentage points of per-capita consumption growth. Since v(c) empirically ranges from about 2 (Ethiopia) to 5 (rich countries), and averages 2.7 over 1960–2019 across 101 countries, population growth receives a substantial weight in the CE welfare measure. Without this amplification (i.e., if v = 1 so CE welfare equals aggregate consumption growth), population would still account for 36% of all growth.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-distinguish-its-approach-from-simply-using-aggregate-total-consumption-growth"&gt;Q3. How does the paper distinguish its approach from simply using aggregate (total) consumption growth?&lt;/h3&gt;
&lt;p&gt;Using aggregate consumption growth is equivalent to setting v(c) = 1 in the CE welfare formula — that is, weighting population growth and consumption growth equally. The paper shows that, under a total utilitarian welfare function with diminishing marginal utility, the correct weight on population growth is v(c) &amp;gt; 1, not 1. So aggregate consumption growth systematically understates the contribution of population growth to welfare: in a country with average v(c) = 2.7, a percentage point of population growth should receive 2.7 times the weight of a percentage point of consumption growth, not equal weight.&lt;/p&gt;
&lt;h3 id="q4-what-threats-to-the-baseline-calibration-of-vc-does-the-paper-address"&gt;Q4. What threats to the baseline calibration of v(c) does the paper address?&lt;/h3&gt;
&lt;p&gt;The paper addresses four main threats. First, VSL uncertainty: it considers halving and raising the baseline VSL by 50%, yielding ū = 2.4 and ū = 7.3 respectively. Population&amp;rsquo;s share of CE welfare growth remains 38% even under the low VSL. Second, functional form: it considers CRRA utility with risk-aversion γ = 2 rather than log (γ = 1), which lowers the population share to 40% (from 53% baseline, population-weighted). Third, whether v(c) should be constant rather than income-varying: rows 7–9 of Table 3 test constant v = 4.87, v = 2.7, and v = 1. Even v = 1 (aggregate consumption growth) gives population a 36% share. Fourth, whether the marginal VSL used to calibrate the model overstates the average value of a birth (since a birth produces a new life from the start, not an added year for a middle-aged person). The paper acknowledges this concern but treats the calibration as a natural baseline and explores lower ū as a robustness check.&lt;/p&gt;
&lt;h3 id="q5-how-does-within-country-inequality-affect-the-results"&gt;Q5. How does within-country inequality affect the results?&lt;/h3&gt;
&lt;p&gt;Under log utility and a log-normal distribution of individual consumption, CE welfare growth becomes gλ = [ū + log(c_t) - (1/2)σ²_t]·gN + gc - σ²_t·gσ, where σ² is the cross-sectional variance of log consumption. Inequality enters in two ways: it reduces the weight on population growth (because average utility is lower than utility of average consumption under concavity), and increases in inequality directly reduce CE welfare growth. Implementing this for 90 countries over 1980–2007, the mean adjustment is -10 basis points (6.1% to 6.0%), with a mean absolute deviation of 18 basis points. The adjustment is sizable for South Africa (−0.83 pp, due to very high inequality relative to U.S. 2006 baseline) and small or positive for Brazil (falling inequality over the period) and Ethiopia.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-treat-migration-and-does-it-matter"&gt;Q6. How does the paper treat migration, and does it matter?&lt;/h3&gt;
&lt;p&gt;The baseline credits population growth to the country of residence. The migration-adjusted measure reassigns migrants to their country of birth: it adds the flow utility of out-migrants (at destination-country consumption levels) and subtracts the flow utility of in-migrants (at destination-country consumption levels) from each country&amp;rsquo;s welfare. Using the World Bank Global Bilateral Migration Database for 81 countries over 1960–2000, migration-adjusted and baseline CE welfare growth rates have a correlation of 0.92. The adjustment matters most for specific countries — it raises welfare growth for net out-migrant countries like Mexico and the Philippines (since their emigrants consume more abroad) and lowers it for net in-migrant countries — but it does not alter the broad conclusion that population growth matters greatly.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-decomposition-of-population-growth-into-fertility-and-longevity-effects"&gt;Q7. What is the decomposition of population growth into fertility and longevity effects?&lt;/h3&gt;
&lt;p&gt;For a 24-country subsample (from the Human Mortality Database combined with World Bank migration data), the authors compute counterfactual population growth holding age-specific death rates constant at their initial-period values. Population-weighted, actual annual population growth is 0.72% versus a counterfactual of 0.53% with fixed longevity. So roughly three-quarters of population growth (and therefore three-quarters of the CE welfare contribution of population growth) reflected an increase in the number of lives lived (births minus deaths under fixed mortality), not gains in longevity. Italy and Japan are outliers: falling death rates (i.e., longevity gains) account for about three-quarters of their population growth. For context, Jones and Klenow (2016) attribute ~1% per year of CE welfare growth to rising longevity for 128 countries over 1980–2007; the total population growth contribution here (~3% per year, population-weighted from Table 1) substantially exceeds the longevity-only benchmark.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-extended-model-with-parental-preferences-add-and-what-are-its-main-results"&gt;Q8. What does the extended model with parental preferences add, and what are its main results?&lt;/h3&gt;
&lt;p&gt;The extended model incorporates adult leisure, parental altruism toward children&amp;rsquo;s consumption and human capital, endogenous fertility, and children&amp;rsquo;s utility as separate welfare contributors. Social welfare is W = N_p·u(c_p, l, c_k, h_k, b) + N_k·ũ(c_k), where b is fertility per adult, l is adult leisure, and h_k is children&amp;rsquo;s human capital. CE welfare growth is computed using first-order conditions from parents&amp;rsquo; utility maximization — specifically, the MRS between leisure/fertility/human capital and consumption can be measured from time-use data, which provides the welfare weights on each term. Key parameters: parental altruism weight α = 2/3 (calibrated to USDA household spending data), diminishing-returns-to-fertility parameter θ = 0.8, and children&amp;rsquo;s human capital elasticity η = 0.21 (from Mincer estimates in Lee, Roys, and Seshadri 2024). Main results: (1) Population growth remains an important contributor to CE welfare growth in most countries. (2) The population share falls somewhat because parental altruism double-counts children&amp;rsquo;s consumption, raising the relative weight on consumption growth. (3) Rising children&amp;rsquo;s quality (human capital, measured via real wage growth) roughly offsets falling fertility in most countries. (4) Mexico is the main exception: CE welfare growth drops from 6.5% to 3.3% due to falling leisure and little offset from rising children&amp;rsquo;s quality. (5) Japan&amp;rsquo;s population share falls further, turning slightly negative in some specifications.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-philosophical-foundation-and-what-is-the-repugnant-conclusion-objection"&gt;Q9. What is the philosophical foundation and what is the &amp;lsquo;repugnant conclusion&amp;rsquo; objection?&lt;/h3&gt;
&lt;p&gt;The total utilitarian social welfare function W = N·u(c) follows from three axioms: same-number Pareto (welfare ordering respects Pareto improvements for fixed populations), non-anti-egalitarianism (society does not prefer inequality), and mere addition (adding a person who values living, holding others&amp;rsquo; utilities constant, does not reduce welfare). These axioms, as surveyed by Kuruc, Budolfson, and Spears (2022), together imply total utilitarianism and rule out diminishing-returns-to-population approaches (e.g., W = N^α·u(c) for α &amp;lt; 1). The repugnant conclusion (Parfit 1984) holds that total utilitarianism could justify very large populations of people whose lives are barely worth living. The authors respond that their calculations are local — reflecting only actual births and deaths over 1960–2019 — not arbitrary expansions. They also note that 29 philosophers and economists (Zuber et al. 2021) have argued the repugnant conclusion is not a reason to reject totalism. The per-capita approach has its own problems: it implies one should remove people whose utility is valuable but below average (and implies the &amp;lsquo;sadistic conclusion&amp;rsquo; under certain conditions).&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-jones-and-klenow-2016"&gt;Q10. How does this paper relate to Jones and Klenow (2016)?&lt;/h3&gt;
&lt;p&gt;Jones and Klenow (2016) is the closest predecessor. That paper computes CE welfare measures incorporating consumption, leisure, life expectancy, and inequality, but in a per-capita framework — it measures individual living standards, not aggregate social welfare. The key difference here is moving from per-capita utility to total utilitarian welfare by multiplying individual utility by population, which introduces the v(c)·gN term. The current paper&amp;rsquo;s baseline is also simpler (consumption only) with an extended version that adds leisure and parental preferences. Jones and Klenow attribute ~1% per year of CE welfare growth to rising longevity for 128 countries over 1980–2007; the present paper shows total population growth (birth + longevity channels combined) contributes ~3% per year (population-weighted), substantially more.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper explicitly states it cannot make policy recommendations because it says nothing about the production side of the economy or about externalities (pollution, ideas externalities, human capital spillovers). Whether fertility rates are &amp;rsquo;too low&amp;rsquo; or the demographic transition raised or reduced social welfare requires estimating these externalities, which is beyond the paper&amp;rsquo;s scope. The paper is a measurement exercise, not an optimal policy analysis. Nonetheless, the results have implications for policy questions that depend on which welfare criterion is adopted: optimal fertility policy, the welfare cost of HIV/AIDS or other mortality shocks, the assessment of China&amp;rsquo;s One Child Policy, the welfare calculus of climate change mitigation, and the social returns to nonrival knowledge (which benefit a larger future population under totalism). The scope condition throughout is that the paper evaluates actual births and deaths over a historical period; the results do not directly speak to the desirability of population expansion beyond what occurred.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-main-robustness-checks-run-and-what-do-they-show"&gt;Q12. What are the main robustness checks run and what do they show?&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;VSL calibration: halving (ū = 2.4) or raising by 50% (ū = 7.3) the baseline VSL — population share falls to 38% or rises to higher levels, but population remains important in all cases. 2. CRRA utility with γ = 2 (more concave): population share falls to 40% population-weighted (from 53%). 3. Constant v(c): results with v = 4.87 (U.S. 2006 level), v = 2.7 (world average), and v = 1 (aggregate consumption growth) all confirm that population growth matters, with v = 1 still giving a 36% population share. 4. Inequality: mean absolute adjustment of 18 basis points; largest adjustment for South Africa (−0.83 pp). 5. Migration: correlation 0.92 between adjusted and baseline. 6. Birth vs. longevity decomposition: ~75% of population growth (population-weighted) is from net new lives, not longevity. 7. Extended preferences (time-use data): qualitative results survive; population share falls modestly except for Japan, South Korea, and Mexico.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="q13-what-heterogeneity-across-countries-and-time-periods-is-documented"&gt;Q13. What heterogeneity across countries and time periods is documented?&lt;/h3&gt;
&lt;p&gt;Cross-country: CE welfare growth ranges from just above 2% per year for the slowest-growing countries to more than 10% per year for the fastest. The correlation between CE welfare growth and per-capita consumption growth is 0.51; with population growth, 0.29. The value of v(c) ranges from about 2 (Ethiopia) to 5 (U.S.) in 2019, tracking consumption levels. Countries with high population growth (Mexico, Brazil, South Africa, Kenya, Sub-Saharan Africa more broadly) move up sharply in the growth rankings; countries with slow population growth (Germany, Japan, China) fall sharply. Within time: Japan shows CE welfare growth falling from 9.7% per year in the 1960s to −0.3% in the 2010s as both consumption growth and population growth slowed and then turned negative. China&amp;rsquo;s CE welfare growth fell more modestly from a 7.0% peak in the 1990s to 5.7% in the 2010s because rising v(c) partly offset slower population growth. Sub-Saharan Africa maintained stable population growth (~2.5% per year across all decades) and saw rising consumption in the 2000s and 2010s, producing CE welfare growth above 8% in the 2010s. Extended-model results (six-country sample): Mexico is a major outlier with falling leisure driving CE welfare growth down to 3.3% (from 6.5% baseline); Japan and South Korea have very small or near-zero population shares under extended preferences.&lt;/p&gt;
&lt;h3 id="q14-what-does-the-very-long-run-historical-exercise-show"&gt;Q14. What does the very long-run historical exercise show?&lt;/h3&gt;
&lt;p&gt;Using Maddison Project data (de Pleijt and van Zanden 2020) from 1500 to 2018 for the world as a whole, per-capita consumption rose by a factor of 20 (0.6% per year) and aggregate consumption rose by a factor of 163 (1.1% per year). CE welfare rose by a factor of 3,700 — a 1.6% per year average annual growth rate. The power of compounding over 500 years causes a difference of only 1 percentage point per year between CE welfare growth (1.6%) and per-capita consumption growth (0.6%) to produce a ratio of 185:1 in cumulative outcomes (3,700-fold versus 20-fold). Population growth accounts for 61% of CE welfare growth over this very long run.&lt;/p&gt;
&lt;h3 id="q15-what-data-sources-are-used-and-what-are-the-sample-restrictions"&gt;Q15. What data sources are used and what are the sample restrictions?&lt;/h3&gt;
&lt;p&gt;Baseline: Penn World Table 10.0 (Feenstra, Inklaar, and Timmer 2015) for 101 countries over 1960–2019 (starting from 111 countries and dropping those flagged as outliers in any year). Consumption is private plus government consumption. Inequality data: Jones and Klenow (2016), 90 countries, 1980–2007. Migration data: World Bank Global Bilateral Migration Database (1960, 1970, 1980, 1990, 2000), 81 countries. Birth/death decomposition: Human Mortality Database combined with World Bank migration data, 24 countries. Long run: Maddison Project (de Pleijt and van Zanden 2020), 1500–2018, with consumption proxied as 0.8 times per-capita GDP. Extended model: time-use surveys for U.S. (2003–2019), Netherlands (1975–2006), Japan (1991–2016), South Korea (1999–2019), Mexico (2006–2019), South Africa (2000–2010); Penn World Table for population, consumption, and hours worked; World Bank for number of children (0–19 years); USDA spending-on-children data (Lino 2011) for parental altruism calibration; Lee, Roys, and Seshadri (2024) Mincer estimates for human capital elasticity.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Consumption-Equivalent (CE) Welfare Growth&lt;/strong&gt;: The rate at which per-capita consumption would need to grow, holding population constant, to produce the same increase in total utilitarian social welfare as the observed combination of population growth and per-capita consumption growth. Formally gλ = v(c)·gN + gc. It is analogous to GDP in being a flow measure (not a present-discounted sum across generations).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Value of a Year of Life, v(c)&lt;/strong&gt;: The ratio u(c)/[u&amp;rsquo;(c)·c], equal to individual utility divided by the marginal utility of consumption times consumption. It converts the value of being alive for one year into consumption-equivalent units. Under log utility u(c) = ū + log(c), v(c) = ū + log(c), so it rises with living standards. It is calibrated from Value of Statistical Life estimates: v(c_US,2006) ≈ 4.87, meaning a year of American life in 2006 was worth approximately 4.87 years&amp;rsquo; worth of per-capita consumption.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Total Utilitarian Social Welfare Function&lt;/strong&gt;: W = N·u(c): social welfare is the sum of all individuals&amp;rsquo; flow utilities. It treats every person&amp;rsquo;s utility symmetrically and linearly in population, so adding a person who values living always increases welfare. This contrasts with the per-capita (average utilitarian) approach (which implicitly sets the welfare weight on population to zero) and intermediate approaches that weight population with diminishing returns (W = N^α·u(c), α &amp;lt; 1).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mere Addition (Axiom)&lt;/strong&gt;: One of three axioms (with same-number Pareto and non-anti-egalitarianism) whose conjunction implies the total utilitarian social welfare function for variable populations. It states that, holding the utilities of existing persons constant, adding a new person who values living does not reduce social welfare. The axiom directly rules out the average (per-capita) utilitarian criterion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Repugnant Conclusion&lt;/strong&gt;: Parfit&amp;rsquo;s (1984) critique of total utilitarianism: maximizing the sum of utilities could in principle justify an extremely large population of people whose lives are just barely worth living (positive but tiny utility), since total utility could exceed that of a smaller population with high per-capita welfare. The paper responds that its welfare calculations are local (reflecting actual historical births and deaths), not global maximization exercises, and cites the Zuber et al. (2021) consensus that this is not a decisive objection.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parental Altruism Weight (α, θ)&lt;/strong&gt;: Parameters governing how parents value children&amp;rsquo;s consumption relative to their own in the extended model. Under Assumption 1, u includes the term αb^θ·log(c_k): α governs the overall weight on children&amp;rsquo;s consumption and θ governs diminishing returns as the number of children rises. Calibrated to α = 2/3 (from USDA household spending ratios with two children) and θ = 0.8 (from cross-family variation in per-child spending). Parental altruism causes double-counting of children&amp;rsquo;s consumption in the social welfare function, upweighting consumption growth and reducing the relative contribution of population growth to CE welfare.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Double-Counting of Children&amp;rsquo;s Consumption&lt;/strong&gt;: When parents are altruistic, their utility depends on children&amp;rsquo;s consumption as well as their own; and children receive direct utility from their consumption too. In the CE welfare growth formula, this means a rise in children&amp;rsquo;s consumption raises welfare through two channels simultaneously (parental and child utility), so each unit of consumption growth is &amp;lsquo;worth more&amp;rsquo; relative to population growth. This is why the extended model&amp;rsquo;s population term is smaller than the baseline&amp;rsquo;s: consumption growth is valued more heavily under parental altruism, scaling down the consumption-equivalent weight on population growth.&lt;/p&gt;</description></item><item><title>Property rights, fiscal capacity, and social capacity: The lasting impact of the Taiping Rebellion</title><link>https://macropaperwarehouse.com/papers/property-rights-fiscal-capacity-and-social-capacity-the-lasting-impact-of-the-taiping-rebellion/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/property-rights-fiscal-capacity-and-social-capacity-the-lasting-impact-of-the-taiping-rebellion/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: How do civil wars affect long-term development, and through which institutional mechanisms? The paper studies the Taiping Rebellion (1850-1864) in Qing China, one of history&amp;rsquo;s deadliest civil wars (at least ~20 million deaths, with some estimates of 70-100 million), as a critical juncture in China&amp;rsquo;s path to modernity. It matters because the rebellion generated large, persistent regional institutional variation that can help explain what the authors call the &amp;ldquo;Intra-China Divergence&amp;rdquo; — regional GDP-per-capita gaps as large as 27-to-1 (Dongguan vs. Tianshui, 2010) that rival the world&amp;rsquo;s largest inter-regional gaps.&lt;/p&gt;
&lt;p&gt;Data and design: A prefecture-level (occasionally county-level) panel covering 266 prefectures in China proper (1820 delineation). 55 prefectures fell under Taiping control (treatment) — split into 37 &amp;ldquo;Early Taiping&amp;rdquo; prefectures (occupied up to 1859, in Anhui/Jiangxi/Hubei, ambiguous land rights) and 18 &amp;ldquo;Late Taiping&amp;rdquo; prefectures (occupied from 1860, in Jiangsu/Zhejiang, stronger land rights) — and 211 control prefectures. Population is observed at seven points (1820, 1851, 1880, 1910, 1953, 1982, 2000). The core strategy is difference-in-differences (1820 reference year, prefecture and year fixed effects), supplemented by propensity-score matching (135-prefecture matched sample), a spatial autoregressive (SAR) model, and an instrumental-variable strategy using the longitude of the prefectural seat (motivated by the Taiping Navy&amp;rsquo;s eastward-along-the-Yangtze military strategy; first-stage F-statistics above 20).&lt;/p&gt;
&lt;p&gt;Main quantitative findings (with scope conditions): (1) Population: The rebellion caused large, permanent population losses. The Taiping DID coefficient is -0.45 in 1880 (a 36% lower population growth rate vs. control) and -0.51 in 1953 (40% lower) — no convergence. Crucially, in the matched sample Late Taiping areas recovered (no significant long-run population gap vs. control) while Early Taiping areas did not (an immediate ~30% drop in 1880 plus further decline). (2) Property rights: In 1915 county data, the idle-land share is 3.6 percentage points higher in Early Taiping than control counties, while Late Taiping is not significantly different from control — supporting the property-rights hypothesis. (3) Fiscal capacity (likin): Taiping areas collected ~12 times (e^2.5) as much likin per 1,000 sq km as control areas in 1869-1879, still 3.7 times as much in 1922-1925. Late Taiping areas had even higher intensity (22.2x in 1869-1879; 6.1x in 1922-1925) than Early Taiping (9.0x; 2.7x). (4) Social capacity (charities): On average the rebellion had no significant effect, but Late Taiping areas saw charity growth ~56 percentage points (44 log points) above control by 1880, rising to ~78 percentage points (58 log points) by mid-20th century. (5) Long-term development: Driven entirely by Late Taiping areas — 1982 agricultural+industrial output per capita 90% higher (64 log points), 2010 GDP per capita 87% higher (63 log points), and 2010 fiscal revenue per capita 203% higher (111 log points) than control; Early Taiping is statistically indistinguishable from control. Late Taiping counties also show higher post-1895 industrial firm entry. (6) Civic outcomes and resilience: Using CGSS 2010, Late Taiping residents show higher trust in personal networks and greater civic engagement (political attention, local participation). During the Great Famine (1959-1961), Taiping areas had 6.9% larger survivor cohorts; the effect is 28% stronger in Late Taiping (8.4%) than Early Taiping (6.5%).&lt;/p&gt;
&lt;p&gt;Implications: Violent conflict can leave lasting positive institutional imprints — through property rights, decentralized local fiscal capacity (&amp;ldquo;war made the state&amp;rdquo; at the local level), and elite-led social capacity — conditional on favorable initial conditions (strong gentry, wealthier commercial regions). The authors argue cultivating civil society and social capacity could yield large payoffs given China&amp;rsquo;s strong-state/weak-society configuration.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the core identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The baseline is a difference-in-differences comparing Taiping vs. control prefectures over 1820-2000, with prefecture and year fixed effects and 1820 as the reference year. Identification rests on parallel pre-trends: the Taiping coefficient in 1851 (pre-rebellion) is small and insignificant, indicating no differential selection conditional on controls. The main threats are: (i) the binary Taiping measure aligning with provincial boundaries and picking up broad regional dynamics; (ii) control-group contamination because some control prefectures were temporarily conquered (but not governed) by the Taiping Army; (iii) spatial spillovers between neighbors (Tobler&amp;rsquo;s law / Kelly 2019 critique); (iv) omitted subsequent historical events; and (v) omitted variables differing systematically between treated and control areas. The authors address these with dosage measures (battles, occupation months), matching, a SAR model, an IV (longitude), explicit controls for the Taiping conquest, an adjacent-treatment indicator, leave-one-province-out checks, and controls for many other historical events.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-instrumental-variable-strategy-work-and-why-might-longitude-be-valid"&gt;Q2. How does the instrumental-variable strategy work and why might longitude be valid?&lt;/h3&gt;
&lt;p&gt;Longitude of the prefectural seat instruments for the Taiping dummy. Relevance: the Taiping leaders&amp;rsquo; July 1852 military plan was to march eastward along the Yangtze, capture Jiangning (Nanjing), and expand from there using their dominant navy — so eastern (higher-longitude) prefectures were far more likely to fall under Taiping rule (Table 1 confirms Taiping prefectures have significantly larger longitudes; first-stage F-statistics above 20, Shea&amp;rsquo;s partial R-squared above 0.1). Exclusion: prefecture fixed effects absorb time-invariant geographic advantages, and year-dummy interactions with key geography (distances to coastline, Grand Canal, Yangtze) allow flexible time-varying geographic effects; conditional on these, longitude is argued to be excludable. IV estimates are larger in magnitude than OLS but qualitatively confirm a persistent negative population effect (robust to Anderson-Rubin weak-IV inference). The authors caution that omitted determinants correlated with longitude cannot be fully ruled out.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-four-hypotheses-and-how-are-they-distinguished-empirically"&gt;Q3. What are the four hypotheses and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;(1) Property-rights hypothesis: Late Taiping areas (post-1860 &amp;lsquo;direct tenant payment&amp;rsquo; system creating de facto/de jure tenant ownership) had better-defined land rights than Early Taiping areas (collapsed landlord system, lost deeds, anti-rent movements), so should have less idle land and faster population recovery — tested via the 1915 idle-land cross-section and the Early-vs-Late population DID. (2) Likin-as-fiscal-capacity hypothesis: Qing fiscal decentralization and the likin tax (introduced 1853) strengthened local fiscal capacity, persistently higher in Taiping (especially Late Taiping) areas — tested via the likin-intensity DID. (3) Social-change hypothesis: elite-led militias and reconstruction spurred charities (&amp;lsquo;benevolent halls&amp;rsquo;/shantang) as bridging social capital, especially in Late Taiping areas — tested via charity-stock DID and by adding charities as a mediator in long-term regressions. (4) Social-cohesion-and-civic-engagement hypothesis: forged social capital persists, raising modern trust/civic engagement and reducing Great Famine deaths — tested via CGSS 2010 and famine-survivor cohort ratios.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented"&gt;Q4. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;The central heterogeneity is Early vs. Late Taiping. Early Taiping areas (Anhui/Jiangxi/Hubei) suffered permanent population loss, higher idle land (+3.6pp), only modest likin gains, no charity growth, no long-term development advantage, and weaker famine resilience. Late Taiping areas (Jiangsu/Zhejiang) recovered population, had no excess idle land, far higher likin intensity (22x early period), large charity growth (+56 to +78pp), strong long-term development gains (90%/87%/203% in output/GDP/fiscal revenue), higher modern trust and civic engagement, and the strongest famine resilience (8.4% vs 6.5%). Industrialization heterogeneity is also temporal: no Early/Late firm-entry difference before 1895, but after the 1895 Treaty of Shimonoseki liberalized private industry, Late Taiping counties had more entry and Early Taiping fewer.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;For the population results: dosage interactions (log battles, log occupation months); excluding six most-intense-fighting prefectures (Wuchang, Songjiang, Anqing, Jiangning, Suzhou, Hangzhou); controlling for newly selected jinshi (civil-service quota channel); a SAR spatial model (after Pesaran cross-sectional-dependence tests); PSM matched sample; longitude IV with Anderson-Rubin inference; controls for seven other historical events (Guangxu Drought, Hui Revolt, Nian Rebellion, early-Republic conflicts, Sino-Japanese War, Chinese Civil War, missionary activity); explicit controls for Taiping conquest vs. regime; an adjacent-treatment indicator (Butts 2021) for spillovers; and leave-one-province-out exclusion. Long-term development results add SAR, matching, historical-event controls including the Cultural Revolution, and an &amp;lsquo;intermediate-term&amp;rsquo; 1930s industrialization check. Famine results are robust to alternative famine-severity measures, SAR, matching, and historical-event controls.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-mediation-analysis-handled-and-what-does-it-show"&gt;Q6. How is the mediation analysis handled and what does it show?&lt;/h3&gt;
&lt;p&gt;The authors add likin intensity (1880) and average charities (1880-1941) to cross-sectional long-term regressions, explicitly flagging these as endogenous &amp;lsquo;bad controls&amp;rsquo; (Angrist-Pischke 2009; Imai et al. 2011) to be interpreted cautiously as descriptive mediation. Findings: a one-SD increase in likin intensity is associated with +1.7pp middle-school completion, +4.8pp literacy, +5.3% schooling, and +12.2% (11.5 log points) GDP per capita in 2010. A one-SD increase in charities is associated with +15% 1982 output, +20% 2010 GDP, and +55% 2010 fiscal revenue per capita. Once charities are netted out, Late Taiping advantages in output, GDP, and fiscal revenue are attenuated by about 17%, 14%, and 22% respectively — highlighting the social-capacity channel.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-great-famine-resilience-result-connect-to-the-rebellion"&gt;Q7. How does the Great Famine resilience result connect to the rebellion?&lt;/h3&gt;
&lt;p&gt;Famine severity is measured by &amp;lsquo;Famine Control&amp;rsquo; = ratio of cohort size born during the famine (1959-1961) to cohort size born pre-famine (1954-1957) from the 1990 census 1% sample (higher = less severe). Taiping areas had a 6.9% larger survivor cohort than non-Taiping; the effect is 8.4% in Late Taiping vs. 6.5% in Early Taiping. Back-of-envelope, the Late Taiping experience would have &amp;lsquo;saved&amp;rsquo; ~31,374 people in an average prefecture (17% of the 1959-1961 cohort) vs. ~24,145 (13%) for Early Taiping. Controlling for political radicalism (reverse party-member density, -1*PMD, after Yang 1996) does not change the result. The mechanism: higher social capital made local officials more sympathetic/less radical in grain procurement and citizens better able to act collectively (paralleling Cao-Xu-Zhang 2022 on clan density and Hu-Yao-You 2023 on home-county officials).&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q8. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Prior Taiping studies examined narrower consequences: civil-service exam quotas (Li 2014), demographic and industrialization effects (Li and Ma 2016), migration and public goods (Hao and Xue 2017), and late-Qing power distribution (Bai, Jia, and Yang 2023). None addressed the rebellion&amp;rsquo;s enduring impacts on modern development, social trust, and Great Famine responses, nor the property-rights/fiscal-capacity/social-capacity mechanism triad. It complements Xue (2021) on Qing charities, generalized trust, and political participation, but extends to development outcomes. Against the European state-building literature (war strengthens central state capacity via centralization), this paper&amp;rsquo;s distinctive claim is that the Taiping Rebellion strengthened LOCAL fiscal capacity through DECENTRALIZATION, and expanded local social capacity that constrained the central state.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The benefits of war-induced institutions are conditional, not universal: they appeared chiefly in Late Taiping areas with a strong gentry class and favorable initial conditions for modern sectors (the wealthier, more commercial Lower Yangtze). The likin/fiscal-capacity benefits are explicitly stated to be conditional on strong gentry and good modern-sector initial conditions. The broad implication is that, given China&amp;rsquo;s very strong state but still weak society today, cultivating civil society and strengthening social capacity could yield particularly large long-term payoffs. The authors also caution (Appendix F.1) that likin could be distortionary taxation rather than fiscal capacity, arguing the fiscal-capacity interpretation is more relevant for long-term development.&lt;/p&gt;
&lt;h3 id="q10-what-significant-caveats-does-the-paper-acknowledge"&gt;Q10. What significant caveats does the paper acknowledge?&lt;/h3&gt;
&lt;p&gt;Long-term mechanisms cannot be exhaustively identified — likin and charities are endogenous outcomes, so mediation magnitudes are descriptive, not causal. History contains near-infinite interrelated events, so confounding cannot be fully eliminated (a fundamental limitation of all history-based work). The IV may have omitted correlates of longitude. Some 2SLS estimates for development outcomes were largely insignificant. The charity-stock measure assumes charities persisted once founded (no closure dates in the data). On property-rights persistence: using 2005 World Bank Enterprise Survey data they find no association between modern firms&amp;rsquo; perceived property-rights protection and Taiping regimes, suggesting the channel works through income effects rather than persistence of property rights per se.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Early vs. Late Taiping areas&lt;/strong&gt;: Early Taiping = prefectures occupied by the rebels up to 1859 (Anhui, Jiangxi, Hubei), where the old landlord system collapsed and land rights stayed ambiguous; Late Taiping = prefectures occupied from 1860 (Jiangsu, Zhejiang), where the Taiping introduced a &amp;lsquo;direct tenant payment&amp;rsquo; (作佃交粮) system and issued new deeds, granting tenants de facto/de jure ownership. This distinction is the paper&amp;rsquo;s central source of institutional variation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Likin (lijin)&lt;/strong&gt;: A local tax on trade and commerce introduced in 1853 (a transit tax on travelling merchants&amp;rsquo; goods plus a business tax on resident merchants), collected in a decentralized, province-specific way. In the paper it is the operational measure of local fiscal capacity (likin revenue per 1,000 sq km), not central state capacity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social capacity&lt;/strong&gt;: In the paper&amp;rsquo;s sense, the ability of society to act collectively, constrain the state, and empower its members — operationalized empirically by the stock of local charity organizations (&amp;lsquo;benevolent halls&amp;rsquo;/shantang) that functioned as bridging social capital across classes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Likin-as-fiscal-capacity hypothesis&lt;/strong&gt;: The claim that the rebellion-induced likin system durably raised LOCAL fiscal capacity (an instance of Tilly&amp;rsquo;s &amp;lsquo;war made the state&amp;rsquo; operating locally rather than centrally), which improved public-goods provision and long-run development — conditional on strong gentry and favorable modern-sector initial conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stationary bandit (applied to Late Taiping rulers)&lt;/strong&gt;: Borrowing Olson (1993): in Late Taiping areas the consolidated, longer-horizon Taiping regime behaved like a stationary bandit, lowering effective tax rates, encouraging land registration, and securing tenant property rights to expand the tax base and promote production, unlike the looting/confiscation of the early stage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Famine Control&lt;/strong&gt;: The paper&amp;rsquo;s local famine-severity measure: the ratio of the cohort born during the Great Famine (1959-1961) to the cohort born pre-famine (1954-1957) in the 1990 census; a higher value means less severe famine and more survivors, and it is less vulnerable to government understatement of famine deaths.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intra-China Divergence&lt;/strong&gt;: The authors&amp;rsquo; term for China&amp;rsquo;s persistent, very large regional disparities in economic performance (up to 27-to-1 in GDP per capita) despite all regions historically sharing similar Malthusian income levels — the macro puzzle the rebellion&amp;rsquo;s institutional legacy helps explain.&lt;/p&gt;</description></item><item><title>Remote Work and City Structure</title><link>https://macropaperwarehouse.com/papers/remote-work-and-city-structure/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/remote-work-and-city-structure/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Monte, Porcher, and Rossi-Hansberg ask why remote work surged abruptly and permanently after COVID-19 despite information-technology advances raising it only marginally between 1980 and 2019, why the change was so heterogeneous across cities, and what the welfare consequences are. Their answer is a coordination mechanism: working downtown (the CBD) yields productive interactions with other in-office workers but entails commuting/congestion costs, while remote work avoids those costs but forgoes agglomeration benefits. Because workers do not internalize the spillovers they confer, a worker prefers the office only if others commute too — generating, in a dynamic discrete-choice model with idiosyncratic preferences and fixed switching costs, the possibility of MULTIPLE stationary equilibria with different permanent commuter shares. A temporary shock (the pandemic) that drives commuters near zero can then select the low-commuting equilibrium permanently.&lt;/p&gt;
&lt;p&gt;The model is a dynamic monocentric city (disk-shaped, radially symmetric CBD, absentee landlords, Cobb-Douglas utility, Gumbel idiosyncratic shocks). Multiplicity arises (Proposition 4.3) when agglomeration forces are strong enough — the net strength delta + xi exceeds a threshold above theta + gamma/(2mu) — AND remote-work productivity relative to office productivity z/A lies in an intermediate &amp;ldquo;cone of multiplicity&amp;rdquo; (neither too low nor too high). The authors quantify city-specific parameters for U.S. CBSAs using pre-2019 data (Census/ACS 1980-2023, NLSY79 panel of 4,147 individuals 1998-2022, SafeGraph cell-phone mobility, Zillow ZHVI zip-code house prices). Estimation: transition elasticity s = 0.30 (elasticity of transitions into remote work = 3.09), fixed switching cost F = 1.78 (equivalent to giving up 83% of a year&amp;rsquo;s earnings); agglomeration externality delta with mean 0.067 (SD 0.022, 619 CBSAs); the amenity-vs-congestion difference xi - theta is statistically insignificant and set to zero.&lt;/p&gt;
&lt;p&gt;Stylized facts. Predicted remote-work share (controlling for composition) rose in the ACS from under 1% (1980) to 2.6% (2019), jumped to 12% (2020), peaked at 15% (2021), and fell to 11% (2023); NLSY shows a parallel path (1.4% in 1998 to 3.7% in 2018, 9.2% in 2020, 7.8% in 2022). The remote-work wage premium rose steadily but did NOT jump post-2018: ACS discount of 44.5% in 1980 became a 6.5% premium by 2022; NLSY discount fell from 18.5% (2000) to 3.1% (2022). A stable premium alongside a sudden quantity jump argues against pure productivity/preference shocks.&lt;/p&gt;
&lt;p&gt;Mobility/housing facts. All cities dropped to ~20% of pre-pandemic CBD trips in spring 2020 (about a 75% drop, unrelated to city size). Recoveries diverged: the 25 largest CBSAs (employment &amp;gt; 1.5M) stabilized at ~60% of January-2020 trips, while the 663 smallest (&amp;lt; 150K) returned fully to pre-pandemic levels by early 2021. New York and San Francisco stabilized near 40%; Madison, WI recovered fully. House-price distance gradients flattened ~0.01 everywhere by January 2021; the flattening persisted and stabilized around 0.095 by end-2024 in large cities but reversed in small ones.&lt;/p&gt;
&lt;p&gt;Results and welfare. Of 278 estimated CBSAs, 208 were inside their cone of multiplicity pre-pandemic; larger cities are systematically more likely to be inside (probit on log employment significant). The cone indicator predicts trip shortfalls (R-squared 0.144 alone, retaining significance with controls) and gradient flattening. Welfare: comparing high- vs low-commuting stationary equilibria for the 208 cone cities, the loss from switching is positive but modest — mean 2.3%, median 2.2%, range 1.2% to 4.0% (Table 3). Average wages fall sharply (15-35%) but option-value and commuting-cost savings offset most of it; net strength delta - gamma/(2mu) predicts the loss with R-squared 0.85. Cities with trips at 60% or less of pre-pandemic levels have an average welfare loss of 2.7%.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-economic-mechanism-and-how-does-it-generate-multiple-equilibria"&gt;Q1. What is the core economic mechanism, and how does it generate multiple equilibria?&lt;/h3&gt;
&lt;p&gt;Office work confers productivity spillovers and CBD amenity value that rise with the mass of in-office workers (L-tilde-c), but workers do not internalize these external benefits. So each worker prefers the office only if enough others commute. In a dynamic setting with idiosyncratic Gumbel preference shocks and fixed switching costs F, this coordination can produce multiple stationary equilibria: a high-commuting and a low-commuting one (with an unstable equilibrium E2 between them). Multiplicity requires (Prop 4.3) static agglomeration forces (delta + xi) above a threshold eta_min &amp;gt; theta + gamma/(2mu), AND relative remote productivity z/A in an intermediate interval Z — the &amp;lsquo;cone of multiplicity.&amp;rsquo; If z/A is too low, the high-commuting equilibrium is unique; if too high, only the remote equilibrium survives.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identificationquantification-strategy-and-its-main-threats"&gt;Q2. What is the identification/quantification strategy and its main threats?&lt;/h3&gt;
&lt;p&gt;To avoid taking a stand on which equilibrium generated the data, the authors rely ENTIRELY on pre-2019 data (when every city was plausibly in the high-commuting equilibrium) and on model relationships that hold in any equilibrium. Four steps: (1) transition elasticity s and cost F from NLSY79 transition probabilities via a CCP/log-linear regression (eq. 21), using past wage ratios as an instrument for future ratios to address measurement error / forward-looking expectations (IV eta0 = -0.47, eta1 = 3.09); (2) agglomeration externality delta_j from commuter-wage changes instrumented by 1980 occupational composition interacted with economy-wide occupation-specific commuter-share changes (shift-share IV, eq. 26-28), with five industry groups; (3) remote/office productivity z_j, A_j from occupation-level remote-work premia (NLSY, 22 occupation groups) reweighted by city occupation shares; (4) transport-cost elasticity gamma_j from CBSA-specific housing rent-distance gradients (ACS block-group rents 2015-2019). Main threats: selection of workers into remote work on unobservables (addressed by NLSY individual fixed effects), endogeneity of commuter shares to local productivity shocks (addressed by the shift-share IV), and the assumption that all cities were in the high-commuting equilibrium in 2019; tau_j is calibrated to match each city&amp;rsquo;s 2019 Lc/L.&lt;/p&gt;
&lt;h3 id="q3-how-do-the-authors-rule-out-competing-explanations-pure-productivitypreference-shocks-congestion-establishment-size-occupational-shift"&gt;Q3. How do the authors rule out competing explanations (pure productivity/preference shocks, congestion, establishment size, occupational shift)?&lt;/h3&gt;
&lt;p&gt;National productivity/preference shocks: would be expected to leave some lasting imprint even in small cities, but small CBSAs reverted fully, and at least 34% of jobs remain teleworkable even in fully-reverting cities (Dingel-Neiman teleworkable share ranges 25-55% across CBSAs), so low telework capacity cannot explain reversion; cities with permanent 40%+ trip declines have only a modestly higher 43% teleworkable share. The wage premium shows no differential evolution across high- vs low-teleworkable occupations over the pandemic. Congestion: if congestion drove the shift, large cities should show lower CBD propensity pre-pandemic, but the opposite holds (30.6% of trips to CBD in large vs 15.6% in small CBSAs in late 2019). Establishment concentration: employment is LESS concentrated in smaller cities, so big-employer return-to-office decisions cannot explain reversion. Occupational shift: teleworkable employment share rose only ~5% post-pandemic, and rose MORE in smaller CBSAs (7.9%) than larger (5.8%) by end-2023, the wrong direction to explain the heterogeneity.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-across-cities-is-documented-and-how-does-it-map-to-the-theory"&gt;Q4. What heterogeneity across cities is documented and how does it map to the theory?&lt;/h3&gt;
&lt;p&gt;Large cities (high agglomeration, high net strength delta - gamma/(2mu), which rises with size: doubling size raises net strength ~0.004 off a mean 0.049) are disproportionately inside the cone of multiplicity (208 of 278 estimated cities in-cone; probit on log employment positive and significant). These cities show permanent CBD-trip declines (stabilizing ~60% for the 25 largest) and persistent gradient flattening (~0.095 by 2024). Small cities are mostly outside the cone, with unique equilibria, and revert fully. The cone indicator is also positively associated with delta_j and z_j/A_j and negatively with gamma_j, as the theory predicts.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Estimates of s and F are similar using restricted-use county-geocoded NLSY and under an alternative city-partition definition (two days/week remote). Main results are robust to lower delta_j and higher gamma_j calibrations (Appendix A.17). A CES production function in remote/in-person labor yields very large substitution elasticities, motivating the linear specification. An endogenous-housing-supply model yields a nearly identical rent gradient (because commuters were a high share of employment pre-2020). Office-trip-only versions of the mobility figures (workplace visits) show similar patterns. The cone indicator retains significance in Table 2 after adding teleworkable share, pre-pandemic CBD-trip share, industry value-added shares, and total employment; results hold for an alternative binary &amp;lsquo;returned to office&amp;rsquo; indicator 1back(5,20). Multiple DYNAMIC equilibria were not found in numerical exercises (Appendix B.6).&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-differ-from-closely-related-prior-work"&gt;Q6. How does this paper differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Unlike Davis, Ghent &amp;amp; Gregory (2024) (remote productivity via adoption externalities), Parkhomenko &amp;amp; Delventhal (2024) (amenity value of remote work), and Duranton &amp;amp; Handbury (2023) (exogenous changes in who may work remotely), this paper does NOT rely on exogenous productivity or amenity/preference shocks to explain the large persistent jump. Instead a temporary commuter shock SELECTS among pre-existing multiple equilibria. Liu &amp;amp; Su (2023) document a falling urban wage premium for remote-amenable occupations (consistent with weaker agglomeration). The paper&amp;rsquo;s documented divergence of residential rent-distance gradients between large and small cities is, to the authors&amp;rsquo; knowledge, a new fact, interpreted structurally. Owens, Rossi-Hansberg &amp;amp; Sarte (2020) similarly use coordination/residential externalities (Detroit neighborhoods).&lt;/p&gt;
&lt;h3 id="q7-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q7. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Because the coordination failure operates partly OUTSIDE firm boundaries, individual firms&amp;rsquo; return-to-office mandates may be insufficient to restore the high-commuting equilibrium. City-level interventions — taxing remote work or subsidizing commuting — could in principle move a city back, since the only active externality in the quantification is a positive agglomeration externality (implying too little commuting relative to the efficient benchmark in all equilibria). However, the authors stress these welfare effects and the effectiveness of policy remain open questions; their welfare numbers depend on estimation details and the abstraction from a system-of-cities with migration.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-main-caveats-and-abstractions"&gt;Q8. What are the main caveats and abstractions?&lt;/h3&gt;
&lt;p&gt;The model treats each city as a CLOSED economy: no inter-city migration, trade, or investment links, though the authors note large cities show a small differential population drop (Appendix A.9), attributed to low migration elasticities. Remote work is &amp;lsquo;partial&amp;rsquo; with a FIXED fraction mu = 3/5 of days at home, not chosen. Occupational heterogeneity is abstracted from (justified by rare occupation transitions). The amenity (xi) vs congestion (theta) externalities are not separately identified and set to zero (difference insignificant). Spillovers are not internalized by firms in the model. The welfare ranking (high-commuting preferred) is intuited from the single positive externality rather than formally proven.&lt;/p&gt;
&lt;h3 id="q9-why-is-there-a-discrepancy-between-the-abstracts-welfare-figures-and-per-city-numbers"&gt;Q9. Why is there a discrepancy between the abstract&amp;rsquo;s welfare figures and per-city numbers?&lt;/h3&gt;
&lt;p&gt;The abstract and revised Table 3 report a mean welfare loss of 2.3% (median 2.2%, range 1.2%-4.0%) across the 208 cone cities, and state cities with permanently low commuting (60% or less of pre-pandemic trips) experience average losses of 2.3% (2.7% in the text). The introduction additionally quotes specific city losses (about 3.7% for Los Angeles and San Jose, 3.2% for New York, 2.8% for San Francisco, 2% for Phoenix); these are the largest cities and lie within or near the upper part of the distribution, consistent with welfare loss rising in net agglomeration strength (R-squared 0.85 of loss on delta - gamma/(2mu)).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;!-- flags: Welfare magnitudes: the final/revised headline figures are mean 2.3%, median 2.2%, range 1.2-4.0% (Table 3, 208 cities). The Introduction also cites larger per-city losses (3.7% LA/San Jose, 3.2% NYC, 2.8% SF, 2% Phoenix); these are consistent with the distribution (loss rises with net agglomeration strength) but appear to be from a specific large-city calibration table, not the summary distribution. Reported both, flagged for reviewer., Paper is a Nov 2025 revision of NBER WP 31494 (orig. July 2023); some figures span data through end-2024/Nov-2024, later than the original draft. --&gt;</description></item><item><title>Resource Misallocation in European Firms: The Role of Constraints, Firm Characteristics and Managerial Decisions</title><link>https://macropaperwarehouse.com/papers/resource-misallocation-in-european-firms-the-role-of-constraints-firm-characteristics-and-managerial-decisions/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/resource-misallocation-in-european-firms-the-role-of-constraints-firm-characteristics-and-managerial-decisions/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates why firms in the European Union exhibit wide dispersion in marginal revenue products (MRP) of capital and labor — a direct indicator of resource misallocation — and asks how much aggregate productivity the EU forfeits as a result. The research question is motivated by the persistent productivity gap between the EU and the United States, by evidence that within-country MRP dispersion in Europe has been trending upward since the mid-1990s, and by an institutional context in which the EU single market (launched in 1993) has not eliminated cross-country factor market frictions even three decades later.&lt;/p&gt;
&lt;p&gt;The primary data source is the EIB Investment Survey (EIBIS), a stratified random survey of non-financial enterprises conducted annually since 2016 across all 28 EU member states, covering manufacturing, services, utilities, and construction (NACE categories C–J). The analysis uses three waves (2016–2018), with approximately 12,500 firms per wave and a panel component of roughly 2,000 firms appearing in all three waves. Survey responses are matched to Orbis administrative data; the correlation between log employment in EIBIS and Orbis is 0.91, confirming data quality. MRP of capital (MRPK) is measured as the capital cost share times revenue divided by fixed assets; MRP of labor (MRPL) is the labor cost share times revenue divided by employment. Cost shares are calibrated from OECD STAN and Eurostat national accounts at the country–year–industry level.&lt;/p&gt;
&lt;p&gt;The theoretical framework is a dynamic model of a profit-maximizing firm with Cobb-Douglas production, isoelastic demand, and quadratic adjustment costs. Under the assumption that pure economic profits are small and that the labor output distortion is negligible (following Hsieh-Klenow 2009), the model implies that log MRPK and log MRPL can be approximated by observable average revenue products. The empirical strategy is a Mincerian regression of log MRPK (and log MRPL) on a rich vector of firm-level characteristics — firm demographics, input quality, capacity utilization, investment constraints, dynamic adjustment variables, and financing sources — plus country, industry, and year fixed effects (and their interactions). Because regressors are endogenous, the R² from OLS is interpreted as an upper bound on the share of MRP variance attributable to each factor (formally shown to dominate the IV R²). Marginal R² increments when a variable block is added identify the contribution of that block to the variance in MRP, which is then mapped into productivity gains via the Hsieh-Klenow formula.&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. Raw dispersion is large: the standard deviation of log MRPK is 1.43 and of log MRPL is 1.19 (and 1.63 for log MRPL minus log MRPK), all substantially exceeding comparable US figures (0.98 for capital and 0.58 for labor from Asker et al. 2014 and Bartelsman et al. 2013). The R² in the full regression is 0.14 (without fixed effects) and 0.49 (with country × industry × year fixed effects) for MRPK, and 0.29 and 0.74 respectively for MRPL. Among firm-characteristic blocks, the &amp;ldquo;adjustment&amp;rdquo; (dynamic investment and employment growth) and &amp;ldquo;demographics&amp;rdquo; (firm size, age, subsidiary and exporter status) blocks carry the largest marginal R² contributions; the &amp;ldquo;obstacles to investment&amp;rdquo; block (direct reports of constraints) contributes modestly by comparison. Country fixed effects alone explain R² = 0.052 for MRPK and R² = 0.445 for MRPL, while industry fixed effects alone explain R² = 0.239 for MRPK and R² = 0.268 for MRPL. The combined country–industry–year fixed-effects R² reaches 0.275 for MRPK and 0.611 for MRPL; adding the full interaction yields 0.492 and 0.736 respectively.&lt;/p&gt;
&lt;p&gt;Treating the &amp;ldquo;distortions&amp;rdquo; block of variables as genuine frictions, removing them would raise EU aggregate productivity by more than 40 percent (computed as 1.5 × 1.42 × 0.186 + 0.13 × 2.66 × 0.134 = 0.442). If all variables in X are treated as distortions, the implied gain is approximately 72 percent (0.715 in log points). Removing cross-country inequality in average MRPs (equalizing country fixed effects) would imply a 102 percentage log-point gain in productivity under the Hsieh-Klenow formula; removing barriers between industries and countries could raise productivity by at least 143 percentage log points.&lt;/p&gt;
&lt;p&gt;A Machado-Mata distributional decomposition comparing Germany (σ(log MRPK) = 0.92, σ(log MRPL) = 0.61) and Greece (σ(log MRPK) = 1.64, σ(log MRPL) = 0.91) reveals that the primary driver of Greece&amp;rsquo;s higher dispersion is the &amp;ldquo;prices&amp;rdquo; (regression coefficients reflecting institutional and policy environment), not the &amp;ldquo;endowments&amp;rdquo; (firm characteristics). Giving Greece German institutional &amp;ldquo;prices&amp;rdquo; reduces the counterfactual standard deviation of Greek MRPK from 1.66 to 0.94. This pattern generalizes across EU countries: German b (coefficients) tends to reduce MRPK dispersion for most countries, while German X (firm characteristics) tends to increase it, because Germany has more heterogeneous firms but an environment that prices those characteristics in a way that equalizes returns. This finding constitutes large-scale microeconomic evidence that institutions matter — cross-country differences in MRP dispersion reflect how business, institutional, and policy environments translate firm heterogeneity into outcomes, more than they reflect differences in firm characteristics per se.&lt;/p&gt;
&lt;p&gt;The policy implication is that deep institutional reform — not merely changes in firm composition — is required to narrow EU resource misallocation. The scope condition is that these estimates are upper bounds, and some observed MRP dispersion likely reflects compensating differentials (e.g., higher-quality capital commanding a higher MRPK) rather than pure distortions.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper does not attempt causal identification. Instead, it uses OLS to estimate equilibrium (Mincerian-type) regressions of log MRPK and log MRPL on firm characteristics plus fixed effects. The key insight is that OLS R² provides an upper bound on the share of MRP variance causally attributable to each regressor, because simultaneity or omitted variables can only inflate OLS R² above the true IV R². The main threats are: (1) endogeneity of regressors — a growing firm facing red tape will have high MRPK and a binding constraint simultaneously, inflating the R² attributed to constraints; (2) classical measurement error in survey responses, which attenuates R² toward zero (so OLS actually understates causal effects in this direction); (3) omitted variable bias via unobserved firm quality (managerial talent, etc.); (4) use of same variables (employment, fixed assets) on both left and right sides, addressed by cross-checking with Orbis data as instruments. The authors argue these threats are mostly conservative — they overstate, not understate, the upper bound.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-theoretical-justification-for-using-average-revenue-products-to-measure-marginal-revenue-products"&gt;Q2. What is the theoretical justification for using average revenue products to measure marginal revenue products?&lt;/h3&gt;
&lt;p&gt;Under the assumption that the share of pure economic profits is small (following Basu and Fernald 1997), the optimality conditions of the dynamic model imply that MRPK ≈ (capital cost share) × (revenue / capital) and MRPL ≈ (labor cost share) × (revenue / employment). These are average revenue products scaled by factor cost shares, matching Hsieh and Klenow (2009). The distortion framework further implies that the variance of log MRPK and log MRPL, when distortions are log-normally distributed and uncorrelated, maps directly into the Hsieh-Klenow productivity-loss formula, linking the regression R² to quantitative welfare calculations.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-role-of-compensating-differentials-versus-true-distortions-in-interpreting-the-results"&gt;Q3. What is the role of compensating differentials versus true distortions in interpreting the results?&lt;/h3&gt;
&lt;p&gt;The paper emphasizes that not all dispersion in MRPs reflects inefficient distortions. Some dispersion — particularly from &amp;lsquo;quality of capital,&amp;rsquo; &amp;lsquo;capacity utilization,&amp;rsquo; and &amp;lsquo;dynamic adjustment&amp;rsquo; — may reflect compensating differentials: firms that invest in higher-quality capital rationally face higher costs, demanding a higher MRPK in equilibrium, analogous to how more educated workers earn higher wages in a Mincerian framework. If these variables reflect compensating differentials rather than frictions, using &amp;lsquo;raw&amp;rsquo; MRP dispersion overstates misallocation. Conversely, if all variables proxy for distortions, the productivity gains from reform are even larger (72 percent versus 40 percent). The paper presents both interpretations explicitly, making the framework &amp;lsquo;highly portable&amp;rsquo; for different views of what drives observed dispersion.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-in-mrp-dispersion-is-documented-across-eu-countries-and-industries"&gt;Q4. What heterogeneity in MRP dispersion is documented across EU countries and industries?&lt;/h3&gt;
&lt;p&gt;Dispersion is notably lower in Germany (σ(log MRPK) = 0.92, σ(log MRPL) = 0.61) than in Greece (1.64 and 0.91) or smaller countries such as Malta, Luxembourg, and Cyprus. Country fixed effects explain R² = 0.445 of MRPL variation but only R² = 0.052 of MRPK variation, meaning labor is more segmented across countries than capital. Industry fixed effects explain R² = 0.239 for MRPK versus R² = 0.268 for MRPL, indicating capital is more segmented across industries than across countries. Core EU countries (France, Denmark) are relatively insensitive to counterfactual substitution of German coefficients, while periphery countries (Portugal, Ireland) show large movements. Romania, which resembles Slovenia in raw MRPK dispersion, looks much more like the Netherlands after controlling for firm characteristics — illustrating that observed dispersion rankings can be misleading without adjustment.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-machado-mata-decomposition-reveal-and-how-is-it-implemented"&gt;Q5. What does the Machado-Mata decomposition reveal, and how is it implemented?&lt;/h3&gt;
&lt;p&gt;The Machado-Mata (2005) decomposition separates the distribution of MRP into an &amp;rsquo;endowments&amp;rsquo; component (due to the values of firm characteristics X) and a &amp;lsquo;prices&amp;rsquo; component (due to the regression coefficients b, which capture how the institutional and policy environment translates X into outcomes). The decomposition draws B = 10,000 bootstrap samples from the empirical distribution of X for each country, combines them with quantile regression coefficients estimated separately for each country, and constructs counterfactual distributions. Applying Greek X with German b reduces Greece&amp;rsquo;s counterfactual σ(log MRPK) from 1.66 to 0.94 — close to Germany&amp;rsquo;s actual 0.92 — while applying German X with Greek b increases dispersion. The main finding is that differences in &amp;lsquo;prices&amp;rsquo; (institutional environment) dominate differences in &amp;rsquo;endowments&amp;rsquo; (firm characteristics) in explaining cross-country variation in within-country MRP dispersion. This pattern holds generally across EU countries: gains from &amp;lsquo;importing&amp;rsquo; German institutions are correlated with poor World Bank Governance Indicators and International Country Risk Guide scores.&lt;/p&gt;
&lt;h3 id="q6-how-do-the-papers-estimates-of-eu-misallocation-compare-to-us-benchmarks"&gt;Q6. How do the paper&amp;rsquo;s estimates of EU misallocation compare to US benchmarks?&lt;/h3&gt;
&lt;p&gt;The EU standard deviations of log MRPK (1.43) and log MRPL (1.19) substantially exceed comparable US figures of 0.98 for capital (Asker et al. 2014) and 0.58 for labor (Bartelsman et al. 2013). The paper discusses three caveats for this comparison: (1) EIBIS uses revenue rather than value added, which affects dispersion (approximately +0.16 log points for MRPL, -0.21 for MRPK) — insufficient to explain the full gap; (2) survey measurement error is present but small — averaging over multiple waves reduces the standard deviation of log MRPK by only 8–12 percent; (3) EIBIS measures firms (not plants), and since about two-thirds of within-firm MRPK variance occurs across plants within firms (Kehrig and Vincent 2017), the EU–US comparison likely understates the true difference. Qualitatively, the greater EU dispersion is consistent with lower EU aggregate TFP relative to the US.&lt;/p&gt;
&lt;h3 id="q7-what-specific-regression-results-are-reported-for-individual-variable-blocks"&gt;Q7. What specific regression results are reported for individual variable blocks?&lt;/h3&gt;
&lt;p&gt;The full R² (without / with country × industry × year fixed effects) is 0.14 / 0.49 for MRPK and 0.29 / 0.74 for MRPL. Among variable blocks, the &amp;lsquo;adjustment&amp;rsquo; (investment, employment growth, past and planned investment) and &amp;lsquo;demographics&amp;rsquo; (size, age, subsidiary, exporter) blocks have the largest marginal R². The &amp;lsquo;obstacles to investment&amp;rsquo; (direct constraint reports) block contributes modestly, with some coefficients not statistically significant. Within regression coefficients (from Table A.4): older, exporting, high-utilization firms have higher MRPK and MRPL; investment is strongly negatively associated with MRPK (movement down the MRPK curve as capital rises) and positively with MRPL (labor becomes relatively scarcer); employment growth is positively associated with MRPK and negatively with MRPL (symmetric logic); credit-constrained status is negatively correlated with both MRPK and MRPL.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-are-run"&gt;Q8. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper reports: (1) &amp;lsquo;between&amp;rsquo; regressions on multi-year firm averages to reduce transitory variation and measurement error — results are qualitatively similar with slightly larger productivity gains; (2) restricting the sample to firms appearing in all three survey waves (Appendix Table A.5) — qualitatively similar results; (3) estimating equation (4) for each wave separately — similar results; (4) using Orbis employment and investment as regressors instead of EIBIS responses to address mechanical measurement-error correlation — nearly identical results (Appendix Table A.17); (5) replacing log(1+investment) with an indicator for positive investment (Appendix Table A.7) — similar results; (6) using industry-specific rather than country–year–industry cost shares — similar results; (7) confirming that measurement error can account for only a portion of the EU–US dispersion difference (8–12 percent reduction in standard deviation when averaging over waves). The paper also reports separate coefficient estimates for three blocs of EU countries (North/West, South, Center/East) in Appendix Tables A.10–A.16.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-relate-to-and-differ-from-hsieh-and-klenow-2009-and-related-prior-work"&gt;Q9. How does the paper relate to and differ from Hsieh and Klenow (2009) and related prior work?&lt;/h3&gt;
&lt;p&gt;The paper extends Hsieh and Klenow (2009) in several directions. First, while Hsieh-Klenow use administrative census-type data for India and China restricted to manufacturing, this paper uses a consistent cross-country survey covering all sectors in 28 EU countries, enabling direct cross-country comparison. Second, Hsieh-Klenow implicitly assume all MRP dispersion reflects distortions; this paper explicitly distinguishes distortions from compensating differentials and shows the distinction matters quantitatively. Third, this paper develops the Mincerian regression approach to apportion the variance in MRPs across observable factors — analogous to labor economists decomposing wage dispersion — and shows OLS R² provides a valid upper bound without requiring exogenous variation. Fourth, unlike country-level distortion measures (Gamberoni et al. 2016), tight theoretical restrictions (David and Venkateswaran 2017), or specific reforms (Rotemberg 2019), this paper draws on firm-level survey data with minimal restrictions and maintains high external validity. Fifth, the Machado-Mata distributional decomposition adds a new dimension absent from Hsieh-Klenow: decomposing cross-country differences into endowments vs. institutional &amp;lsquo;prices.&amp;rsquo;&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary policy implication is that EU productivity could rise by more than 40 percent if distortions to resource allocation were removed — and up to 72 percent if all observed MRP variation is attributed to distortions. A more modest goal of equalizing within-industry MRP dispersion across countries (i.e., making Germany and Greece similar within industries) implies gains of approximately 31–53 percent depending on interpretation. The decomposition evidence implies that institutional reform (changing how environments price firm characteristics) is more important than directly changing firm composition. The scope conditions are: (1) these are upper bounds derived from OLS; (2) some dispersion reflects compensating differentials that should not be counted as losses; (3) the EIBIS covers firms with at least 5 employees, so very small firms are excluded; (4) the framework assumes log-normal, uncorrelated distortions and constant returns to scale — relaxing these can increase estimated losses further (Jones 2011); (5) the estimates do not account for firm-level markup heterogeneity, which could overstate or understate other channels.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-paper-contribute-to-the-literature-on-measurement-error-in-mrp-studies"&gt;Q11. What does the paper contribute to the literature on measurement error in MRP studies?&lt;/h3&gt;
&lt;p&gt;The paper shows formally (Appendix D) that classical measurement error in regressors attenuates OLS R² toward zero, so OLS provides a conservative upper bound from this direction. It also shows that averaging across multiple survey waves reduces measurement error while also attenuating transitory adjustment-cost variation, so multi-year averages likely overstate the role of measurement error. Crucially, the paper validates EIBIS against Orbis administrative data, finding a 0.91 correlation for log employment, similar standard deviations of log MRPK (1.44 in Orbis vs. 1.37 in EIBIS) and log MRPL (1.07 in Orbis vs. 1.30 in EIBIS) for matched firms, and a mean absolute log difference in standard deviations of approximately 2 percent across countries. This contributes to the debate initiated by Bils et al. (2017) on whether measured MRP dispersion reflects mismeasurement, and corroborates that surveys can be reliable substitutes for census-type administrative data in cross-country analysis.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-paper-find-about-the-role-of-credit-constraints-specifically"&gt;Q12. What does the paper find about the role of credit constraints specifically?&lt;/h3&gt;
&lt;p&gt;Credit constraint status (defined as loan rejection, discouragement from applying, or receiving a loan that was too small or too expensive) is negatively correlated with both MRPK and MRPL in the full regression. This is consistent with credit-constrained firms being unable to invest to the point where MRPK is equalized with the cost of capital, but the negative sign also raises the interpretive caveat noted by the authors: cross-sectional equilibrium relationships can have signs inconsistent with causal priors because constraints may be more binding for firms that are already performing poorly. The &amp;lsquo;source of funds&amp;rsquo; block (share of investment from internal vs. external sources, and credit constraint) is grouped with &amp;lsquo;distortions&amp;rsquo; in the paper&amp;rsquo;s preferred decomposition.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Marginal Revenue Product (MRPK/MRPL)&lt;/strong&gt;: In this paper, the marginal revenue product of capital (MRPK) and labor (MRPL) are measured as observable average revenue products — the capital or labor cost share times revenue divided by the stock of capital or employment. Under the paper&amp;rsquo;s model assumptions, these approximate the shadow cost of inputs and serve as the primary measure of firm-level resource allocation efficiency. A firm with a high MRPK relative to its cost of capital is under-capitalized; dispersion of MRPK across firms signals misallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Compensating differentials (in the MRP context)&lt;/strong&gt;: The paper adapts the Mincerian concept of compensating differentials from labor markets to the firm side: some observed dispersion in MRPK and MRPL may reflect optimal responses to heterogeneity in input quality, capital utilization, or adjustment dynamics — not inefficient distortions. For example, a firm with state-of-the-art machinery may face a higher MRPK reflecting the quality premium, not a barrier to investment. Because such dispersion is rational, it should be subtracted from productivity-loss calculations rather than counted as welfare-reducing misallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Machado-Mata decomposition&lt;/strong&gt;: A distributional decomposition technique (Machado and Mata 2005) applied here to attribute cross-country differences in the dispersion of MRPK and MRPL to two components: &amp;rsquo;endowments&amp;rsquo; (the empirical distribution of firm characteristics X in a given country) and &amp;lsquo;prices&amp;rsquo; (the regression coefficients b, which capture how the country&amp;rsquo;s business, institutional, and policy environment translates those characteristics into marginal revenue products). The decomposition constructs counterfactual MRP distributions by combining one country&amp;rsquo;s X with another country&amp;rsquo;s b.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mincerian productivity regression&lt;/strong&gt;: The paper&amp;rsquo;s core empirical framework, modeled explicitly on Mincer&amp;rsquo;s (1958) wage regression: just as wages are regressed on worker characteristics (education, experience) to decompose earnings dispersion, log MRPK and log MRPL are regressed on firm characteristics (demographics, quality, utilization, adjustment, constraints, financing) to decompose MRP dispersion. OLS R² in this regression is an upper bound on the share of MRP variance attributable to each regressor.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;EIB Investment Survey (EIBIS)&lt;/strong&gt;: An annual firm-level survey administered by Ipsos MORI on behalf of the European Investment Bank since 2016, covering all 28 EU member states with a stratified random sample of approximately 12,500 non-financial enterprises per wave (minimum 5 employees, NACE C–J). Unique features include consistent cross-country design, merger with Orbis administrative data, and questions on investment plans, capital quality, capacity utilization, perceived obstacles, and financing sources — all directly informative about sources of MRP variation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Institutional &amp;lsquo;prices&amp;rsquo; on firm characteristics&lt;/strong&gt;: In the Machado-Mata framework as applied here, &amp;lsquo;prices&amp;rsquo; refer to the country-specific regression coefficients b in the MRP regression — how steeply a country&amp;rsquo;s environment (regulations, institutions, policies) translates a given unit of firm heterogeneity in X into a difference in marginal revenue products. Countries with smaller b magnitudes (like Germany) achieve more equalization of MRPs across heterogeneous firms, reflecting an efficient institutional environment; countries with large b (like Greece) amplify firm-level heterogeneity into large MRP dispersion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Upper-bound R² approach to productivity gains&lt;/strong&gt;: The paper&amp;rsquo;s portable method for quantifying productivity gains from removing a friction: the marginal R² increment in an OLS regression of log MRPK (or log MRPL) when a friction variable is added is an upper bound on the share of MRP variance attributable to that friction. This bound, multiplied by the variance of log MRP and the Hsieh-Klenow productivity-loss formula parameters, gives an upper-bound estimate of the aggregate TFP gain from eliminating that friction. The method does not require exogenous variation or tight structural assumptions.&lt;/p&gt;</description></item><item><title>Selection, Structural Transformation, and the Cost Disease of Services</title><link>https://macropaperwarehouse.com/papers/selection-structural-transformation-and-the-cost-disease-of-services/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/selection-structural-transformation-and-the-cost-disease-of-services/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether worker self-selection, rather than slow technological progress, can explain the low measured labor productivity growth in the U.S. service sector — a phenomenon known as Baumol&amp;rsquo;s cost disease. The conventional view, associated with Young (2014), is that as workers reallocate from manufacturing into services, the incoming workers are less skilled than incumbents, mechanically depressing measured productivity; on that view, the cost disease might be a transient mismeasurement artifact rather than a permanent technological fact. Shu challenges this interpretation by showing that the selection pattern differs sharply across service sub-sectors and is far weaker in aggregate than the conventional model predicts.&lt;/p&gt;
&lt;p&gt;The empirical foundation is the Outgoing Rotation Group of the U.S. Current Population Survey (1989–2020), linked longitudinally to track workers who switch sectors between consecutive years. The sample contains 1,406,674 matched worker-year observations. Cross-country patterns from the GGDC 10-Sector Database (nine developed countries, 1989–2009) provide motivating evidence: over that period, labor productivity grew 56 log points in manufacturing, 77 log points in professional services (finance, real estate, professional and business services), and only 5 log points in EHP (education, health, and public administration). The cross-country correlation between employment-share growth and labor productivity growth is +0.48 for professional services — the opposite sign from the conventional selection story — and −0.14 for EHP, which conforms to it.&lt;/p&gt;
&lt;p&gt;At the micro level, a regression of log real weekly earnings on previous-sector dummies (with year and county fixed effects, standard errors clustered by county) yields a key asymmetry: workers who move from manufacturing into professional services earn 4.8 log points (approximately 4.9%) more than incumbent professional services workers (coefficient 0.048, se 0.010), while workers who move from EHP into professional services earn 14.3 log points less (coefficient −0.143, se 0.008). Workers switching from manufacturing into EHP earn 8.7 log points less than EHP incumbents (coefficient −0.087, se 0.023). The first fact — that incoming workers from manufacturing outperform incumbents in professional services — cannot be generated by conventional Roy models based on independent Fréchet skill distributions, which force skill levels in an expanding sector to fall.&lt;/p&gt;
&lt;p&gt;To accommodate these patterns, Shu builds a three-sector general-equilibrium Roy model with a non-homothetic CES demand structure (following Comin, Lashkari and Mestieri 2021). The skill distribution is parameterized by allowing absolute advantage in professional services to depend on comparative advantages in manufacturing (parameter αm) and EHP (αe), conditional on the comparative advantage quantiles following a Gumbel distribution. The model is estimated via simulated method of moments, targeting the three observed earnings premia and the variance of log income. The estimated parameters confirm αm = 0.055 &amp;gt; 0 (workers with higher comparative advantage in manufacturing also have higher absolute productivity in professional services) and αe = −0.123 &amp;lt; 0 (workers with higher comparative advantage in EHP are less productive in professional services).&lt;/p&gt;
&lt;p&gt;The main quantitative results for the full 1990–2020 sample are: selection raises labor productivity in professional services by 1.2 log points and lowers it in EHP by 0.7 log points, for a net effect of zero on aggregate services. By contrast, the conventional independent Fréchet model predicts selection effects of −8.7 log points for professional services and −3.0 log points for EHP, summing to −5.2 log points for aggregate services. The discrepancy for professional services alone is 9.9 log points — a difference of more than seven-fold in magnitude and opposite in sign. Consequently, the conventional model overpredicts true technology growth in professional services by over one-third relative to the baseline. The implied true technology growth rates over 1990–2020 are 88.1 log points for manufacturing, 27.3 for professional services, and −0.6 for EHP, leaving a large and unexplained productivity gap between manufacturing and services that selection cannot close. This directly refutes Young&amp;rsquo;s (2014) claim that selection accounts for virtually all of the measured gap, and confirms that Baumol&amp;rsquo;s cost disease reflects genuinely low technology growth in EHP and moderately lower growth in professional services.&lt;/p&gt;
&lt;p&gt;A forward-looking simulation extending the implied technology growth rates (2.9% p.a. for manufacturing, 0.9% for professional services, 0% for EHP) over fifty years produces similar welfare gains under both specifications (29.4 vs. 29.2 log points), but through very different mechanisms: the conventional model reaches its welfare estimate through counterfactually large selection effects in both directions that cancel, while the baseline model generates more modest and empirically grounded reallocation dynamics.&lt;/p&gt;
&lt;p&gt;The unexplained portion of the manufacturing-to-professional-services earnings premium is explored through an extensive set of micro-regressions controlling for education, experience, hours, occupation, age, race, and gender. Gender composition is the single most important observable channel: workers switching from manufacturing into professional services are 17.7 percentage points more male than the incumbent professional services workforce, and male workers earn roughly 40% more, implying a composition-driven premium of about 7.1 log points. Even after controlling for all observables, approximately one-quarter of the 4.8 log-point premium remains unexplained. Among college-educated female workers, the unexplained manufacturing premium is 4.5 log points — as large as the unconditional estimate — which Shu flags for future investigation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-micro-level-selection-patterns-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy for the micro-level selection patterns, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses the longitudinal structure of the CPS Outgoing Rotation Group to observe the same worker in two consecutive years and identify their origin sector and destination sector. The income gap between incoming workers and incumbents in the same sector-year cell (conditional on year and county fixed effects, with county-clustered standard errors) provides the key moments. The main threats are: (1) workers may self-select into switching for unobserved reasons correlated with productivity (e.g., those with better outside options move), but the direction of such bias is ambiguous; (2) the paper explicitly focuses on direct sector-to-sector transitions to isolate long-run structural reallocation from short-run labor supply fluctuations — a design choice distinguishing it from Young (2014), who used aggregate defense spending as an IV but thereby conflated unemployment and non-participation dynamics with genuine sector reallocation. The paper does not employ a separate instrument for the selection into switching; instead, it uses the income-gap moments as identified empirical objects to discipline the structural model.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-differ-from-young-2014-and-why-does-it-reach-opposite-conclusions"&gt;Q2. How does the paper differ from Young (2014), and why does it reach opposite conclusions?&lt;/h3&gt;
&lt;p&gt;Young (2014) uses industry-level employment and output data and estimates a uniform, negative elasticity of &amp;lsquo;worker efficacy&amp;rsquo; with respect to employment share across all industries, concluding that selection explains away essentially all of the manufacturing–services productivity gap. Three key differences drive Shu&amp;rsquo;s opposite conclusion. First, Shu uses worker-level panel data that allow distinct selection patterns to be estimated separately for professional services versus EHP, rather than imposing a common pattern. Second, Shu documents that the conventional pattern (incoming workers earn less than incumbents) holds for EHP but fails for professional services, where workers from manufacturing earn about 4.9% more than incumbents — a fact Young&amp;rsquo;s approach cannot detect. Third, Young&amp;rsquo;s IV (defense spending-to-GDP ratio) is used for demand shocks on aggregate employment, which mixes short-run unemployment and non-participation adjustments with the long-run structural reallocation that is relevant for selection; Shu&amp;rsquo;s design isolates workers who transition directly between sectors and thus captures only the long-run phenomenon.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-role-of-the-relationship-between-absolute-and-comparative-advantages-in-the-model-and-how-does-the-paper-generalize-prior-work"&gt;Q3. What is the role of the relationship between absolute and comparative advantages in the model, and how does the paper generalize prior work?&lt;/h3&gt;
&lt;p&gt;Standard Roy models (including those using independent Fréchet distributions as in Lagakos and Waugh 2013, Bryan and Morten 2019, and Hsieh et al. 2019) implicitly assume that workers&amp;rsquo; absolute advantage in a sector increases with their comparative advantage in the same sector. This restriction forces labor productivity of any expanding sector to fall. Adão (2016) and Alvarez-Cuadrado, Amodio and Poschke (2019) made the theoretical point that the sign of αm (the correlation between comparative advantage in manufacturing and absolute advantage in professional services) is the key determinant of whether selection helps or hurts professional services productivity. Shu&amp;rsquo;s paper generalizes Adão&amp;rsquo;s two-sector log-linear framework to three sectors, introduces the explicit parameterization via the Gumbel conditional distribution, and crucially provides a parametric method to quantify the contribution of selection to measured labor productivity by estimating αm and αe from worker-level moments. The estimated αm = 0.055 &amp;gt; 0 is what generates the positive selection effect for professional services.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-calibrated-technology-growth-rates-implied-by-the-model-and-what-do-they-imply-for-baumols-cost-disease"&gt;Q4. What are the calibrated technology growth rates implied by the model and what do they imply for Baumol&amp;rsquo;s cost disease?&lt;/h3&gt;
&lt;p&gt;Over 1990–2020, the calibrated model implies cumulative technology growth of 88.1 log points in manufacturing, 27.3 log points in professional services, and −0.6 log points in EHP. These numbers confirm that technology growth in EHP has been essentially zero over three decades, and that professional services, despite having high measured labor productivity growth, has grown at roughly one-third the rate of manufacturing in true technology terms. The 93.5 log-point difference in measured output per worker between manufacturing and aggregate services is broken down as: 15.6 log points attributable to the selection effect on manufacturing (outgoing workers are below-average) and essentially zero attributable to selection in aggregate services, leaving a true technology gap of approximately 77.9 log points. The conclusion is that the cost disease — specifically the stagnation of EHP — is a real technological phenomenon, not a mismeasurement artifact.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-conventional-independent-fréchet-model-compare-quantitatively-to-the-baseline-and-where-do-the-specifications-diverge-most"&gt;Q5. How does the conventional independent Fréchet model compare quantitatively to the baseline, and where do the specifications diverge most?&lt;/h3&gt;
&lt;p&gt;The comparison is presented in Table 7. For professional services, the baseline finds a selection effect of +1.2 log points while the conventional model finds −8.7 log points — a difference of 9.9 log points, more than seven-fold in magnitude and reversed in sign. For EHP the baseline finds −0.7 versus −3.0 under the conventional model. For aggregate services the baseline finds 0.0 versus −5.2 for the conventional model. In the implied technology growth, the conventional model overpredicts professional services technology growth by over one-third relative to the baseline (37.3 versus 27.3 log points), and for aggregate services overpredicts by more than 50% (15.4 versus 10.2 log points). In the 50-year forward projection, both models produce nearly identical welfare changes (29.4 vs. 29.2 log points) but through opposite and partially offsetting selection effects in manufacturing versus services under the Fréchet model — a result Shu flags as an artifact of the conventional model&amp;rsquo;s internally inconsistent mechanism.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-in-selection-patterns-is-documented-at-the-micro-level"&gt;Q6. What heterogeneity in selection patterns is documented at the micro level?&lt;/h3&gt;
&lt;p&gt;Three dimensions of heterogeneity are documented. First, the direction of selection differs by sub-sector: incoming manufacturing workers earn more than incumbents in professional services (+4.9%) but less than incumbents in EHP (−8.7%). Second, the role of observables differs: in professional services, none of the standard controls (education, experience, hours, occupation, age, race) eliminate the manufacturing premium, while gender composition accounts for roughly three-quarters of it. In EHP, the same set of controls explains the income gaps well, consistent with conventional selection. Third, the premium within professional services is concentrated among college graduates: among workers with college degrees, the manufacturing premium is 2.7%; among those without degrees, it is statistically indistinguishable from zero. College-educated female workers from manufacturing show a particularly strong premium of 4.5 log points, larger than most subgroups. Male workers switching from manufacturing constitute over 60% of the inflow for most of the sample, compared to roughly 50% male share among incumbents (the male share of incumbents rises over time as the inflow changes the composition).&lt;/p&gt;
&lt;h3 id="q7-what-role-does-gender-play-in-explaining-the-manufacturing-earnings-premium-in-professional-services"&gt;Q7. What role does gender play in explaining the manufacturing earnings premium in professional services?&lt;/h3&gt;
&lt;p&gt;Gender is the quantitatively dominant observable channel. Workers reallocating from manufacturing into professional services are on average 17.7 percentage points more male than the incumbent professional services workforce. Male workers earn roughly 40% (log 0.407) more than female workers within professional services. A back-of-envelope calculation: a 17.7 percentage-point male-share gap times a 40% earnings premium implies a composition-driven premium of approximately 7.1 log points, which matches the difference between the unconditional coefficient (0.048) and the gender-conditioned coefficient (−0.022). Adding the gender dummy to the regression turns the manufacturing premium negative and marginally significant (−0.022, Table 10 column 4), confirming that the premium is largely a composition effect. However, Table 11&amp;rsquo;s full specification (including all observable controls) still leaves a positive residual of 1.2 log points (statistically significant), suggesting approximately one-quarter of the original 4.8 log-point premium is genuinely unexplained. The paper identifies non-pecuniary sorting preferences (Goldin 2014; Faberman, Mueller and Şahin 2025) and sector-specific human capital as candidate explanations for future research.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-are-run-and-what-is-the-sensitivity-of-results-to-parameter-choices"&gt;Q8. What robustness checks are run, and what is the sensitivity of results to parameter choices?&lt;/h3&gt;
&lt;p&gt;The paper compares the baseline model to an independent Fréchet specification (shape parameter 2.7, consistent with Bryan and Morten 2019 and Lagakos and Waugh 2013) as the main alternative parameterization. It notes in a footnote that lower shape parameters (Hsieh et al.&amp;rsquo;s ~2, or Young&amp;rsquo;s implied ~1.33) would produce even stronger negative selection effects, making the Fréchet comparison conservative. At the micro level, the earnings regressions are extended through five successive specifications in Tables 9, 10, 11, and 12, each adding further controls, to verify the robustness of the manufacturing premium in professional services. The premium survives across all specifications for workers with college degrees. The paper also notes that its selection effect is identified entirely from worker-level income data and does not depend on the measured numbers of labor productivity, so measurement errors in sectoral output data (discussed in Triplett and Bosworth 2004) do not contaminate the core finding. The paper excludes workers under 25 to ensure the sector choices are long-run-oriented rather than early-career experiments.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-future-projection-exercise-show-and-what-are-its-scope-conditions"&gt;Q9. What does the future projection exercise show, and what are its scope conditions?&lt;/h3&gt;
&lt;p&gt;The exercise projects structural transformation over 2020–2070 by feeding the sample-period-implied technology growth rates (2.9% p.a. for manufacturing, 0.9% for professional services, 0% for EHP) into both specifications, starting from 2020 equilibrium conditions. Under the baseline model, manufacturing employment share declines by 9.1 percentage points, professional services by 5.2 points, and EHP rises by 14.3 points — reflecting that stagnant EHP technology must absorb more workers to meet demand. The conventional Fréchet model produces less contraction in professional services (−2.2 points) and more in manufacturing (−11.5 points). Both specifications predict similar welfare gains (~29 log points). The scope condition is that these projections treat technology growth rates as exogenous and constant at their sample-period averages; they abstract from endogenous innovation, feedback between human capital reallocation and technology, and from demand-side shifts (which Duernecker, Herrendorf, and Valentinyi 2024 and Sen 2021 emphasize).&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-and-differ-from-the-broader-structural-roy-model-literature"&gt;Q10. How does this paper relate to and differ from the broader structural Roy model literature?&lt;/h3&gt;
&lt;p&gt;The paper is in direct dialogue with Lagakos and Waugh (2013), who use a two-sector Roy model with independent Fréchet marginals to explain cross-country agricultural/non-agricultural productivity gaps; Bryan and Morten (2019) and Hsieh et al. (2019), who use multivariate Fréchet to evaluate productivity gains from reducing labor market frictions; and Adamopoulos et al. (2022), Pulido and Świecki (2019), and Gai et al. (2025), who use multivariate normal distributions for similar questions. All these papers find that sector expansion is accompanied by falling average worker quality — a consequence of the parametric restriction that comparative advantage aligns positively with absolute advantage in the same sector. Adão (2016) and Alvarez-Cuadrado, Amodio and Poschke (2019) showed theoretically that this alignment is the key sufficient condition for the conventional result, and found non-parametric evidence against it in some sectors. Shu&amp;rsquo;s contribution is to provide a tractable parametric framework (Gumbel conditional on quantile ranks) that relaxes this restriction, estimate it with the relevant micro moments (earnings gap between incumbents and switchers), and show quantitatively that the relaxation matters enormously — reversing the sign of the selection effect for professional services.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary implication is that policies aimed at accelerating technology growth in EHP (education, health, public administration) are warranted because the cost disease there is genuine and not a mismeasurement artifact. The paper explicitly confirms that low labor productivity growth in services reflects slow true technology growth, especially in EHP where the calibrated 30-year technology growth is essentially zero. The positive selection effect for professional services (1.2 log points over 30 years) is quantitatively small and does not materially offset the technology disadvantage. A secondary implication is that conventional models used in trade and development economics (with independent Fréchet skill distributions) systematically overstate the adverse selection effect of sectoral expansion, leading to overprediction of implied technology growth in professional services by over one-third. Studies using such models to evaluate, for example, gains from reducing labor market frictions should interpret their implied technology parameters with caution. Scope conditions: the model takes technology as exogenous and abstracts from endogenous responses of innovation to worker quality, from demand-side dynamics studied elsewhere, and from industry-level heterogeneity within the broad sub-sectors.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Selection effect (on labor productivity)&lt;/strong&gt;: In this paper, the change in average skill level of workers in a sector induced by reallocation — measured as the difference between measured labor productivity growth and true technology growth. A positive selection effect means incoming workers are more skilled than incumbents on average; a negative effect means they are less skilled. The paper distinguishes the selection effect from the conventional presumption that expansion always produces negative selection.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Absolute advantage (in professional services)&lt;/strong&gt;: A worker&amp;rsquo;s log skill level in professional services, a(i) ≡ ln z_p(i), which determines output contribution to that sector independently of what the worker could earn elsewhere. In the model, absolute advantage is distributed Gumbel conditional on the worker&amp;rsquo;s comparative advantages, with mean α(q_m, q_e) = α_m ln q_m + α_e ln q_e.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Comparative advantage (between sectors)&lt;/strong&gt;: The log ratio of a worker&amp;rsquo;s skill in one sector relative to professional services: s_m(i) ≡ ln(z_m(i)/z_p(i)) for manufacturing and s_e(i) ≡ ln(z_e(i)/z_p(i)) for EHP. A worker&amp;rsquo;s comparative advantage determines which sector they choose when wage rates are equalized, while the relationship between comparative and absolute advantage determines the productivity of workers on the margin of switching.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;α_m and α_e parameters&lt;/strong&gt;: The key parameters governing whether incoming workers from manufacturing (α_m) or EHP (α_e) are more or less productive in professional services than incumbents. When α_m &amp;gt; 0, workers with a high comparative advantage in manufacturing also have high absolute advantage in professional services, so that reallocation from manufacturing raises average quality in professional services. When α_e &amp;lt; 0, workers with high comparative advantage in EHP have low absolute advantage in professional services, so inflows from EHP lower quality. Estimated values: α_m = 0.055, α_e = −0.123.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Baumol&amp;rsquo;s cost disease&lt;/strong&gt;: Used in this paper to refer to the phenomenon whereby the service sector&amp;rsquo;s true technology growth is persistently low relative to manufacturing — implying that resources must continuously be reallocated to services to maintain consumption of service output, raising the relative price of services. The paper confirms this is a genuine technology fact, not a mismeasurement artifact from selection, especially for EHP where 30-year cumulative technology growth is calibrated at essentially −0.6 log points.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Income premium of switching workers&lt;/strong&gt;: The difference in log real weekly earnings between workers who transitioned from a given source sector in the prior year and workers who were already in the destination sector (incumbents), estimated by regression with year and county fixed effects. This premium is the paper&amp;rsquo;s primary empirical moment and the main target for identifying the skill-distribution parameters. Positive premium (MFG→PROF: +0.048) indicates incoming workers are more productive; negative premium (EHP→PROF: −0.143; MFG→EHP: −0.087) indicates they are less productive.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Independent Fréchet specification (benchmark)&lt;/strong&gt;: The conventional parametric Roy model in which each worker&amp;rsquo;s sector-specific skills are drawn independently from Fréchet marginal distributions. This specification implies that workers&amp;rsquo; absolute advantage in a sector is negatively correlated with their comparative advantage — an implicit restriction that forces average skill in any expanding sector to decline with employment share. The paper uses this as the comparison case, with shape parameter 2.7 following Bryan and Morten (2019) and Lagakos and Waugh (2013), and shows it mispredicts the selection effect for professional services by 9.9 log points and reverses its sign.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-homothetic CES preference&lt;/strong&gt;: The demand structure from Comin, Lashkari and Mestieri (2021) used in the model, which allows income elasticities to differ across sectors and vary with aggregate consumption. It governs how structural transformation proceeds on the demand side as incomes grow. Calibrated parameters imply professional services demand is most income-elastic (ξ_p = 1.382) and EHP demand is least income-elastic (ξ_e = 0.644), so growth shifts expenditure toward professional services and eventually toward EHP as incomes rise further.&lt;/p&gt;</description></item><item><title>Self-Fulfilling Prophecies in the Transition to Clean Technology</title><link>https://macropaperwarehouse.com/papers/self-fulfilling-prophecies-in-the-transition-to-clean-technology/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/self-fulfilling-prophecies-in-the-transition-to-clean-technology/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper by Smulders and Zhou challenges the standard lock-in narrative for the slow green transition. The conventional explanation — path dependency in directed technical change (DTC) — is hard to reconcile with forward-looking investors who anticipate an eventual move to clean technology. The authors propose an alternative: strategic investment complementarities among innovators can produce self-fulfilling prophecies that delay the low-carbon transition even when all agents foresee it will ultimately occur.&lt;/p&gt;
&lt;p&gt;The framework is a continuous-time general equilibrium DTC model in the tradition of Acemoglu et al. (2012), modified in two key ways: patents last forever (rather than one period), and labor is mobile between production and R&amp;amp;D. The economy has a clean and a dirty final-goods sector with substitution elasticity σ between them. A continuum of monopolistic intermediate goods suppliers in each sector invest in R&amp;amp;D to improve product quality. The key mechanism is a demand externality: when goods are gross substitutes (σ &amp;gt; 1), innovation in a sector reduces the relative price of that sector&amp;rsquo;s output, shifting consumer expenditure toward it. This raises the return to all innovation in the sector. For σ &amp;gt; 2, this demand externality outweighs the intra-sector business-stealing effect, making within-sector innovations strategic complements — each firm&amp;rsquo;s R&amp;amp;D raises the payoff to R&amp;amp;D for all others in the same sector. The threshold σ &amp;gt; 2 is necessary and sufficient for a coordination problem to arise in the unregulated economy.&lt;/p&gt;
&lt;p&gt;The paper establishes three steady states: two saddlepath-stable corner steady states (one with innovation only in the clean sector, one only in the dirty sector) and an unstable interior steady state with simultaneous R&amp;amp;D. When σ &amp;gt; 2, there exists a range of initial clean market shares θc,0 (the &amp;ldquo;overlap&amp;rdquo;) from which both corner steady states are reachable under rational expectations. The overlap grows with σ and shrinks with impatience ρ (Proposition 3). Furthermore, for any initial condition within the overlap, multiple transition paths to the same corner steady state exist: a &amp;ldquo;fast&amp;rdquo; path with immediate concentration of R&amp;amp;D in one sector, and &amp;ldquo;delayed&amp;rdquo; paths in which firms temporarily innovate in the competing sector before finally converging. For higher σ values, these delays may involve regime switches between the clean-only and dirty-only innovation regimes (σ ∈ [σ-bar, σ-bar-bar)) or even stagnation periods with zero R&amp;amp;D (σ &amp;gt; σ-bar-bar), producing non-monotonic patterns of clean innovation — rises followed by falls before eventual clean dominance (Proposition 4).&lt;/p&gt;
&lt;p&gt;The welfare-maximizing path always leads to the clean steady state: a dirty steady state violates the transversality condition on the carbon stock because unbounded climate damages accumulate. The paper calibrates to 2019 data: initial clean sector share θc,0 = 0.177 (matching the 17.7% renewable energy share in global final energy consumption), world GDP per capita of $11,019 (constant 2015 USD), per capita carbon emissions of 1.22 metric tons, emission intensity ad = 0.198 tonnes per thousand USD, and σ = 1.5. Under this calibration, three distinct equilibrium paths coexist under an optimal Pigouvian carbon tax — one with clean-only innovation from the start and two involving temporary dirty R&amp;amp;D — all converging to the clean steady state but at different speeds and with different amounts of stranded dirty assets.&lt;/p&gt;
&lt;p&gt;The central policy finding (Proposition 7) is that a Pigouvian carbon tax set equal to the social cost of carbon at all times eliminates the dirty steady state but does not pin down a unique transition path. Multiple equilibria with different durations of dirty innovation persist under the first-best carbon tax. Effective coordination requires a second instrument that directly controls relative innovator profitability: a minimum clean revenue guarantee, an emission cap, a dirty R&amp;amp;D tax, or a contingent super-Pigouvian carbon tax all qualify. A clean R&amp;amp;D subsidy works but is an inferior device because it distorts labor allocation between production and research. Crucially, commitment is required: unless the government commits to maintaining the coordination instrument until the economy exits the multiple-equilibria region, delayed transitions remain possible.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-mechanism-generating-multiple-equilibria-and-why-does-it-require-σ--2"&gt;Q1. What is the core mechanism generating multiple equilibria, and why does it require σ &amp;gt; 2?&lt;/h3&gt;
&lt;p&gt;Intermediate good monopolists in each sector earn profits proportional to their sector&amp;rsquo;s expenditure share, which rises with relative quality when σ &amp;gt; 1 (demand shift effect). But a firm&amp;rsquo;s share of sector profits falls as rivals innovate (business-stealing effect). From equation (24), the relative marginal profit of clean versus dirty innovation scales as (Qc/Qd)^(σ-2). The demand shift effect dominates the business-stealing effect if and only if σ &amp;gt; 2. When σ &amp;gt; 2, innovations within a sector are strategic complements: any firm&amp;rsquo;s R&amp;amp;D raises all other firms&amp;rsquo; marginal return to R&amp;amp;D in the same sector. This complementarity means beliefs about which sector will be large in the future become self-reinforcing: if investors expect the clean sector to grow, clean innovation is profitable, and the expectation is validated.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-two-modifications-from-acemoglu-et-al-2012-affect-the-results"&gt;Q2. How do the two modifications from Acemoglu et al. (2012) affect the results?&lt;/h3&gt;
&lt;p&gt;First, infinite (rather than one-period) patents allow future expected profits to influence innovation decisions, giving expectations a more direct role. Second, labor mobility between production and R&amp;amp;D makes the speed of innovation endogenous alongside its direction. However, the paper shows (OA3.2 and Section 3.3) that neither modification is necessary for the qualitative result: the overlap and strategic complementarity arise even with finite patent length and segmented labor markets. Longer patent length has an effect similar to lower impatience — it increases the overlap. OA4 shows that a segmented labor market model has essentially identical dynamics but requires a third state variable (an effective savings-rate proxy), so it is no simpler than the baseline.&lt;/p&gt;
&lt;h3 id="q3-what-types-of-transition-delays-are-possible-and-how-do-they-depend-on-σ"&gt;Q3. What types of transition delays are possible and how do they depend on σ?&lt;/h3&gt;
&lt;p&gt;Proposition 4 identifies three regimes of delay: (a) for 2 &amp;lt; σ &amp;lt; σ-bar, only temporary simultaneous R&amp;amp;D is possible as a delay; (b) for σ ∈ [σ-bar, σ-bar-bar), delay must include temporary regime switches between the clean-only and dirty-only innovation regimes; (c) for σ &amp;gt; σ-bar-bar, delay must include a stagnation period with no R&amp;amp;D at all. The numerical example shows that for σ = 2.5 and σ = 3, delayed paths involve a flat simultaneous-research segment (mc = 1/2). For σ = 5 and σ = 7, equilibrium paths involve switches between clean-only and dirty-only regimes. For σ = 8 and σ = 9, paths contain vertical stagnation sections and multiple regime switches, with clean innovation peaking, falling, then rising again before converging to the clean steady state.&lt;/p&gt;
&lt;h3 id="q4-what-does-the-welfare-analysis-reveal-about-the-costs-of-delayed-transition"&gt;Q4. What does the welfare analysis reveal about the costs of delayed transition?&lt;/h3&gt;
&lt;p&gt;Under the calibrated model (σ = 1.5, θc,0 = 0.177), three equilibrium paths coexist under the Pigouvian carbon tax, corresponding to no delay, short delay, and long delay in clean innovation. Paths with delay accumulate more dirty capital (Qd,∞ &amp;gt; Qd,0), creating more stranded assets in the long run. Figure 4 shows that, at calibrated emission intensity (ad = 0.198), the clean-only path dominates in welfare whenever multiple equilibria arise. However, at a counterfactually low pollution intensity (ad = 0.0198, one-tenth of calibrated), the planner may prefer some temporary dirty innovation when the clean sector starts small, because investment complementarities in the (larger) dirty sector generate higher short-run consumption growth that outweighs the smaller pollution cost.&lt;/p&gt;
&lt;h3 id="q5-why-does-a-pigouvian-carbon-tax-fail-to-coordinate-the-transition-and-what-instruments-can-succeed"&gt;Q5. Why does a Pigouvian carbon tax fail to coordinate the transition, and what instruments can succeed?&lt;/h3&gt;
&lt;p&gt;A Pigouvian tax changes the marginal cost of emissions and affects relative profitability, but it does not fully control relative innovation profitability because strategic complementarities within a sector persist: total innovation in a sector still raises marginal returns for all firms in it, and the complementarity can dominate the tax effect. An emission cap, by contrast, fixes the quantity of dirty output (given the Leontief emissions-to-output structure), which mutes the complementarity: expanding dirty productivity no longer pays if the quantity cap is binding. A minimum clean revenue guarantee sets a floor on clean firms&amp;rsquo; profits that controls relative profitability directly without taxing the dirty sector. A dirty R&amp;amp;D tax raises the marginal cost of dirty research, shifting the innovation regime border and eliminating dirty equilibrium paths. A contingent super-Pigouvian carbon tax (above the social cost of carbon) that activates only when the economy innovates in the dirty sector also works. All of these require policy commitment over the duration of the multiple-equilibria region; without commitment they fail.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-relate-to-and-differ-from-acemoglu-et-al-2012"&gt;Q6. How does the paper relate to and differ from Acemoglu et al. (2012)?&lt;/h3&gt;
&lt;p&gt;The model starts from Acemoglu et al. (2012) but reaches a qualitatively different policy conclusion. Acemoglu et al. (2012) acknowledge the multiplicity of equilibria in their appendix but restrict their analysis to initial conditions and policies that make equilibrium unique, concluding that a Pigouvian tax combined with an R&amp;amp;D subsidy is sufficient for the optimal transition. This paper shows that when forward-looking expectations and investment complementarities are fully accounted for, the coordination failure is separate from the pollution and monopoly externalities, and a Pigouvian tax — even when optimal — does not resolve it. The paper also differs by using infinite patent length (vs. one-period) and an integrated labor market (vs. segmented), though Appendices OA3.2 and OA4 show the qualitative conclusions are robust to these modeling choices.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-paper-relate-to-the-stranded-asset-literature"&gt;Q7. How does the paper relate to the stranded asset literature?&lt;/h3&gt;
&lt;p&gt;Van der Ploeg and Rezai (2020) and Kalkuhl et al. (2020) explain asset stranding through policy uncertainty, distributional effects, or disordered transition. This paper provides a complementary explanation: excess dirty investment and asset stranding can occur even under a committed, fully optimal Pigouvian tax — not because of uncertainty, but because of rational coordination failure. Firms continue investing in polluting technologies, knowing a clean steady state is inevitable, because strategic complementarities make the dirty sector temporarily attractive when the dirty sector is larger. The amount of stranded assets varies across equilibria: the longer the delay in clean innovation, the larger the accumulated stock of ultimately worthless dirty technology capital (Qd,∞ &amp;gt; Qd,0).&lt;/p&gt;
&lt;h3 id="q8-what-role-do-knowledge-spillovers-and-cross-sectoral-knowledge-externalities-play"&gt;Q8. What role do knowledge spillovers and cross-sectoral knowledge externalities play?&lt;/h3&gt;
&lt;p&gt;The baseline model assumes knowledge spillovers within sectors (quality in sector j benefits from sector-wide average quality Qj). The Online Appendix (OA3) shows that inter-sectoral knowledge spillovers (parameter χ) do not affect complementarities at all, because knowledge stock is predetermined and current rival innovation cannot affect one&amp;rsquo;s own value through the knowledge channel. Learning-by-doing production spillovers (parameter ε) strengthen complementarities. The general condition for self-fulfilling prophecies in the extended model is ψ &amp;gt; max{0, -η}, where ψ = (1+ε)(σ-1)(1-α)/(1-ωα) - 1 and η measures own-sector knowledge advantage in innovation productivity. The baseline model (ε=0, ω=1) gives ψ = σ-2, recovering the σ &amp;gt; 2 condition.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The main policy implication is that a single Pigouvian carbon tax is insufficient for the optimal green transition even if credibly committed to; a coordination device is necessary as a second instrument. Scope conditions: (1) This conclusion holds whenever σ &amp;gt; 1 under optimal industry policy (which internalizes monopoly and spillover externalities) — the threshold is lower than σ &amp;gt; 2 in the unregulated economy. (2) The preferred coordination device (revenue guarantee, emission cap, dirty R&amp;amp;D tax, or contingent super-Pigouvian tax) depends on institutional constraints. (3) All coordination devices require policy commitment for the duration of the multiple-equilibria region. (4) The conclusion that the clean-only path is welfare-superior when multiple equilibria arise holds at calibrated emission intensity; at very low pollution intensity the planner might prefer some temporary dirty innovation. (5) The analysis abstracts from uncertainty, heterogeneous beliefs, large players, multiple abatement options, and physical capital — directions for future quantitative work.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-role-of-impatience-ρ-and-patent-length-in-the-size-of-the-coordination-problem"&gt;Q10. What is the role of impatience (ρ) and patent length in the size of the coordination problem?&lt;/h3&gt;
&lt;p&gt;Proposition 3 shows that the overlap (the range of initial conditions admitting multiple equilibria) decreases with impatience ρ. When ρ is large, investors discount future profits heavily, limiting how far ahead expectations can drive current investment choices. In the limit of infinite impatience, only current profit matters and the game collapses to a static one-period coordination problem (Section 3.3). Shorter patent length, modeled as a Poisson patent infringement risk ι (OA3.2), acts identically to higher ρ in the equilibrium dynamics: the dynamics of the model with infringement risk ι are identical to the baseline with ρ replaced by ρ + ι. Hence shorter patents shrink the overlap, and policy must subsidize R&amp;amp;D to compensate for the excessively short investment horizon.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Strategic investment complementarity&lt;/strong&gt;: Within-sector R&amp;amp;D is a strategic complement when σ &amp;gt; 2: one firm&amp;rsquo;s innovation raises the return to other firms&amp;rsquo; innovation in the same sector, because the demand shift effect (innovation increases sector expenditure share) outweighs the business-stealing effect (innovation dilutes rivals&amp;rsquo; profit share). This is not a knowledge spillover but a demand externality operating through the market size of the innovating sector.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Overlap&lt;/strong&gt;: The range of initial clean market shares θc,0 from which both the clean and dirty corner steady states can be reached in a rational expectations equilibrium. The overlap exists if and only if σ &amp;gt; 2 in the unregulated economy (σ &amp;gt; 1 under optimal industry policy), grows with the substitution elasticity σ, and shrinks with impatience ρ or shorter patent length.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Market valuation share (mc)&lt;/strong&gt;: The share of the clean sector in the total marginal value of innovation across sectors, defined as mc = Qcλc / (Qcλc + Qdλd). When mc &amp;gt; 1/2, the economy is in the clean-only innovation regime; when mc &amp;lt; 1/2, in the dirty-only regime; when mc = 1/2, simultaneous research is active. Because mc is a forward-looking, continuous variable, it captures investors&amp;rsquo; collective expectation about future market conditions and directly determines the direction of technical change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Self-fulfilling prophecy (in innovation)&lt;/strong&gt;: An equilibrium in which investors&amp;rsquo; shared belief about the future direction of innovation is rational precisely because all investors, acting on that belief, make it come true. If all investors expect the dirty sector to remain large, they concentrate R&amp;amp;D there, the dirty sector grows, and the belief is confirmed. The same logic applies to clean beliefs. In the paper&amp;rsquo;s context, self-fulfilling prophecies extend to the speed of transition: even if firms agree the economy will eventually go clean, pessimistic beliefs about timing can rationally support periods of dirty innovation before the switch.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Delayed transition&lt;/strong&gt;: An equilibrium path in which the economy ultimately converges to the clean steady state but investors temporarily concentrate R&amp;amp;D in the dirty sector before switching permanently to clean. The delay generates more stranded dirty assets (a higher terminal dirty technology stock Qd,∞) and higher short-run growth (via dirty-sector complementarities) relative to the fast-transition path. Multiple delayed paths may coexist, distinguished by the length of the dirty innovation period and the amount of accumulated dirty capital.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Coordination device&lt;/strong&gt;: A policy instrument that directly controls the relative profitability of clean versus dirty innovation, thereby eliminating the undesired equilibrium paths without relying solely on price incentives. The paper identifies four classes: (1) minimum clean revenue guarantee, (2) emission cap (quantity-based), (3) dirty R&amp;amp;D tax or clean R&amp;amp;D subsidy, and (4) contingent super-Pigouvian carbon tax. All require government commitment for the duration of the multiple-equilibria region. A clean R&amp;amp;D subsidy is inferior because it distorts labor allocation toward innovation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stranded assets&lt;/strong&gt;: In this paper, the dirty technology capital that becomes economically worthless in the clean steady state. The amount of stranding is determined by the dirty technology stock at the moment the economy permanently switches to clean innovation (Qd,∞). Different equilibrium paths — fast vs. delayed transitions — imply different terminal dirty stocks and hence different quantities of stranded assets. Excess stranding relative to the social optimum is a welfare cost of coordination failure.&lt;/p&gt;</description></item><item><title>Serial Entrepreneurship in China</title><link>https://macropaperwarehouse.com/papers/serial-entrepreneurship-in-china/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/serial-entrepreneurship-in-china/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper studies entrepreneurship and new firm creation in China through the lens of serial entrepreneurs (SEs) — individuals who establish more than one firm — contrasting them with non-serial entrepreneurs (Non-SEs). The central question is whether serial entrepreneurs are selected on persistent productive skill or on non-skill advantages such as preferential access to finance, because the two mechanisms have opposite implications for resource allocation: skill-driven serial entrepreneurship raises aggregate productivity, while favoritism-driven serial entrepreneurship generates misallocation.\n\nThe empirical foundation is two administrative datasets for Chinese firms: the Business Registry of China (SAIC), covering the universe of all firms since 1949 with a 2015 snapshot, used for the period 1995–2015; and the Inspection Database (SAIC), providing firm-level income-statement and balance-sheet data, used for 2008–2012 due to data quality constraints. The sample focuses on individually-owned firms (with the largest shareholder being a natural person), covering roughly 17 million entrepreneurs and 20 million firms by 2015. SE firms constitute approximately one-third of all individual-owned firms throughout the period and hold nearly half of all registered capital — making serial entrepreneurship quantitatively central to the Chinese private sector. SE firms have on average about twice the registered capital of Non-SE firms (e.g., 3.22 million yuan vs. 1.91 million yuan in 1995).\n\nTo organize empirical findings the authors develop a two-period Hopenhayn (1992)-style model with collateral-constrained borrowing (k ≤ λe, where k is capital and e is equity). The model generates two competing predictions. If TFP draws across firms started by the same entrepreneur are persistent (AR(1) with autocorrelation ρ), SEs outperform Non-SEs on TFP and the second firm outperforms the first. If instead some entrepreneurs are &amp;ldquo;favored&amp;rdquo; with a less binding collateral constraint (higher λ) and persistence is low, favored entrepreneurs enter more readily, pushing SE TFP below Non-SE TFP while installing more capital conditional on TFP.\n\nEmpirically, the average evidence favors persistent skills: 1st-SE firms are 9% more productive than Non-SE firms (within 2-digit industry, province, and year) and 2nd-SE firms are 18% more productive, both significant at the 1% level. In terms of assets, 1st-SE firms are 40% larger and 2nd-SE firms are 66% larger than Non-SE firms.\n\nThis average premium, however, conceals critical heterogeneity driven by industry-switching behavior. Two-thirds of SEs (67%) start the second firm in a different 2-digit input-output industry (switchers); one-third stay in the same industry (stayers). Stayers&amp;rsquo; 1st-SE and 2nd-SE firms are respectively 49% and 70% more productive than Non-SE firms — accounting for the entire average SE premium. Switchers&amp;rsquo; 1st-SE and 2nd-SE firms are respectively 9% and 11% less productive than Non-SE firms. Despite their TFP deficit, switchers hold at least 7% more capital in both firm generations than stayers. TFP persistence (autocorrelation of log TFP across 1st- and 2nd-SE firms) is twice as high for stayers (0.29) as for switchers (0.14), confirming the model&amp;rsquo;s key identifying assumption that within-industry persistence exceeds cross-industry persistence. The model interprets switchers&amp;rsquo; low-TFP/high-capital profile as the empirical signature of favored entrepreneurs.\n\nThe model further predicts that equity-constrained entrepreneurs should close the first firm when the second is substantially more productive (opportunity cost of capital). Consistently, 1st-SE firms that are shut when the 2nd starts have 32% lower TFP and 13% lower equity than those run concurrently; 2nd-SE firms operated non-concurrently have 8% higher TFP and 22% lower equity than those run alongside the first.\n\nBeyond learning, the paper documents two additional industry-choice motives for switchers. First, a diversification motive: a one-standard-deviation increase in the covariance of returns between the 1st- and 2nd-SE firm industries raises 2nd-SE TFP by 20%, consistent with entrepreneurs demanding a risk premium to enter correlated industries. Second, an input-output complementarity motive: serial entrepreneurs are significantly more likely to choose industries that are upstream-integrated (coefficient 0.46), downstream-integrated (0.47), or complementary (0.41) with the first industry (all significant at 1%), consistent with transaction-cost motives for co-owning trading partners.\n\nThe policy implication is that China&amp;rsquo;s private sector harbors both dynamism — embodied in highly productive stayer SEs driven by persistent skills — and distortion — embodied in low-productivity switcher SEs who enter and accumulate capital through preferential credit access. Since SE firms account for roughly one-third of all firms and nearly half of all capital, the aggregate productivity costs of favoritism-driven serial entrepreneurship are likely significant. Results apply to individually-owned private firms in China over 1995–2015 and may not extend to settings with more uniform financial markets or state-owned firm dynamics.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper does not use a natural experiment or instrumental variables for the main TFP comparisons. It relies on a structural model to interpret conditional correlations, with TFP measured relative to province-industry-year cell averages (2-digit industry, province, and year fixed effects). The theoretical identification comes from the fact that two distinct mechanisms — persistent skills and favoritism — generate opposite predictions on the joint TFP/capital relationship: skill dominance predicts higher TFP for SEs while favoritism predicts lower TFP combined with higher capital. The paper shows both signatures in data for distinct subgroups (stayers and switchers respectively), lending internal consistency. The concurrent/non-concurrent distinction provides an additional layer: the model predicts concurrency depends on equity and the TFP gap between firms, and the data confirm these predictions precisely (Table 7). The main threat is selection on unobservables: entrepreneurs who choose to start second firms may differ from non-SEs along dimensions not captured by the model, such as risk preferences, managerial talent, or social connections, and these could confound the TFP comparisons even within industry-province-year cells.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;Two mechanisms are posited. (1) Persistent skills (ρ &amp;gt; 0 in an AR(1) for TFP across an entrepreneur&amp;rsquo;s firms): positive selection makes SEs more productive and the 2nd-SE more productive than the 1st-SE. (2) Favoritism/credit access heterogeneity (heterogeneous collateral multiplier λ): favored entrepreneurs enter at lower TFP thresholds, so they are over-represented among SEs but have lower TFP and more capital conditional on TFP. The mechanisms are empirically distinguished by using industry switching as a proxy for favoritism. The learning model predicts low-first-period-TFP entrepreneurs switch industry (they do better by searching elsewhere), so favored individuals, who also have low TFP, should be concentrated among switchers. The data show switchers have both lower TFP than Non-SEs and more capital — a pattern only rationalized by favoritism. Stayers exhibit high TFP consistent with persistent skills. TFP persistence (autocorrelation) is twice as high within-industry (stayers, 0.29) as across-industry (switchers, 0.14), confirming the structural assumption separating the two mechanisms.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented-across-se-types"&gt;Q3. What heterogeneity is documented across SE types?&lt;/h3&gt;
&lt;p&gt;First, stayer vs. switcher heterogeneity is the dominant finding: stayers&amp;rsquo; 1st-SE TFP is 49% above Non-SE and 2nd-SE TFP is 70% above Non-SE; switchers&amp;rsquo; 1st-SE TFP is 9% below Non-SE and 2nd-SE TFP is 11% below Non-SE. Switchers have more assets, equity, and registered capital than stayers despite lower TFP (at least 7% more capital). Second, concurrent vs. non-concurrent heterogeneity: 47.5% of SE firms in the 2008–2012 sample are operated concurrently. Non-concurrent 1st-SE firms have 32% lower TFP and 13% lower equity; non-concurrent 2nd-SE firms have 8% higher TFP and 22% lower equity, consistent with equity-constrained optimal capital reallocation. Third, generational heterogeneity: 2nd-SE firms are consistently larger and more productive than 1st-SE firms across all measures (TFP +18% vs. +9%; assets +66% vs. +40%), consistent with high ρ and positive selection into the second firm. Fourth, geographic stability: 72.3% of SEs locate the 2nd firm in the same prefecture as the first, suggesting local knowledge and networks matter for firm creation.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-and-data-restrictions-are-applied"&gt;Q4. What robustness checks and data restrictions are applied?&lt;/h3&gt;
&lt;p&gt;The paper trims the top and bottom 1% of assets and TFP before computing relative TFP. It excludes the 2007–2008 period from return-to-capital calculations (financial crisis concern). It excludes post-2014 registry data because of a registry reform that inflated new registrations and depressed measured exit. It confirms the covariance-TFP diversification result holds when including SE firms not run concurrently. It excludes entrepreneurs who established more than 20 firms (542 individuals, 188,266 firms) to avoid chain-store effects. The paper does not report instrumental-variable estimates, placebo tests, or alternative TFP measures as formal robustness exercises.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Prior work on serial entrepreneurship (Holmes and Schmitz 1990, 1995; Lafontaine and Shaw 2016 for US; Rocha et al. 2015 for Portugal; Shaw and Sørensen 2019, 2022 for Denmark; Felix et al. 2021) uniformly finds SEs are more productive or larger than Non-SEs and attributes this to ability or learning. This paper confirms the average finding but is the first to demonstrate that the premium fully disappears and reverses for industry switchers, and to link this reversal to capital market distortions and favoritism rather than skill. The use of a comprehensive universe of firms (not manufacturing-only or survey-based samples) distinguishes it empirically. The misallocation literature (Hsieh and Klenow 2009; Buera, Kaboski, Shin 2011; Midrigan and Xu 2014; Moll 2014) analyzes distortions across all firms but does not analyze serial entrepreneurship. Song, Storesletten and Zilibotti (2011) and Hsieh and Song (2015) focus on state vs. private sector differences; this paper shows distortions exist within the private sector among individual-owned firms. Contemporaneous work by Shaw and Sørensen (2022) on Denmark documents similar properties of SE firms to the Chinese average findings.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-models-key-structural-propositions"&gt;Q6. What are the model&amp;rsquo;s key structural propositions?&lt;/h3&gt;
&lt;p&gt;Proposition 1: entrepreneurs enter iff TFP z ≥ z*(e), where the entry threshold is decreasing in equity e. Proposition 2: without financial frictions and with ρ &amp;gt; 0, 1st-SE and 2nd-SE firms have higher expected TFP than Non-SE, and 2nd-SE &amp;gt; 1st-SE for sufficiently large ρ. Proposition 3: with frictions, the 2nd-period entry threshold Z(z1, e) is increasing in z1 (opportunity cost of first firm&amp;rsquo;s capital) and decreasing in e. Proposition 4: with frictions and Assumption 1 (equity monotone in TFP) and sufficiently large ρ, SE firms are more productive than Non-SE. Proposition 5: with ρ = 0 and heterogeneous λ, favored entrepreneurs are over-represented among SEs, which then have lower average TFP but more capital conditional on TFP. Proposition 6: concurrent operation is increasing in equity and decreasing in |z2 − z1|. Proposition 7: entrepreneurs stay in the same industry iff 1st-firm TFP exceeds the unconditional mean; stayers have higher TFP than switchers for both SE firms. Proposition 8: with a risk diversification motive, the probability of choosing industry s&amp;rsquo; for the 2nd firm is decreasing in Cov(δs&amp;rsquo;, δs); conditional on choosing s&amp;rsquo;, 2nd-SE TFP is increasing in Cov(δs&amp;rsquo;, δs).&lt;/p&gt;
&lt;h3 id="q7-what-are-the-diversification-and-input-output-linkage-findings"&gt;Q7. What are the diversification and input-output linkage findings?&lt;/h3&gt;
&lt;p&gt;For diversification, the authors construct an industry-level return-on-assets covariance matrix using 2010–2012 Inspection Data (excluding the financial crisis year). A one-standard-deviation increase in the covariance of returns between 1st and 2nd SE firm industries increases 2nd-SE TFP by 20% (significant at 1%), meaning entrepreneurs require a TFP risk premium to enter a correlated industry. In the excess-probability regression for industry choice, the covariance has a coefficient of -0.11 (significant at 1%), confirming switchers prefer industries negatively correlated with their first industry. For linkages, using 2007 Chinese Input-Output tables and Fan-Lang (2000) methodology, the authors find excess probability of industry choice is significantly higher for downstream-integrated industries (0.47), upstream-integrated industries (0.46), and complementary industries (0.41), all at the 1% level in a joint regression. These results hold controlling for 1st-SE industry fixed effects and year of establishment.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper implies that China&amp;rsquo;s private sector suffers from a specific type of misallocation: entrepreneurs with preferential credit access (favored individuals, proxied by industry switchers) establish and expand firms despite lower productivity, crowding out more productive entrepreneurs. Reducing distortions in credit access — leveling the collateral constraint across entrepreneurs — would shift resources toward skill-driven serial entrepreneurs (stayers) and raise aggregate productivity. The scale of the problem is meaningful: SE firms hold roughly half of all capital in the individual-owner sector. Scope conditions: these findings apply to individually-owned private firms in China during 1995–2015, a period characterized by rapid private-sector growth, underdeveloped financial markets, and significant political-economic favoritism. The results abstract from cross-regional and cross-industry variation in financial frictions; if such variation matters (as Brandt, Kambourov and Storesletten 2023 suggest), the aggregate distortion estimates could differ. The paper does not quantify the aggregate TFP losses from misallocation in a counterfactual exercise.&lt;/p&gt;
&lt;h3 id="q9-what-data-limitations-and-caveats-apply"&gt;Q9. What data limitations and caveats apply?&lt;/h3&gt;
&lt;p&gt;The Inspection Data lack employment information, so the authors impute labor input from the labor first-order condition under competitive wages within province-industry-year cells — a valid proxy only if factor market prices are equalized within cells. Revenue is used as a proxy for value added, valid only if intermediate input shares are constant within industry-province-year cells. The registry snapshot is from end-2015, so ownership history must be inferred; the authors note that for over 80% of individual-owned firms the founding owner coincides with the exit-period or current owner. Post-2014 data are excluded due to registry reform contamination. The analysis excludes entrepreneurs who established more than 20 firms (542 individuals, 188,266 firms) to avoid chain-store effects. The analysis excludes SEs who start a 2nd firm through an enterprise they control (expanding the definition would add 300,400 such cases). Concurrent/non-concurrent classification uses the Inspection Data&amp;rsquo;s 2008–2012 window, which may misclassify some firms. The TFP measure is relative within province-industry-year cells, so cross-cell TFP comparisons are not made.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Serial entrepreneur (SE)&lt;/strong&gt;: In this paper, an individual investor who is or has been the largest shareholder in at least two separate firms over the observation period, not necessarily concurrently; 1st-SE refers to the entrepreneur&amp;rsquo;s first firm and 2nd-SE to all subsequent firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-serial entrepreneur (Non-SE)&lt;/strong&gt;: An individual investor who is or was the largest shareholder in exactly one firm over the entire observation window; the benchmark category for TFP and size comparisons.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stayer&lt;/strong&gt;: A serial entrepreneur whose 2nd-SE firm is in the same 2-digit input-output industry as the 1st-SE firm; interpreted in the model as evidence of high industry-specific comparative advantage and high TFP persistence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Switcher&lt;/strong&gt;: A serial entrepreneur whose 2nd-SE firm is in a different 2-digit input-output industry from the 1st-SE firm; interpreted as evidence of either low first-period TFP (learning/Jovanovic motive) or preferential credit access (favoritism motive); empirically identified by lower TFP than Non-SEs combined with more capital.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Favored entrepreneur&lt;/strong&gt;: In the model, an entrepreneur with a less binding collateral constraint (higher λ), representing individuals with preferential access to bank credit or other non-skill advantages; they enter at lower TFP thresholds, are over-represented among SEs, and display the signature pattern of lower TFP combined with more capital conditional on TFP.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Collateral constraint&lt;/strong&gt;: A borrowing limit of the form k ≤ λe, where k is installed capital, e is equity, and λ ≥ 1 is the collateral multiplier; the central financial friction in the model, generating the observed co-movement between TFP, assets, and debt-equity ratios in the data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Concurrent vs. non-concurrent SE operation&lt;/strong&gt;: Whether the entrepreneur&amp;rsquo;s 1st and 2nd firms are both operating simultaneously (concurrent) or the 1st firm is closed before or when the 2nd begins (non-concurrent); the model predicts non-concurrent operation is optimal when equity is scarce and the TFP gap between firms is large, rationalizing the observed pattern that non-concurrent 2nd-SE firms have higher TFP and lower equity.&lt;/p&gt;</description></item><item><title>Strapped for Cash: The Role of Financial Constraints for Innovating Firms, Misallocation and Aggregate Productivity Growth</title><link>https://macropaperwarehouse.com/papers/strapped-for-cash-the-role-of-financial-constraints-for-innovating-firms-misallocation-and-aggregate-productivity-growth/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/strapped-for-cash-the-role-of-financial-constraints-for-innovating-firms-misallocation-and-aggregate-productivity-growth/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Firms that invest heavily in intangible assets — patents, R&amp;amp;D, software — face a structural financing disadvantage: intangibles offer limited collateral value to banks, so intangible-intensive firms can be cut off from credit even when their marginal revenue product of capital (MRPK) exceeds the going interest rate. The paper asks how binding this collateral constraint is in practice, what relaxing it does to firm behavior, and how large the aggregate productivity and misallocation consequences are.&lt;/p&gt;
&lt;p&gt;The empirical setting is a 2015 Norwegian legal reform that, for the first time, allowed firms to pledge patents as stand-alone collateral. Before the reform, a patent could serve as collateral only in conjunction with a physical asset or if it was actively generating revenue; the reform removed both conditions as of 1 July 2015. The change was introduced specifically to ease financing for innovative firms and was narrow in scope — not part of a broader financial reform.&lt;/p&gt;
&lt;p&gt;The empirical analysis draws on matched administrative panel data covering the universe of Norwegian private non-financial joint-stock companies (about 85 percent of all firms with employees) over 2005–2018. The five linked data sets provide annual firm accounts, loan-level bank lending records (firm-bank-year), shareholder and equity issuance records, and the universe of patent applications to the Norwegian Patent Office. The pre-reform window runs 2010–2015; the post-reform window 2015–2018; the 2005–2010 period is used for placebo tests.&lt;/p&gt;
&lt;p&gt;The identification strategy is difference-in-differences. The treatment group consists of firms with at least one patent application in the five years before the reform (2010–2015); the control group consists of firms without a patent portfolio but with similar observable characteristics (size, tangible assets, intangible intensity, profitability, public-funding status), all within the same 2-digit NACE industry. Firm fixed effects and industry-by-year fixed effects are included throughout; control variables are measured pre-reform and interacted with year dummies.&lt;/p&gt;
&lt;p&gt;Firm-level results confirm that treated firms were collateral constrained: (i) the probability of having a bank loan rose by 5.1 percentage points; (ii) the bank debt-to-sales ratio rose by 1.5 percentage points; (iii) the share of short-term debt fell by 2.7 percentage points, consistent with conversion to longer-term collateralized debt; (iv) the number of bank connections rose by 0.144; and (v) the interest rate was unchanged. Simultaneously, the capital stock (total fixed assets) rose by 0.20 log points, employment rose by 0.051 log points, and MRPK fell significantly (–0.224), satisfying the necessary and sufficient conditions for collateral constraint under the theoretical framework. Sales showed no significant change, which the authors attribute to the short post-reform window (only three years). Pre-trend tests using placebo reform years (2010) and pre-2010 periods yield insignificant estimates, supporting parallel trends.&lt;/p&gt;
&lt;p&gt;For young firms (six years old or younger in 2015), there are additional effects: a larger employment response (+0.181 log points for the interaction term) and positive effects on equity issuance (the equity issue dummy rises by 0.137 for young treated firms) and number of shareholders (+0.225 log points). The improvement in debt access appears to have signaled creditworthiness and improved terms of access to equity for young firms. Innovation also rose: the probability of filing at least one patent in 2016–2018 increased by 21.7 percentage points for treated firms relative to the control group, and the count of patent applications increased by 0.936.&lt;/p&gt;
&lt;p&gt;For aggregate quantification, the authors develop a model of monopolistic competition with heterogeneous firms and credit constraints (following Hsieh and Klenow, 2009 and Melitz, 2003). Each constrained firm faces an implicit capital cost of τ times the market interest rate, where τ ≥ 1. The model is solved in changes using exact hat algebra. Under the small-open-economy assumption (capital supply infinitely elastic), removing the constraint raises labor productivity through two channels: (1) reduced within-industry misallocation as firms equalize MRPKs, and (2) capital deepening as constrained firms invest more. The key advantage of the methodology is that the friction τ is identified directly from the DiD capital stock estimate (0.20 log points) combined with observed capital shares (mean α = 0.30) and an elasticity of substitution σ = 4 (from Broda and Weinstein, 2006), sidestepping the need to estimate revenue TFP.&lt;/p&gt;
&lt;p&gt;The median treated firm faces a credit friction of τ = 1.12, implying an implicit capital cost 12 percent above the market rate. Industry output per worker increases by up to 3 percent, concentrated in sectors where treated (innovative) firms hold a large initial market share. The dominant source of this gain is capital deepening: the ratio of economy-wide labor productivity growth to TFP growth is 39:1, meaning within-industry misallocation reduction accounts for only a small fraction of the productivity gain. The aggregate price index falls by 0.6 percent (P-hat = 1.006 in output-per-worker terms), translating to an increase in total output of 6.4 billion NOK (approximately 0.62 billion USD). A back-of-the-envelope calculation using the implicit cost r(τ-1)K yields 7.5 billion NOK, consistent with the model estimate. For comparison, Norway&amp;rsquo;s main innovation subsidy agency disbursed 5.3 billion NOK in 2021, putting the collateral reform&amp;rsquo;s welfare gain in the same order of magnitude.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses difference-in-differences: the treatment group is firms with at least one patent application in 2010–2015; the control group is all other firms matched on size, tangible assets, intangible intensity, profitability, and public-funding status within the same 2-digit NACE industry. Identification requires parallel trends in the absence of the reform. Three tests are conducted: (1) visual inspection of pre-reform trends in the bank loan dummy after residualizing on controls and fixed effects shows broadly similar trajectories; (2) a placebo regression using 2010 as the fake reform year over 2005–2015 yields insignificant coefficients across most credit access measures; (3) a second placebo uses the same 2010–2015 treatment group but compares the pre-2010 period against 2010–2015, again finding insignificant pre-trends. A residual threat is that treated and control firms may differ in unobservable ways that generate differential post-2015 trends unrelated to the reform. The authors address this by conditioning on a rich set of pre-reform firm characteristics interacted with year dummies, but general equilibrium spillovers (e.g., control firms affected by increased competition from treated firms) mean the DiD cannot cleanly capture the aggregate effect, which is why the structural model is needed.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-authors-establish-that-observed-effects-reflect-collateral-constraints-rather-than-mere-debt-substitution"&gt;Q2. How do the authors establish that observed effects reflect collateral constraints rather than mere debt substitution?&lt;/h3&gt;
&lt;p&gt;The theoretical framework makes a sharp prediction: if a firm is unconstrained, an increase in available funding will leave the capital stock and MRPK unchanged (the firm simply substitutes between funding sources). Only a constrained firm will simultaneously (i) increase borrowing, (ii) increase the capital stock, and (iii) show a decline in MRPK as capital is brought closer to its optimal level. The paper documents all three outcomes for treated firms — 5 pp higher probability of bank debt, 0.20 log-point higher capital, and –0.224 significant decline in MRPK — satisfying the necessary and sufficient conditions for collateral constraint. The unchanged interest rate rules out credit becoming cheaper as a confound.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-two-channels-through-which-removing-collateral-constraints-raises-aggregate-productivity-and-how-large-is-each"&gt;Q3. What are the two channels through which removing collateral constraints raises aggregate productivity, and how large is each?&lt;/h3&gt;
&lt;p&gt;The model decomposes industry labor productivity growth (Ys-hat/Ls-hat) into two multiplicative components: (1) TFP growth (TFPs-hat) reflecting reduced within-industry misallocation as capital is reallocated toward previously constrained firms with high MRPK, and (2) capital deepening (Ks-hat/Ls-hat)^alpha reflecting an increase in the aggregate capital-labor ratio as constrained firms invest more. Quantitatively, capital deepening dominates: economy-wide labor productivity growth is 39 times larger than TFP growth. This is because Norway is treated as a small open economy where capital supply is elastic at a fixed world interest rate, so aggregate capital expands substantially when constraints are removed. Under the alternative closed-economy assumption (capital supply fixed, interest rate endogenous), capital deepening would be muted and misallocation reduction would play a larger relative role.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-by-firm-age-is-documented-and-why-does-it-arise"&gt;Q4. What heterogeneity by firm age is documented, and why does it arise?&lt;/h3&gt;
&lt;p&gt;Young firms (six years old or younger in 2015) show larger employment responses (the triple interaction P_t x P_i x Young_i is 0.181, significant at 5%) and are the primary drivers of the shift from short-term to long-term debt (triple interaction –0.114, significant at 1%). Young treated firms also gain more in equity access: equity issuance probability rises by 0.137 (significant at 1%) and number of shareholders rises by 0.225 log points (significant at 10%) compared to older treated firms. The authors argue that for young firms the collateral constraint is more binding — consistent with the broader literature — and that improved bank access signals creditworthiness to equity investors, alleviating information asymmetries. For innovation outcomes, there is no strong differential effect by age.&lt;/p&gt;
&lt;h3 id="q5-how-is-the-structural-credit-friction-τ-identified-from-the-reduced-form-estimates"&gt;Q5. How is the structural credit friction τ identified from the reduced-form estimates?&lt;/h3&gt;
&lt;p&gt;From the structural model, the capital stock of a treated firm changes relative to a control firm as K-hat_si = τ^[α_s(σ-1)+1] x P-hat_s^(σ-1). Inverting this expression (Proposition 1 in the paper) yields τ as a function of the observed capital growth K-hat (from the DiD estimate of 0.20 log points), the capital share α_s (measured from the data as 1 minus wage costs over total costs, mean 0.30), and the elasticity of substitution σ (set to 4 from Broda and Weinstein, 2006). Because the DiD estimate is well-identified from a quasi-natural experiment, τ is identified directly from causal variation rather than from cross-sectional dispersion in MRPK as in the traditional misallocation literature (Hsieh-Klenow). This avoids the measurement error and production function estimation problems inherent in that approach.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-distribution-of-the-credit-friction-τ-across-treated-firms"&gt;Q6. What is the distribution of the credit friction τ across treated firms?&lt;/h3&gt;
&lt;p&gt;Since τ in Proposition 1 varies only with the industry capital share α_s (the other inputs — the DiD estimate and σ — are uniform), variation in τ across firms is entirely driven by cross-industry variation in α_s. The density of τ is concentrated between roughly 1.06 and 1.14. The median treated firm has τ = 1.12, implying an implicit capital cost 12 percent above the market interest rate.&lt;/p&gt;
&lt;h3 id="q7-how-are-aggregate-gains-computed-and-how-large-are-they"&gt;Q7. How are aggregate gains computed and how large are they?&lt;/h3&gt;
&lt;p&gt;The aggregate output gain is computed as 1 minus the aggregate price index P-hat. Using initial expenditure shares β_s and the industry price indices from equation (5), the authors obtain P-hat = 1.006 — a 0.6 percent fall in the aggregate price level, equivalently a 0.6 percent rise in output per worker and real wages. Multiplied by aggregate value added in the data, this yields 6.4 billion NOK (approximately 0.62 billion USD). A separate back-of-the-envelope calculation using the formula r(τ-1)K — the total implicit cost of the constraint — gives 7.5 billion NOK (approximately 0.73 billion USD), with median r = 0.07 and median τ = 1.12. The proximity of the two estimates is offered as a consistency check. These gains accrue over the three post-reform years (2015–2018) and are described as substantial, comparable in magnitude to Norway&amp;rsquo;s main innovation subsidy program (5.3 billion NOK in 2021).&lt;/p&gt;
&lt;h3 id="q8-what-does-the-paper-find-regarding-the-impact-on-innovation-and-why-is-the-innovation-regression-different-from-the-other-regressions"&gt;Q8. What does the paper find regarding the impact on innovation, and why is the innovation regression different from the other regressions?&lt;/h3&gt;
&lt;p&gt;Post-reform innovation (2016–2018) is measured using a patent dummy (equals 1 if the firm files at least one application) and a patent count. The paper finds a 21.7 percentage point increase in the patent dummy and a 0.936 increase in the patent count for treated firms. These regressions are cross-sectional (estimated on the 2015 cross-section) rather than panel DiD, because using patenting pre-reform to define treatment and then examining patenting post-reform as an outcome would create a mechanical correlation. There is no strong age heterogeneity in the innovation response (the interaction with Young is negative for patent count at –0.469, marginally significant, but the patent dummy interaction is insignificant at 0.054).&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-differ-methodologically-from-the-standard-hsieh-klenow-misallocation-approach"&gt;Q9. How does this paper differ methodologically from the standard Hsieh-Klenow misallocation approach?&lt;/h3&gt;
&lt;p&gt;Hsieh and Klenow (2009) infer capital misallocation from cross-sectional dispersion in MRPK across firms, computed from observed factor shares and revenue. This approach requires estimating production functions and is subject to measurement error in capital stock and revenue TFP. The present paper instead identifies the credit friction τ from a quasi-natural experiment (the DiD capital growth estimate), which directly measures the within-sector relative capital response for constrained firms. This sidesteps production function estimation, avoids TFPR measurement issues, and produces a transparent mapping from reduced-form estimates to model primitives. The trade-off is that results are specific to the type of friction being studied (collateral constraints on intangible-intensive firms) rather than summarizing aggregate misallocation.&lt;/p&gt;
&lt;h3 id="q10-what-capital-market-assumption-is-used-in-the-baseline-and-what-is-the-alternative"&gt;Q10. What capital market assumption is used in the baseline, and what is the alternative?&lt;/h3&gt;
&lt;p&gt;The baseline assumes that Norway is a small open economy with an infinitely elastic capital supply at a fixed world interest rate r (exogenous r). Under this assumption, relaxing constraints allows constrained firms to expand their capital stock without crowding out capital from unconstrained firms, generating large capital-deepening gains. The appendix solves the model under the alternative closed-economy assumption where aggregate capital supply is fixed and the interest rate adjusts endogenously. Under the closed-economy assumption, capital deepening is muted (constrained firms can expand only at the expense of unconstrained ones), and the misallocation reduction channel plays a larger relative role. The authors argue the small open economy assumption is more appropriate for Norway.&lt;/p&gt;
&lt;h3 id="q11-what-complementarities-between-debt-and-equity-funding-are-documented-and-what-mechanism-is-proposed"&gt;Q11. What complementarities between debt and equity funding are documented, and what mechanism is proposed?&lt;/h3&gt;
&lt;p&gt;For young treated firms, improved access to bank debt (pledging patents as collateral) is associated with a higher probability of equity issuance (coefficient 0.137) and more shareholders (0.225 log points). The proposed mechanism has two parts: (1) the investment financed by bank loans improves firm profitability and return on equity, attracting investors; (2) obtaining a bank loan credibly signals firm quality to equity investors who face information asymmetries about intangible-intensive firms, facilitating equity access that would not have occurred without the debt catalyst. This complementarity is concentrated in young firms, consistent with information asymmetries being most severe early in the firm life cycle.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-paper-find-about-the-funding-structure-beyond-total-borrowing"&gt;Q12. What does the paper find about the funding structure beyond total borrowing?&lt;/h3&gt;
&lt;p&gt;Beyond the extensive margin (probability of having bank debt, +5.1 pp) and intensive margin (bank debt-to-sales ratio, +1.5 pp), the paper documents a shift in debt maturity: the share of short-term debt in total debt falls by 2.7 percentage points. This is interpreted as firms converting short-term unsecured debt into long-term debt backed by patent collateral. The number of bank connections also rises by 0.144, indicating that treated firms gained access to additional lenders (credit lines) after the reform. The interest rate on bank debt shows no significant change, ruling out a price effect — the reform operated through quantity of credit rather than its cost.&lt;/p&gt;
&lt;h3 id="q13-how-does-this-paper-relate-to-the-broader-intangible-capital-finance-literature"&gt;Q13. How does this paper relate to the broader intangible-capital finance literature?&lt;/h3&gt;
&lt;p&gt;Mann (2018) studies the US, where patent pledging is already common, and finds that strengthened creditor rights over patents raise debt and innovation. Hochberg et al. (2018) show that thicker secondary markets for patents improve debt access. Farre-Mensa et al. (2020) find that getting a patent granted raises the probability of a patent-backed loan. Falato et al. (2022) show that rising intangible intensity explains the trend decline in US corporate debt capacity. Brown et al. (2009) document the importance of financial constraints for R&amp;amp;D financing among young US firms. The present paper differs by: (a) using a reform-based quasi-experiment rather than exploiting existing cross-sectional variation; (b) covering the universe of firms including startups rather than only listed or patent-filing firms; (c) quantifying the aggregate implications for misallocation and growth, which prior work does not; and (d) documenting complementarities with equity funding and innovation.&lt;/p&gt;
&lt;h3 id="q14-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q14. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The main policy implication is that legal reform to improve the pledgeability of intangible assets — specifically patents — can substantially ease financing constraints for innovative firms, with economy-wide productivity gains of comparable magnitude to direct innovation subsidies. The scope conditions are: (1) gains are concentrated in sectors where innovative, intangible-intensive firms hold large initial market shares; (2) the capital-deepening channel — which dominates — requires an elastic capital supply, making the results most directly applicable to small open economies integrated into global capital markets; (3) the reform&amp;rsquo;s effectiveness depended on the prior absence of patent collateral rights (Norway was late relative to other OECD countries where 38% of patenting US firms had already pledged patents by 2013); (4) the short post-reform observation window (three years) may understate long-run effects on sales and productivity, since capital investment takes time to translate into revenue. The results underscore the importance of financial regulation — beyond direct subsidy programs — as a tool for promoting innovation and growth.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Collateral constraint&lt;/strong&gt;: In this paper&amp;rsquo;s framework, a firm is collateral constrained if it holds less capital than it would choose at the interest rate it currently pays — formally K_si &amp;lt; K*_si — because limited pledgeable collateral restricts its access to bank credit. The constraint is parameterized as an implicit capital cost markup τ ≥ 1 above the market rate r, so the firm equates MRPK to τr rather than r.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stand-alone patent collateral&lt;/strong&gt;: The legal status introduced by Norway&amp;rsquo;s 2015 reform under which a firm can pledge patents as collateral independently of any physical asset and regardless of whether the patent is generating current revenue. Before the reform, Norwegian law required patents to be bundled with physical assets or actively used in production before they could serve as collateral.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implicit capital cost (τ)&lt;/strong&gt;: The paper&amp;rsquo;s measure of the severity of a firm&amp;rsquo;s credit constraint: the ratio of the firm&amp;rsquo;s effective cost of capital (MRPK) to the market interest rate r. A firm with τ = 1 is unconstrained (MRPK = r); τ &amp;gt; 1 implies the firm would invest more if it could obtain capital at the prevailing rate. The median treated firm has τ = 1.12, meaning a 12% implicit cost premium.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital deepening (as a source of productivity growth)&lt;/strong&gt;: In the model, removing credit constraints allows previously constrained firms to expand their capital stock, raising the aggregate capital-to-labor ratio without proportionally reducing unconstrained firms&amp;rsquo; capital (under elastic capital supply). This increase in capital intensity per worker raises labor productivity independently of any improvement in allocative efficiency or TFP.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Within-industry misallocation (TFP_s)&lt;/strong&gt;: Following Hsieh and Klenow (2009), the paper defines industry-level TFP as the efficiency loss from heterogeneous MRPKs across firms within a sector. When firms face different implicit capital costs (τ_si), capital is misallocated: some firms use too little capital relative to their productivity. Removing constraints equalizes MRPKs and raises TFP_s, but in the paper&amp;rsquo;s quantitative results this channel is small relative to capital deepening.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pledgeability of intangible assets&lt;/strong&gt;: The extent to which a firm&amp;rsquo;s intangible assets (patents, R&amp;amp;D, goodwill, licenses) can be legally accepted as collateral for bank loans. The paper treats low pledgeability as a market friction specific to intangible-intensive firms — distinct from general credit risk — that results in those firms being systematically credit rationed even when their MRPK exceeds the interest rate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exact hat algebra&lt;/strong&gt;: A solution method due to Dekle, Eaton, and Kortum (2008) in which the model is solved entirely in terms of relative changes (hat variables, e.g., x-hat = x&amp;rsquo;/x) using observed pre-reform values in place of calibrated level parameters. This approach avoids the need to estimate unobservable structural parameters and is used here to compute counterfactual industry and aggregate outcomes after the credit friction is removed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Debt–equity complementarity&lt;/strong&gt;: The paper&amp;rsquo;s term for the finding that improved access to bank debt (via patent collateral) also raises equity issuance and the number of shareholders, especially for young firms. The proposed mechanism is that new bank loans signal creditworthiness to equity investors who face information asymmetries about intangible-intensive firms, making debt and equity complements rather than substitutes in the financing of innovative young firms.&lt;/p&gt;</description></item><item><title>Taxation and Entrepreneurship in the United States</title><link>https://macropaperwarehouse.com/papers/taxation-and-entrepreneurship-in-the-united-states/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/taxation-and-entrepreneurship-in-the-united-states/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates how the level and progressivity of personal income taxes shape entrepreneurial activity in the United States, contributing empirical evidence, theoretical intuition, and a structural quantitative evaluation. The motivation is both descriptive — entrepreneurs own more than 40% of total capital and hire more than half of private-sector workers, yet their share of the population varies substantially across states and time — and normative, given growing policy interest in more redistributive taxation. The central question is whether a more progressive tax system, which simultaneously reduces the risk and the return to entrepreneurship, produces more or fewer entrepreneurs in practice.&lt;/p&gt;
&lt;p&gt;The empirical analysis draws on CPS microdata from 1962 to 2019 (entrepreneurs defined as households where the head or spouse is self-employed, averaging 11.7% of the national population), County Business Pattern data from 1986 to 2018, Business Dynamics Statistics, and NBER TAXSIM. Tax measures — both average tax rates at the 50th, 90th, and 95th percentiles of the national earnings distribution and a parametric Benabou (2002) tax function with a level parameter theta_0 and a progressivity parameter theta_1 — are constructed by applying TAXSIM to a fixed 2010 CPS cross-section across all 51 state-year cells from 1977 to 2019, thereby limiting endogeneity from tax-code-induced changes in the observed income distribution. The benchmark panel regression includes state and year fixed effects, state-level economic and demographic controls, lagged local business cycle variables, and local non-linear time trends; the benchmark outcome is measured two years after the tax change. Instrumental variables — lagged state tax rates plus contemporaneous federal rates — are used to further address endogeneity.&lt;/p&gt;
&lt;p&gt;The core empirical findings are strongly negative across all measures of entrepreneurship and all tax measures. A one-percentage-point increase in the average tax rate at median income reduces the number of entrepreneurs by 4.5% (coefficient -0.0449, significant at 1%); a one-standard-deviation increase in that tax rate (about 2.35 percentage points) implies roughly 9.7% fewer entrepreneurs. Negative effects also hold for college-educated entrepreneurs and for firm-side proxies (number of small establishments, employment at small establishments). For tax progressivity, holding tax level constant, a one-percentage-point increase in the average tax rate at twice average earnings reduces the number of entrepreneurs by about 15%. Using the parametric progressivity measure, an increase in theta_1 of 0.01 (about 60% of the cross-state standard deviation) reduces the total number of entrepreneurs by approximately 10% and the number of small establishments by about 2.5%. These results hold under additional lagged controls, different horizons (negative and significant through about nine years for the count of entrepreneurs, more persistent for firm-side measures), and IV estimation (IV magnitudes are one to three times larger than OLS, with first-stage F-statistics of 136 and 112 for the progressivity instrument). A subsample analysis around major federal tax reform years (1988, 1991–1993, 2001) finds consistent signs but smaller and noisier estimates given the reduced sample size.&lt;/p&gt;
&lt;p&gt;To explain these patterns, the paper develops a life-cycle overlapping-generations incomplete-markets model in the spirit of Quadrini (2000) and Cagetti and De Nardi (2006). Households are heterogeneous in age, innate ability, idiosyncratic labor and entrepreneurial productivity shocks, risk aversion (distributed uniformly over three values), and asset holdings. Entrepreneurs face a collateral constraint (capital bounded by theta times assets), a fixed operating cost each period, and a switching cost when exiting to wage employment. The same progressive tax function applies to both workers and entrepreneurs. The model is calibrated to U.S. data: exogenous parameters include an inverse Frisch elasticity of 1, labor productivity persistence of 0.929 and standard deviation of 0.227 (from Chang and Kim 2007), a 45-year working life, and returns to scale in entrepreneurship of 0.85. Eight parameters — including the discount factor, entrepreneurial productivity persistence and dispersion, operating cost, switching cost, and risk-aversion dispersion — are estimated via simulated method of moments, matching 21 moments including the entrepreneur population share, income and wealth shares of entrepreneurs, fraction of entrepreneurs with negative profits, and aggregate wealth distribution. The model matches the data well on targeted and untargeted moments.&lt;/p&gt;
&lt;p&gt;The main structural counterfactual holds average tax rates constant and varies progressivity. Converting to a flat tax (theta_1 = 0) increases the number of entrepreneurs by about 15% in general equilibrium. Aggregate output rises by about 11% and the capital stock falls by about 27% when progressivity doubles from 0.13 to 0.26 (relative to the benchmark of theta_1 = 0.13). The return effect — more progressive taxes compress the expected return to entrepreneurship relative to wage work — quantitatively dominates the insurance effect (more progressive taxes reduce the variance of entrepreneurial income). The distributional analysis shows that medium-productivity entrepreneurs are more sensitive to tax changes than high-productivity ones; older, wealthier entrepreneurs are also more responsive. For welfare, the socially optimal progressivity level — measured by ex-ante expected lifetime welfare of unborn agents in steady state — is theta_1 = 0.109, only about 16% less progressive than the current U.S. benchmark of 0.13. The welfare gains from this reform are described as tiny. The welfare-optimal policy reflects the trade-off between efficiency losses (from reduced entrepreneurship and output) and distributional gains (from redistribution to below-average-income households, who benefit from more progressive taxation). Raising the average tax level while holding progressivity constant also reduces output and capital, with capital falling by roughly 40% and output by about 10% when the level parameter doubles; these effects interact with progressivity in non-linear ways captured only through the structural model.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-in-the-empirical-analysis-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy in the empirical analysis and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The benchmark strategy is a state-year panel regression with state and year fixed effects, state-level economic and demographic controls (real GDP per capita, sector employment shares), and lagged local GDP growth rates and unemployment rates over four years before the tax measure. The dependent variable is measured two years after the tax change to allow recognition lags. IV instruments are constructed as the sum of the lagged (by two years) state tax rate at the relevant income percentile and the current federal marginal tax rate at that percentile, following Akcigit et al. (2018); for progressivity, lagged theta_1 and theta_0 are used as instruments, with first-stage F-statistics of 136 and 112 respectively, ruling out weak instruments. A further alternative IV constructs hypothetical tax parameters by applying current federal rates to state-level rates lagged by two years via TAXSIM. Main threats are (1) endogeneity of state tax policy to local economic conditions — addressed through the rich set of lagged business cycle controls, state-specific quadratic trends, and IV; (2) income-composition endogeneity in estimating the tax function — addressed by fixing the CPS 2010 sample and scaling incomes by average wage growth rather than using the contemporaneous distribution; (3) short sample periods around major reform years, which make the reform-event analysis underpowered.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-through-which-taxes-affect-entrepreneurial-choice-and-how-are-they-distinguished"&gt;Q2. What are the main mechanisms through which taxes affect entrepreneurial choice, and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;The paper identifies two opposing forces from greater tax progressivity. The return effect: higher progressivity reduces the average after-tax payoff to entrepreneurship, because entrepreneurs earn above-average incomes and the progressive schedule compresses post-tax profits relative to wages. The insurance effect: higher progressivity also reduces the variance of after-tax entrepreneurial income, making entrepreneurship less risky and potentially more attractive to risk-averse agents. The simple theoretical models (mean-variance utility with lognormal profits and CRRA utility) show that the sign of the net effect is theoretically ambiguous. In the quantitative model — and in the data — the return effect dominates: flatter taxes raise entrepreneurial entry. The two effects are separated analytically in the simple model (Section 4) and quantitatively in the structural model by examining partial-equilibrium versus general-equilibrium effects and by isolating the capital demand response (sensitive to progressivity) from the labor demand response (less sensitive).&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-across-entrepreneurs-and-along-the-life-cycle-is-documented"&gt;Q3. What heterogeneity across entrepreneurs and along the life cycle is documented?&lt;/h3&gt;
&lt;p&gt;Empirically, the negative tax effect is larger for college-educated entrepreneurs than for non-college entrepreneurs when measured by high-income tax rates (90th and 95th percentiles), consistent with higher-educated entrepreneurs having higher incomes. In the structural model, medium-productivity entrepreneurs lose the most when progressivity rises: when theta_1 doubles, the medium-productivity group&amp;rsquo;s share falls by 0.84 percentage points from a base of 9.08%, while the high-productivity group falls by only 0.11 points from 3.47%. Older and wealthier households are more sensitive to progressivity changes because the return effect matters more relative to the insurance effect for those who have accumulated wealth. Risk aversion heterogeneity (modeled as uniform dispersion around 2.5) affects saving and occupational choice; more risk-averse households are more sensitive to the variance reduction from progressive taxes, but the model shows this does not reverse the dominance of the return effect in aggregate.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The robustness battery includes: (1) adding state-specific quadratic time trends and longer lags of local business cycle variables; (2) two IV strategies — lagged state tax rates plus current federal rates, and hypothetical tax measures constructed from TAXSIM with lagged state and current federal components; (3) controlling for lagged entrepreneurial activity levels (log number of entrepreneurs and establishments lagged two years); (4) examining effects at horizons from t+0 to t+10 via local projection methods, finding effects most pronounced in the short run and diminishing over about nine years for entrepreneur counts but more persistent for establishment and employment measures; (5) restricting the sample to years around major federal tax reforms (1988, 1991–1993, 2001) and finding consistent negative signs even though magnitudes are weaker given the smaller sample; (6) using alternative measures of progressivity (differences between tax rates at multiples of average earnings) as a robustness check on the parametric theta_1 measure; (7) structural model sensitivity analysis varying each estimated parameter individually to confirm monotonic identification of moments.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-prior-empirical-and-structural-work"&gt;Q5. How does this paper relate to and differ from prior empirical and structural work?&lt;/h3&gt;
&lt;p&gt;Empirically, it extends Gentry and Hubbard (2000), who used PSID data 1978–1993 to document that progressive marginal rates discourage self-employment, and Cullen and Gordon (2007), who used IRS cross-sectional data to study the role of tax incentives in business formation. The current paper uses a much larger micro-level dataset (CPS, CBP, BDS), covers both cross-sectional and time-series variation across all U.S. states from 1962 to 2019, examines a broader set of entrepreneurial outcomes (count, employment, establishment dynamics), and controls rigorously for local trends and business cycles. Structurally, it is in the tradition of Quadrini (2000), Cagetti and De Nardi (2006), and Kitao (2008), but uniquely combines a life-cycle OLG framework with empirically estimated tax progressivity and a novel SMM estimation of key entrepreneurial parameters including risk-aversion dispersion. Unlike Meh (2005), which studies switching from progressive to proportional tax in a similar model, this paper brings empirical discipline via state-level identification and explicitly estimates the optimal progressivity. Unlike Brüggemann (2017), which focuses on optimal top marginal rates, this paper studies the full distribution and links it to state-level quasi-experimental evidence. Scheuer (2014) studies optimal taxation with endogenous entry theoretically; this paper complements that with quantitative general-equilibrium analysis.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The main policy implication is that tax progressivity has a quantitatively large negative effect on entrepreneurship and output: converting to a flat tax (holding average tax revenue constant) would increase the number of entrepreneurs by about 15% and GDP by about 11%. However, the welfare-optimal progressivity is only marginally less than the current U.S. level (optimal theta_1 of 0.109 versus benchmark of 0.13, about 16% less progressive), implying the welfare gains from flattening taxes are tiny. This is because redistribution from high-income entrepreneurs to below-average-income workers and retirees is welfare-improving even as it reduces aggregate output. The results hold in both general equilibrium (where wages and interest rates adjust) and in partial equilibrium (more relevant for state-level comparisons, where PE effects are somewhat stronger). The scope conditions include: the model abstracts from age-dependent taxation, occupational-specific tax treatment, endogenous human capital accumulation by entrepreneurs, wealth taxes, and the distinction between corporate and pass-through taxation. These omitted features could alter the optimal progressivity result.&lt;/p&gt;
&lt;h3 id="q7-what-do-the-general-equilibrium-versus-partial-equilibrium-comparisons-reveal"&gt;Q7. What do the general equilibrium versus partial equilibrium comparisons reveal?&lt;/h3&gt;
&lt;p&gt;Partial equilibrium effects (constant wages and interest rates, approximating the small open economy view of U.S. states) are somewhat stronger than general equilibrium effects. This is consistent with the empirical panel estimates, which more closely correspond to PE since state economies face roughly fixed factor prices from the national market. When progressivity doubles in PE (adjusting average tax), the entrepreneur share falls more than in GE, and the optimal progressivity in PE is higher than in GE because in GE there is an additional channel: lower capital stock from reduced entrepreneurship depresses wages, imposing an additional cost on workers that is absent in PE. This comparison validates using PE as the interpretive benchmark for the empirical regressions.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-model-say-about-the-interaction-between-tax-level-and-tax-progressivity"&gt;Q8. What does the model say about the interaction between tax level and tax progressivity?&lt;/h3&gt;
&lt;p&gt;The model reveals a non-linear interaction that cannot be separated in empirical analysis. When tax progressivity is held at zero (flat tax), the entrepreneur share declines smoothly as the tax level rises. At benchmark progressivity, the entrepreneur share exhibits a non-monotonic relationship with the level: for very low tax levels the share is high, it falls as taxes rise, but at sufficiently high levels the entrepreneur share may rise again because workers&amp;rsquo; wealth effects lead to higher labor supply, partially offsetting the dampening of entrepreneurial returns. At doubled progressivity, the non-monotonicity is more pronounced. Tax revenue also exhibits a Laffer-curve pattern with respect to the level parameter across all progressivity scenarios, though this is not the paper&amp;rsquo;s primary focus.&lt;/p&gt;
&lt;h3 id="q9-what-quantitative-moments-does-the-calibrated-model-match-and-where-does-it-fall-short"&gt;Q9. What quantitative moments does the calibrated model match, and where does it fall short?&lt;/h3&gt;
&lt;p&gt;The model matches an aggregate capital-to-output ratio of 2.716 (data: 2.650), entrepreneur population share of 12.6% (data: 12.1%), employment hired by entrepreneurs of 55.9% (data: 56.0%), share of entrepreneurs with negative profits of 12.2% (data: 11.0%), average exit rate of 9.4% (data: 17.0%, a notable miss), average age of entrepreneurs of 44.4 (data: 49.2, another miss), entrepreneur income and wealth shares across the distribution, and top household wealth shares. The model overshoots capital and wealth shares for the top decile relative to data but matches the middle of the distribution well. The average age and exit rate mismatches are acknowledged; the operating-cost and switching-cost parameters are the primary levers for these, and the paper notes that exit costs (rather than entry costs) are more effective at generating entrepreneurs with negative profits.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Tax progressivity (theta_1)&lt;/strong&gt;: The progressivity parameter in the Benabou (2002) tax function ya/AE = theta_0*(y/AE)^(1-theta_1): a higher theta_1 means after-tax income rises less than proportionally with pre-tax income, implying marginal rates increase with income. In the paper&amp;rsquo;s measure, theta_1 = 0 is a flat tax and the U.S. benchmark is estimated at 0.13. Progressivity is measured separately from the average tax level (controlled by theta_0), allowing the two to vary independently in both empirics and counterfactuals.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Return effect vs. insurance effect&lt;/strong&gt;: The two opposing forces through which tax progressivity affects entrepreneurial choice. The return effect is the compression of average after-tax entrepreneurial profits relative to wages — since entrepreneurs earn above-average incomes, progressive taxes reduce the relative net payoff to entrepreneurship. The insurance effect is the reduction in after-tax income variance for entrepreneurs — progressive taxes act as partial insurance against bad profit realizations. The paper finds the return effect quantitatively dominates in both the simple theoretical models and the calibrated quantitative model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Collateral constraint&lt;/strong&gt;: The restriction k &amp;lt;= Theta*a in the model, where k is the entrepreneur&amp;rsquo;s capital input and a is her asset holdings. This models credit market frictions: an entrepreneur can borrow and invest no more than Theta - 1 times her own wealth in the business. Set to Theta = 0.35 in calibration (following Midrigan and Xu 2014), this constraint links entrepreneurial capital demand to wealth accumulation, making the tax-wealth-capital nexus a central quantitative mechanism.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Entrepreneur switching cost (Gamma_s)&lt;/strong&gt;: A cost paid by an entrepreneur who exits to wage employment in the current period. In the calibrated model, Gamma_s = 1.005 (in units of average earnings). This switching cost generates inertia in occupational choice: entrepreneurs with temporarily low productivity may remain rather than exit, generating the empirical share of entrepreneurs with zero or negative profits. It also contributes to life-cycle patterns of entrepreneurship by raising the bar for exit among older, wealthier incumbents.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ex-ante welfare measure&lt;/strong&gt;: The paper&amp;rsquo;s social welfare criterion: the expected lifetime utility of an unborn agent at the beginning of life (age 1), averaging over all initial states (innate ability, initial labor and entrepreneurial productivity draws), and taking the maximum of the worker and entrepreneur value functions. This differs from ex-post welfare (which conditions on realized occupational choice) and is the basis for the optimal tax progressivity calculation. The welfare-maximizing theta_1 = 0.109 uses this criterion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Progressivity wedge (PW)&lt;/strong&gt;: A summary statistic for tax progressivity defined as PW(y1, y2) = 1 - (1 - T&amp;rsquo;(y2))/(1 - T&amp;rsquo;(y1)) for pre-tax incomes y1 &amp;lt; y2. Under the Benabou tax function, the wedge is uniquely determined by theta_1 and equals zero for a flat tax, approaching 1 as the marginal tax rate at the higher income approaches 100%. This measure allows comparison of progressivity across tax systems independently of the level of tax rates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simulated method of moments (SMM)&lt;/strong&gt;: The estimation procedure used for eight model parameters (discount factor beta, entrepreneurial productivity persistence rho_z and dispersion sigma_z, operating cost Gamma_f, switching cost Gamma_s, labor disutility chi, aggregate productivity A, and risk-aversion dispersion sigma_U). The procedure minimizes the weighted distance between 21 model-implied moments and their data counterparts, with a diagonal weighting matrix that puts larger weights on the aggregate capital-to-output ratio and the overall entrepreneur population share.&lt;/p&gt;</description></item><item><title>Taxation of Capital: Capital Levies and Commitment</title><link>https://macropaperwarehouse.com/papers/taxation-of-capital-capital-levies-and-commitment/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/taxation-of-capital-capital-levies-and-commitment/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Barro and Chari (2024) revisit the long-standing debate over optimal capital income taxation, unifying the Chamley-Judd zero-tax result, the Straub-Werning positive-tax amendment, and the Chari-Nicolini-Teles (2020) commitment-based framework into a single coherent analysis centered on the treatment of the &amp;ldquo;period-zero problem.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The research question is fundamental: under what commitment assumptions is the optimal long-run tax rate on capital income zero, positive, or negative, and does optimal policy require special treatment of the initial period? The paper operates entirely within a deterministic neoclassical growth model with a representative household whose preferences are time-separable, separable between consumption and labor, and homothetic — the &amp;ldquo;standard preferences&amp;rdquo; of Chari et al. (2020). The government&amp;rsquo;s tax instruments are proportional consumption tax rates (τ_t^c), proportional asset-income tax rates (τ_t^k), and possibly a one-time proportional levy on initial assets (l_0 ≤ 1). No empirical estimation is performed; the contribution is analytical and quantitative through calibrated simulation.&lt;/p&gt;
&lt;p&gt;The central theoretical finding is that the transitional dynamics of Chamley-Judd and the fully positive long-run capital taxes of Straub-Werning both derive from the same source: the period-zero Ramsey planner&amp;rsquo;s incentive to impose capital levies on assets that happen to exist at the start of the optimization. In Chamley et al., direct levies are precluded (l_0 = 0) and the capital-income tax rate is capped at 100%, so the planner engineers indirect levies via positive future τ_t^k (possibly forever, as Straub-Werning show) and time-varying consumption taxes. In the Chari-Nicolini-Teles (2020) formulation, the planner instead faces a constraint that household initial wealth in utility units (W_0) must meet a designated threshold (W̃_0). Under this constraint, the optimal policy features a one-time direct capital levy l_0 in period zero, zero asset-income taxes in all periods (τ_t^k = 0 for t ≥ 0), and a uniform consumption tax for all t ≥ 0. The level of l_0 and the consumption tax rate are jointly determined to satisfy the wealth constraint and the government budget.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s main contribution is extending the Chari et al. period-zero commitment to all periods, thereby achieving time-consistency and eliminating period zero&amp;rsquo;s special status. If each period-t policymaker faces a wealth constraint W_t ≥ W̃_t with W̃_t set high enough that the policymaker voluntarily chooses l_t = 0, the full sequence of policies is time-consistent and accords with Woodford&amp;rsquo;s (1999) &amp;ldquo;timeless perspective&amp;rdquo;: period zero is like any other period, capital-income tax rates are always zero, and consumption taxes are constant.&lt;/p&gt;
&lt;p&gt;The appendix provides quantitative validation using a U.S.-calibrated model: government consumption = 20% of output, capital-income tax rate = 38% (initial steady state, from Barro-Furman 2018), public debt = 70% of output, labor-income tax rate = 26%, discount factor β = 0.97 (implying a 3% real interest rate), capital share α = 0.34, and depreciation δ = 0.08. Welfare gains from switching to the Ramsey policy (with the wealth-in-utility constraint set to the pre-reform steady-state value) are 0.82% of steady-state consumption under standard preferences, 0.76% under balanced-growth preferences, and 0.62% under zero-wealth-effect preferences. Under balanced-growth preferences, the capital stock rises monotonically to a new steady state approximately 12% higher, government debt rises about 6 percentage points, the labor-income tax rate stays essentially constant at approximately 30% (roughly 4 percentage points above the old steady state), and the capital-income tax rate is approximately 1% in the first period and then drops quickly to zero. Under zero-wealth-effect preferences, the initial capital-income tax rate is slightly higher at approximately 7% before dropping sharply. Under an extreme scenario with the initial capital stock at half its steady-state level and public debt at twice its normal ratio, the capital-income tax rate starts at approximately 3% and gradually approaches zero. In all three cases, constraining the capital-income tax rate to zero and holding the labor-income tax rate constant yields welfare indistinguishable from the unconstrained Ramsey optimum. The paper concludes that zero taxation of capital income is approximately optimal across all three preference specifications, and that the apparent necessity of positive long-run capital taxes in existing literature is an artifact of the period-zero commitment asymmetry.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-period-zero-problem-and-why-is-it-central-to-the-papers-argument"&gt;Q1. What is the &amp;lsquo;period-zero problem&amp;rsquo; and why is it central to the paper&amp;rsquo;s argument?&lt;/h3&gt;
&lt;p&gt;The period-zero problem refers to the asymmetry in the standard Ramsey formulation whereby the period-zero policymaker can commit to all future tax rates but is not bound by any commitments made in the past. Because assets already in existence at period zero are inelastically supplied ex post, the planner has a strong incentive to expropriate them via a capital levy — directly (l_0) or indirectly through high early tax rates on asset income or non-constant consumption tax rates. Chamley-Judd and Straub-Werning results, while superficially different, both arise from this same incentive. The Barro-Chari paper argues that period zero is in reality just an arbitrary starting point for analysis, not a date on which commitment ability uniquely materializes, and that correctly accounting for this eliminates the period-zero problem.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-chari-nicolini-teles-2020-formulation-differ-from-chamley-et-al-and-what-does-it-imply"&gt;Q2. How does the Chari-Nicolini-Teles (2020) formulation differ from Chamley et al., and what does it imply?&lt;/h3&gt;
&lt;p&gt;Chamley et al. preclude direct capital levies (l_0 = 0) and cap τ_t^k ≤ 1, so the planner engineers indirect capital levies via positive future asset-income taxes and time-varying consumption taxes. Chari et al. (2020) instead constrain the household&amp;rsquo;s initial wealth in utility units (W_0) to be at least a designated threshold W̃_0, but leave all tax instruments unrestricted. Under this constraint, the optimal policy selects a one-time direct capital levy l_0, zero asset-income taxes forever, and uniform consumption taxes. The critical difference is that when l_0 = 0 is the outcome under the Chari et al. formulation, it is an optimizing response to a high W̃_0 rather than an arbitrary restriction, so there is no incentive for indirect levies.&lt;/p&gt;
&lt;h3 id="q3-how-is-time-consistency-achieved-and-what-is-the-timeless-perspective"&gt;Q3. How is time-consistency achieved, and what is the &amp;rsquo;timeless perspective&amp;rsquo;?&lt;/h3&gt;
&lt;p&gt;Time-consistency fails if future policymakers are unconstrained because they will repeat the period-zero capital levy logic for their own &amp;lsquo;initial&amp;rsquo; period. The paper shows that introducing a series of per-period wealth constraints — W_t ≥ W̃_t for all t ≥ 0, where W_t is period-t household wealth in utility units — achieves time-consistency if each W̃_t is set high enough that each policymaker voluntarily chooses l_t = 0. The required sequence of W̃_t corresponds exactly to the wealth path generated by the period-0 policymaker&amp;rsquo;s committed Ramsey plan. When this holds, the analysis conforms to Woodford&amp;rsquo;s (1999) &amp;rsquo;timeless perspective&amp;rsquo;: each policymaker adopts the program that would have been committed to far in the past, period zero is not special, capital-income taxes are always zero, and consumption taxes are constant.&lt;/p&gt;
&lt;h3 id="q4-what-role-do-restrictions-on-tax-instruments-play-and-why-does-the-paper-prefer-wealth-constraints-over-direct-instrument-restrictions"&gt;Q4. What role do restrictions on tax instruments play, and why does the paper prefer wealth constraints over direct instrument restrictions?&lt;/h3&gt;
&lt;p&gt;Direct instrument restrictions — such as banning capital levies (l_t = 0) or forcing τ_t^k = 0 and constant consumption taxes — are vulnerable to circumvention through other instruments. For example, time-varying labor-income tax rates (τ_t^n) introduce intertemporal wedges equivalent to indirect capital levies, so a prohibition on capital-income taxes can be undone by varying labor taxes. Constraints on household wealth in utility units (Eqs. 7 and 8) are robust to this vulnerability because any tax instrument that reduces household utility-unit wealth below the threshold violates the constraint, regardless of which specific instrument is used.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-partial-commitment-interpretation-of-the-per-period-wealth-constraints"&gt;Q5. What is the &amp;lsquo;partial commitment&amp;rsquo; interpretation of the per-period wealth constraints?&lt;/h3&gt;
&lt;p&gt;The paper offers two interpretations. The first is that the sequence of W̃_t was set at the founding of a country (e.g., 1789 for the United States). The more palatable &amp;lsquo;partial commitment&amp;rsquo; interpretation is that each period-t policymaker specifies the wealth commitment W̃_{t+1} for the next policymaker, in exchange for adhering to the commitment W̃_t set by the preceding policymaker. This bilateral exchange generates the same sequence of wealth constraints that would have been set arbitrarily far into the past.&lt;/p&gt;
&lt;h3 id="q6-what-happens-in-the-stochastic-extension-of-the-model"&gt;Q6. What happens in the stochastic extension of the model?&lt;/h3&gt;
&lt;p&gt;In a stochastic setting with fluctuations in government spending, technology, war and peace, etc. (as in Chari et al. 2020, proposition 3), choices of capital levies and tax rates become state-contingent rules, following the Lucas-Stokey (1983) framework. Non-zero direct capital levies are optimal under emergency conditions such as war, pandemic, or major financial crisis, and correspondingly below average during non-emergencies. Consumption and labor-income tax rates follow random-walk-like processes, analogous to the tax-rate smoothing predictions of Barro (1979, 1990) that apply when state-contingent capital levies are unavailable.&lt;/p&gt;
&lt;h3 id="q7-how-is-the-covid-inflation-episode-interpreted-within-this-framework"&gt;Q7. How is the COVID inflation episode interpreted within this framework?&lt;/h3&gt;
&lt;p&gt;The paper interprets the post-2020 rise in the U.S. price level through the fiscal theory of the price level (Cochrane 2023; Barro-Bianchi 2023; Bianchi-Faccini-Melosi 2023). The surge in &amp;lsquo;unfunded&amp;rsquo; government spending during and after the COVID pandemic was financed by the inflation that eroded the real value of nominally-denominated government bonds. This constitutes a state-contingent capital levy on bondholders. A cautionary note is added: the availability of such a mechanism may encourage excessive spending, analogous to Ricardo&amp;rsquo;s (1820) argument for balanced-budget war finance.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-role-of-heterogeneity-among-households-in-potentially-generating-commitment"&gt;Q8. What is the role of heterogeneity among households in potentially generating commitment?&lt;/h3&gt;
&lt;p&gt;The paper discusses two sources. First, drawing on Broner-Martin-Ventura (2010), if the government cares about domestic holders of its bonds but not foreign holders, and if bonds can be traded on secondary markets so the two groups cannot be separated, then default becomes unattractive ex post because it harms domestic residents. This gives the government an incentive to promote secondary markets as a commitment device against sovereign default — potentially extensible to capital taxation commitments. Second, the distinction between old and new capital (e.g., via investment tax credits) partially limits the attractiveness of high capital-income taxes by tying the tax rate on old capital to the rate on new capital, which creates investment disincentives. However, as Straub-Werning demonstrate, this commitment may be too weak to drive the optimal capital-income tax to zero.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-calibration-targets-and-preference-specifications-used-in-the-quantitative-experiments"&gt;Q9. What are the calibration targets and preference specifications used in the quantitative experiments?&lt;/h3&gt;
&lt;p&gt;The model is calibrated to represent the U.S. economy with: government consumption = 20% of output, capital-income tax rate = 38% (from Barro-Furman 2018), public debt = 70% of output, labor fraction of time endowment = 1/3, discount factor β = 0.97 (3% real interest rate), capital share α = 0.34, depreciation δ = 0.08. Three preference specifications are explored: (1) standard preferences (time-separable, separable, homothetic in c and n); (2) balanced-growth preferences with consumption-leisure Cobb-Douglas aggregator and IES = 0.5; (3) zero-wealth-effect preferences. The wealth constraint W̃_0 is set to match the pre-reform steady-state wealth in utility terms.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-detailed-quantitative-results-across-preference-specifications"&gt;Q10. What are the detailed quantitative results across preference specifications?&lt;/h3&gt;
&lt;p&gt;Under standard preferences: capital-income tax rate is always exactly zero, labor-income tax rate is constant, welfare gain = 0.82% of steady-state consumption. Under balanced-growth preferences (IES = 0.5): initial capital-income tax ≈ 1%, quickly drops to zero; capital stock rises ≈ 12% to new SS; government debt rises ≈ 6 pp; labor-income tax ≈ 30% (constant, ≈ 4 pp above old SS of 26%); welfare gain = 0.76%; steady-state public debt under zero-capital-tax policy = 33% of output; initial capital levy l_0 = 0.126; new SS labor tax = 0.297. Under zero-wealth-effect preferences: initial capital-income tax ≈ 7%, drops sharply; welfare gain = 0.62%; l_0 = 0.160; new SS labor tax = 0.301; maximum capital tax rate = 0.070. Under extreme initial conditions (balanced-growth, capital stock at half SS level, debt at twice normal ratio): capital-income tax ≈ 3% initially, approaches zero; l_0 = 0.033; new SS labor tax = 0.400. Across all cases, constraining capital-income tax to zero with constant labor tax yields welfare nearly identical to the unconstrained Ramsey optimum.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-scope-of-the-zero-capital-tax-result-and-what-preference-conditions-support-it"&gt;Q11. What is the scope of the zero-capital-tax result and what preference conditions support it?&lt;/h3&gt;
&lt;p&gt;The zero-capital-tax result holds exactly under standard preferences (time-separable, separable between consumption and labor, and homothetic in consumption and labor), which satisfy the Diamond-Mirrlees-Sandmo-Sadka conditions for uniform taxation of goods. Under balanced-growth preferences, it holds with σ = 1 but not necessarily when σ ≠ 1. Under zero-wealth-effect preferences it does not hold if V is strictly concave. However, the quantitative experiments show that deviations from zero are small and short-lived under all three specifications, so zero capital taxation is approximately optimal across the board.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-relationship-between-the-papers-results-and-tax-rate-smoothing-models"&gt;Q12. What is the relationship between the paper&amp;rsquo;s results and tax-rate smoothing models?&lt;/h3&gt;
&lt;p&gt;Barro (1979, 1990) showed that optimal income-tax rates follow a random walk when capital levies are unavailable. The present paper shows that, once state-contingent capital levies are available (the Lucas-Stokey stochastic extension), consumption and labor-income tax rates also exhibit random-walk-like behavior, as realizations of spending and technology shocks move the optimal tax rates. This provides a unified framework connecting capital levy theory and tax-rate smoothing.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-survivalinstitutional-arguments-for-why-commitment-constraints-might-exist-in-practice"&gt;Q13. What are the survival/institutional arguments for why commitment constraints might exist in practice?&lt;/h3&gt;
&lt;p&gt;The paper suggests a selection argument: societies that fail to maintain commitments of the form W_t ≥ W̃_t severely under-accumulate capital because anticipating capital levies causes households and firms not to invest, potentially causing the economy to effectively disappear. This selection pressure may explain why functioning market economies tend to develop institutions (constitutions, property rights, secondary markets) that approximate the required commitments. Major regime changes, such as the Bolshevik revolution (100% default on Czarist bonds), can destroy these commitments, but many regime changes (e.g., France after World War II) do not fully repudiate prior obligations.&lt;/p&gt;
&lt;h3 id="q14-how-does-this-paper-relate-to-and-differ-from-the-three-main-antecedents-chamley-judd-straub-werning-and-chari-et-al-2020"&gt;Q14. How does this paper relate to and differ from the three main antecedents (Chamley-Judd, Straub-Werning, and Chari et al. 2020)?&lt;/h3&gt;
&lt;p&gt;Chamley (1986) and Judd (1985, 1999) showed zero long-run capital-income tax is optimal under the Ramsey formulation with l_0 = 0 and τ_t^k ≤ 1. Straub-Werning (2020) showed that positive capital-income taxes can be optimal even in the steady state under the same constraints when the IES is below one. Chari et al. (2020) replaced instrument restrictions with a utility-wealth constraint for period zero, obtaining a direct capital levy in period zero plus zero capital-income taxes thereafter. Barro-Chari extend Chari et al.&amp;rsquo;s period-zero constraint to all periods, achieving time-consistency and removing period zero&amp;rsquo;s special status. The novel contribution is the multi-period, time-consistent version of the Chari et al. framework and the quantitative demonstration that zero capital taxation is approximately optimal across preference specifications.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Period-zero problem&lt;/strong&gt;: The asymmetry in the standard Ramsey formulation in which the period-zero policymaker can commit to all future tax rates but faces no commitments from the past, creating a strong incentive to expropriate existing assets via capital levies (direct or indirect); the paper&amp;rsquo;s central target of critique.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital levy&lt;/strong&gt;: A proportional confiscation of asset holdings (l_t), distinct from ongoing taxes on the flow of asset income; a direct capital levy takes a fraction of the stock outright, while indirect capital levies are engineered through high asset-income tax rates or time-varying consumption taxes that reduce the real value of existing wealth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wealth constraint in utility units (W_t ≥ W̃_t)&lt;/strong&gt;: A commitment device, following Chari-Nicolini-Teles (2020) and Armenter (2008), that requires each period&amp;rsquo;s policymaker to leave households with at least a threshold level of wealth measured in units of utility rather than goods; instrumental in eliminating the period-zero problem without directly restricting tax instruments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Timeless perspective&lt;/strong&gt;: Woodford&amp;rsquo;s (1999) principle that the policymaker should adopt the behavior that would have been committed to far in the past contingent on current events, rather than optimizing from the current period taking past expectations as given; the paper shows its Ramsey results conform to this principle once per-period wealth constraints are imposed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time-consistency (in optimal taxation)&lt;/strong&gt;: The property that a tax plan chosen at date 0 will be voluntarily continued by each subsequent policymaker; fails in the Chari et al. (2020) baseline formulation when future policymakers are unconstrained because each will want to re-impose a &amp;lsquo;period-zero&amp;rsquo; capital levy, achieved here only when per-period wealth constraints W_t ≥ W̃_t are sufficient to deter direct levies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Indirect capital levy&lt;/strong&gt;: The engineering of a de facto reduction in the real value of existing wealth through policy instruments other than a direct asset levy — specifically positive tax rates on future asset income (τ_t^k &amp;gt; 0) or non-constant consumption tax rates that alter the present value of after-tax consumption; the mechanism underlying both Chamley-Judd transitional dynamics and Straub-Werning permanent positive capital taxes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Standard preferences&lt;/strong&gt;: Preferences that are time-separable, separable between consumption and labor, and homothetic in consumption and labor (Eq. 1 in the paper: u(c,n) = [c^{1-σ}/(1-σ)] − η·n^{1+Ψ}); the class under which uniform taxation of consumption at all dates and zero tax rates on asset income are exactly optimal, satisfying Diamond-Mirrlees-Sandmo-Sadka conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;State-contingent capital levy&lt;/strong&gt;: In the stochastic extension (following Lucas-Stokey 1983), a capital levy whose magnitude depends on the realized state of the world (e.g., war, pandemic, financial crisis); optimal under emergencies when emergency government spending must be financed, and below average during normal times — the paper interprets post-2020 U.S. inflation as an implicit state-contingent levy on nominal government bonds via the fiscal theory of the price level.&lt;/p&gt;</description></item><item><title>Technology Sophistication Across Establishments</title><link>https://macropaperwarehouse.com/papers/technology-sophistication-across-establishments/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/technology-sophistication-across-establishments/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: How sophisticated are the technologies establishments actually use, and how close are they to the world frontier? Traditional measures (since Ryan-Gross 1943 and Griliches 1957) characterize technology by the presence of one or a few advanced technologies, which (i) cover too few technologies and unrepresentative tasks, (ii) say nothing about how non-adopters produce or how far they are from the frontier, and (iii) ignore the intensity with which a technology is used. The authors argue intensity of use matters for explaining income divergence (Comin-Mestieri 2018), so they build a direct, comprehensive measure of technology sophistication.&lt;/p&gt;
&lt;p&gt;Data and design: The authors construct &amp;ldquo;the grid,&amp;rdquo; a two-dimensional structure with business functions (BF) on the horizontal axis and technologies ranked by sophistication (simplest to world frontier) on the vertical axis. The grid spans 63 business functions (7 general business functions [GBF] relevant to all sectors plus 56 sector-specific business functions [SSBF] across 12 sectors) and a total of 305 technologies. More than 50 industry experts built and ranked the grid before survey administration. The grid is implemented in the Firm Adoption of Technology (FAT) survey, fielded 2019-2023 to 21,055 randomly selected establishments forming nationally representative samples (for establishments with 5+ workers) in 15 countries spanning all income levels (Korea, Poland, Croatia, Chile, Brazil-Ceara, Georgia, Vietnam, four Indian states, Ghana, Bangladesh, Kenya, Cambodia, Senegal, Ethiopia, Burkina Faso), representing a universe of about 2.1 million establishments. The median establishment has 9 workers (mean 34); 20% of workers hold a college degree, 17% are exporters, 18% are multinational-affiliated. FAT records, per BF, which grid technologies are used and which one is &amp;ldquo;most widely used.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Two measures are built at the BF-establishment level on a [1,5] affine scale: MAX (sophistication of the most advanced technology used, reflecting adoption) and MOST (sophistication of the most widely used technology, reflecting both adoption and intensity/diffusion within the firm). Establishment-level measures are simple averages across in-house BFs. Cardinalization is validated three ways: linearity of the sophistication-productivity relationship; correlation above 0.98 with a z-score cardinalization (Bloom-Van Reenen 2007); and median correlation 0.95 with an independent productivity-based (&amp;ldquo;Q&amp;rdquo;) cardinalization for 18 BFs.&lt;/p&gt;
&lt;p&gt;Main findings with magnitudes: (1) Establishments underutilize their most sophisticated adopted technology. In 63% of BFs where multiple technologies are used, MOST is not the most sophisticated available; the MAX-MOST gap appears in 62% of multi-technology BFs. (2) MAX and MOST are distinct upgrading processes: a one-unit rise in the number of technologies (NUM) raises MAX by 0.84 but MOST by only 0.25; MAX explains just 34% of within-establishment MOST variance. (3) Gaps are persistent, not transitory: only weakly related to age (cross-decile correlation -0.29; individual -0.01) and unrelated to time since adoption. (4) Gap frequency falls with income (country-level 51% in Korea to 83% in Burkina Faso; correlation -0.55 with per-capita income) and rises with input scarcity (low human capital, loan denial) and managerial mistakes (perception bias, family ownership, non-exporting). (5) Within-country dispersion in gaps (0.28) is about three times the between-country dispersion (0.09). (6) Establishment-level MAX and MOST average 2.6 and 2.0; both correlate with income (0.78 for MAX, 0.94 for MOST) and with size, human capital, management, exporter and multinational status. (7) Both productivity and profitability rise with sophistication, more strongly for MOST and for agriculture; the association is not smaller in low-income countries, contradicting the &amp;ldquo;appropriate technology&amp;rdquo; hypothesis.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-max-and-most-and-why-are-they-conceptually-distinct"&gt;Q1. What are MAX and MOST, and why are they conceptually distinct?&lt;/h3&gt;
&lt;p&gt;MAX_{f,j} is the sophistication of the most advanced grid technology establishment j uses in business function f; MOST_{f,j} is the sophistication of the most widely used technology in that function. Both lie in [1,5] with MAX &amp;gt;= MOST by construction, and both measure closeness to the world frontier. They are conceptually different: increases in MAX reflect adoption of a new (to the function) more sophisticated technology, whereas increases in MOST can reflect adoption OR the extension/intensification of an already-adopted technology — closer to Mansfield&amp;rsquo;s (1963) concept of intra-firm technology diffusion. The paper&amp;rsquo;s central empirical claim is that these are driven by distinct upgrading processes.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identification-strategy-and-what-does-the-paper-not-claim"&gt;Q2. What is the identification strategy, and what does the paper NOT claim?&lt;/h3&gt;
&lt;p&gt;This is a descriptive/correlational paper, not a causal one. The authors explicitly state their data do not permit causal inference; the productivity, profitability, and characteristic associations are partial correlations from cross-sectional regressions with country and 2-digit sector fixed effects. The BF-level analyses (MAX-NUM, MOST-NUM, MAX-MOST) use establishment and function fixed effects to absorb establishment- and function-specific levels. The main &amp;lsquo;identification&amp;rsquo; work is measurement validity, not causal identification.&lt;/p&gt;
&lt;h3 id="q3-how-are-max-and-most-shown-to-be-distinct-upgrading-processes-empirically"&gt;Q3. How are MAX and MOST shown to be distinct upgrading processes empirically?&lt;/h3&gt;
&lt;p&gt;Three pieces of evidence. First, regressing MAX on NUM (number of technologies) with establishment and function FE yields a coefficient of 0.84 (s.e. 0.01) — near one-to-one — while regressing MOST on NUM yields only 0.25 (s.e. 0.01). Second, regressing MOST on MAX (with FE) shows MAX explains only 34% of within-establishment MOST variance, so MAX is not a sufficient statistic for MOST. Third, MAX and MOST have different distributions (MOST more skewed), different lifecycle profiles, different correlates, and different associations with productivity.&lt;/p&gt;
&lt;h3 id="q4-is-the-max-most-gap-transitory-or-persistent-and-how-is-this-tested"&gt;Q4. Is the MAX-MOST gap transitory or persistent, and how is this tested?&lt;/h3&gt;
&lt;p&gt;Persistent. Three exercises: (i) across age deciles the gap correlates only -0.29 with age (-0.01 at the individual level), with no clear lifecycle pattern by income or size except a decline only among large establishments aged 16+; (ii) the distribution of years since adopting a top-tier technology is similar for BFs with and without a gap, so time does not close it; (iii) splitting top-tier adopters into early vs. recent adopters yields similar MOST distributions. Together these confirm gaps persist long after adoption.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-two-hypothesized-drivers-of-max-most-gaps-and-what-evidence-supports-each"&gt;Q5. What are the two hypothesized drivers of MAX-MOST gaps, and what evidence supports each?&lt;/h3&gt;
&lt;p&gt;(1) Input constraints — scarcity of skilled labor or finance pushes firms to rely on simpler technologies operable by less-educated workers or needing less capital. Supported by the negative coefficient on human capital (college share) and the positive coefficient on the loan-denied dummy. (2) Managerial mistakes — poor management or biased self-perception of one&amp;rsquo;s own sophistication causes suboptimal underuse. Supported by positive correlations with perception bias and family ownership, and a negative correlation with exporter status (competitive pressure narrows the gap); the management z-score association is weak. Across subsamples, input scarcity is more prominent in low-income countries while managerial-mistake proxies are more salient among large establishments (likely from the complexity of managing scale).&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-in-technology-sophistication-is-documented"&gt;Q6. What heterogeneity in technology sophistication is documented?&lt;/h3&gt;
&lt;p&gt;By income: country averages span 1.53 (MAX) and 1.01 (MOST); within-country dispersion (p80-p20) rises with income, more steeply for MOST (0.95 vs 0.33). By sector: agriculture shows greater cross-establishment dispersion in both MAX and MOST than manufacturing or services. Lifecycle: MAX rises gradually with age in all income/size groups, but MOST flattens beyond ~10 years in low-income countries and among small establishments. Size effects on MOST are stronger in high-income countries; on MAX they are similar across income levels. The performance-sophistication link is strongest in agriculture and weakest in services, and is not weaker in low- than high-income countries.&lt;/p&gt;
&lt;h3 id="q7-how-much-of-the-variation-is-across-vs-within-sectors-and-why-does-that-matter"&gt;Q7. How much of the variation is across vs. within sectors, and why does that matter?&lt;/h3&gt;
&lt;p&gt;Following Syverson (2011), sector dummies explain only 14% (2-digit), 20% (3-digit), and 23% (4-digit ISIC) of cross-establishment variance in sophistication — comparable to their explanatory power for productivity (sales per worker). This implies sophistication variation reflects differences in the technologies used to perform similar tasks, not differences in what tasks/goods establishments produce.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-and-validation-checks-are-run"&gt;Q8. What robustness and validation checks are run?&lt;/h3&gt;
&lt;p&gt;Cardinalization: linear approximation of the sophistication-productivity relation; correlation &amp;gt;0.98 with z-score cardinalization; median 0.95 (p25-p75: 0.90-0.98) with a productivity-based Q-cardinalization across 18 BFs; establishment-level baseline-vs-Q correlations of 0.90 (MAX) and 0.91 (MOST). Ranking validity: three-stage expert validation (functionality/integration/automation; novelty and cost; ChatGPT replication) on 14 BFs plus an independent relative-productivity exercise on 18 BFs. Data quality: response rates 15-86% (high for establishment surveys); no significant non-response differences in employment, sophistication, wages, or skill; a Kenya back-check pilot showing 80.6% consistency for technology-use reports; external validation against Korea (KED) and Brazil (RAIS) with cross-establishment correlations above 0.93 for sales/employment and 0.73 for labor productivity; ERP adoption in Korean manufacturing of 32% vs. 40% in Chung-Kim (2021). Establishment-level results are robust to controlling for the in-house fraction of functions.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q9. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;It generalizes the intra-firm diffusion literature (Mansfield 1963; Battisti-Stoneman 2003), which studied a handful of technologies in a few countries, by showing MAX-MOST gaps are widespread and persistent across 63 functions and 15 countries. It parallels Bloom-Van Reenen (2007) on management practices in method (expert rankings, survey scoring, z-scores) and finds supporting evidence for the Bloom-Sadun-Van Reenen (2012) technology-management complementarity. It differs from the US Advanced Business Survey / Acemoglu et al. (2022), which covered five frontier technologies, by being comprehensive and frontier-relative. It contributes new evidence to the agricultural productivity gap (Caselli 2005; Gollin-Lagakos-Waugh 2014) and to the appropriate-technology debate (Basu-Weil 1998; Acemoglu-Zilibotti 2001).&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Because the sophistication-performance association is not smaller in low-income than high-income countries, advanced technologies appear &amp;lsquo;appropriate&amp;rsquo; across income levels — challenging the appropriate-technology hypothesis that poor countries gain little from sophisticated technology. Policy should target not only adoption (MAX) but also the extension of use/intensity (MOST), since MOST is more strongly tied to productivity and profitability. Scope conditions: associations are correlational, not causal; samples are representative only for establishments with 5+ workers; coverage is the 12 surveyed sectors; and the cross-section cannot trace dynamics (the authors plan a longitudinal extension).&lt;/p&gt;
&lt;h3 id="q11-what-do-the-descriptive-technology-use-patterns-show-about-adoption-behavior"&gt;Q11. What do the descriptive technology-use patterns show about adoption behavior?&lt;/h3&gt;
&lt;p&gt;Establishments use about two technologies per function on average; 62.6% of functions use more than one and 28.3% use at least three. Leapfrogging/skipping is rare: among single-technology functions (37.4% of cases), 52.8% use the least sophisticated grid technology, so only about 18% of functions have fully skipped or abandoned simpler technologies. In 70.4% of multi-technology functions one technology used is the least sophisticated available, and sophistication gaps (non-contiguous use) occur in only 25% of functions (27% GBF, 17% SSBF; most common in payments 48%, business administration 34%, sales 28%). Firms thus typically retain dominated technologies rather than abandon them, which is why MAX proxies the full adoption history well. Only 16% of establishments use an ERP system (the most sophisticated business-administration technology).&lt;/p&gt;
&lt;h3 id="q12-any-notable-caveats-about-the-measures-themselves"&gt;Q12. Any notable caveats about the measures themselves?&lt;/h3&gt;
&lt;p&gt;MAX-MOST gaps are ordinal (cardinalization-free), but establishment-level MAX and MOST are cardinal and could be sensitive to the chosen cardinalization — addressed by the validation exercises. Establishment-level measures use only in-house functions (87% of relevant SSBFs and an overwhelming majority of GBFs are in-house; only 3.9% of GBFs not in-house), and results are robust to controlling for the in-house share. The survey deliberately avoided the words &amp;rsquo;technology&amp;rsquo; and &amp;lsquo;sophistication&amp;rsquo; (using &amp;lsquo;methods&amp;rsquo;/&amp;lsquo;processes&amp;rsquo;) to limit social-desirability bias.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The grid&lt;/strong&gt;: A two-dimensional structure mapping each key business function (horizontal axis, task-based) to the range of technologies that can perform it (vertical axis, ranked by sophistication from simplest to the world frontier). Spans 63 business functions (7 general + 56 sector-specific across 12 sectors) and 305 technologies, built and ranked by 50+ industry experts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MAX&lt;/strong&gt;: The sophistication (on a [1,5] affine scale) of the most advanced technology an establishment uses in a given business function. Increases in MAX reflect adoption of a technology new to that function; near one-to-one with the number of technologies used (coefficient 0.84).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MOST&lt;/strong&gt;: The sophistication (on a [1,5] scale) of the most widely used technology in a business function. Changes in MOST reflect both adoption and the intensification/extension of already-adopted technologies — closer to Mansfield&amp;rsquo;s (1963) intra-firm diffusion than to adoption per se; only weakly tied to the number of technologies (coefficient 0.25).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MAX-MOST gap&lt;/strong&gt;: A binary indicator equal to 1 when MAX &amp;gt; MOST in a function with multiple technologies in use — i.e., the most widely used technology is not the most sophisticated one adopted. Present in 62-63% of multi-technology functions, persistent over time, and associated with input scarcity, managerial mistakes, and lower productivity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;FAT survey&lt;/strong&gt;: The Firm Adoption of Technology survey: a cross-section of 21,055 establishments forming nationally representative samples (5+ workers) in 15 countries (2019-2023), implementing the grid plus modules on financials, employment, management practices, and adoption barriers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Appropriate technology hypothesis&lt;/strong&gt;: In this paper&amp;rsquo;s usage, the claim (Basu-Weil 1998; Acemoglu-Zilibotti 2001) that establishments in poor countries underutilize sophisticated technologies because scarce human and physical capital limits the productivity gains those technologies embody. The paper&amp;rsquo;s finding that the sophistication-performance association is not smaller in low-income countries runs counter to this hypothesis.&lt;/p&gt;</description></item><item><title>The Lost Marie Curies and Foregone Economic Growth</title><link>https://macropaperwarehouse.com/papers/the-lost-marie-curies-and-foregone-economic-growth/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-lost-marie-curies-and-foregone-economic-growth/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Women accounted for only 3% of U.S. inventors in 1976 and still just 14% in 2023, a pace of convergence far slower than in law (3% to 49%) or medicine (6% to 46%) over the same period. Under the natural assumption of no innate gender differences in inventive potential, this persistent underrepresentation reveals a misallocation of talent. The paper asks how costly this misallocation is for aggregate productivity and welfare.&lt;/p&gt;
&lt;p&gt;Brouillette develops an overlapping-generations (OLG) model of semi-endogenous growth in the spirit of Jones (1995), in which individuals with heterogeneous innate inventive talent choose sequentially among three decisions: (1) whether to pursue a STEM education (the prerequisite for research), (2) whether to work in research or production, and (3) whether to have children. Three gendered barriers can deter women from their comparative advantage. First, a labor market distortion, modeled as a tax on research earnings, captures discrimination in pay and credit attribution. Second, a child penalty distortion reduces mothers&amp;rsquo; hours in research relative to fathers, amplified by the &amp;ldquo;greedy job&amp;rdquo; nature of research (a premium on long hours). Third, an exposure distortion, modeled as a Bernoulli random variable, captures the probability of ever encountering inventive career opportunities — driven empirically by the absence of female role models.&lt;/p&gt;
&lt;p&gt;The model is calibrated to the U.S. economy using two data sources: PatentsView (all USPTO patents since 1976, covering roughly 1.7 million inventors and 3.7 million patents, with gender inferred from first names) and the U.S. Decennial Census/ACS (demographic and occupational data). Across these sources, female inventors exhibit only marginally higher research productivity than men (consistent with modest positive selection from the earnings tax), while mothers in research work approximately 4.5% fewer hours per week than childless female researchers (fathers work 2.7% more). The small productivity gap and modest hours gap together imply that neither the earnings tax nor the child penalty is the dominant driver; the exposure distortion is inferred as the residual, calibrated to a benchmark female share in research of 23% (average of 19% from PatentsView and 27% from Census/ACS). The resulting distortion estimates are: labor market tax 3.3%, child penalty 7%, and exposure barrier 79%.&lt;/p&gt;
&lt;p&gt;Counterfactual elimination of all three distortions raises U.S. income per person by 14.2% in the long run, compared with only 1.5% from a 30% R&amp;amp;D subsidy in a distortion-free economy. The gain materializes slowly, with a half-life of approximately 76 years, reflecting the semi-endogenous structure (where reallocating talent shifts the level but not the long-run growth rate of living standards) and the OLG structure (where career choices are irreversible, slowing labor reallocation). Aggregate research labor increases by 49% within the first 50 years of the transition — women&amp;rsquo;s research labor more than quadruples while men&amp;rsquo;s shrinks by about 10% — but almost all of the productivity gain operates through the intensive rather than the extensive margin: the aggregate share of inventors barely rises, because exposure barriers blocked many talented women entirely rather than only marginal ones, so lifting them introduces very high-quality new researchers who crowd out less talented men. If the underrepresentation were instead attributed entirely to selection-based barriers (labor market or child penalty), long-run consumption would rise by only 3.6%, less than a quarter of the baseline 14.2%.&lt;/p&gt;
&lt;p&gt;Taking transition dynamics into account, eliminating all distortions is equivalent to permanently raising everyone&amp;rsquo;s consumption by 7.2% (lower than 14.2% because the transition is slow and future gains are discounted back at a rate exceeding the low projected U.S. population growth). Of this welfare gain, 95% comes from higher mean consumption; the remainder comes from reduced consumption inequality and utility from children. The distribution of gains is unequal across time and demographic groups: future cohorts experience an 8.6% permanent consumption increase versus only 1% for surviving cohorts. Among the current generation of inventors, women gain the equivalent of a 1.3% permanent consumption increase while men lose 1.7%, a distributional tension that complicates implementation when current costs are concentrated and future benefits diffuse.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-three-distortions-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy for the three distortions, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The three distortions are identified from three moments, each theoretically linked to a specific distortion through the model&amp;rsquo;s aggregation. The labor market distortion (earnings tax) is identified from the research productivity gender gap: positive selection under this tax implies women should be marginally more productive, and the magnitude of the observed (small) gap pins down a distortion of 3.3%. The child penalty distortion is identified from gender differences in hours worked between parent and non-parent researchers: mothers work 4.5% fewer hours than childless women while fathers work 2.7% more; after normalizing male distortions to zero, the model recovers a child penalty distortion of 7%. The exposure distortion is identified as the residual that explains remaining underrepresentation (23% female share in research) after accounting for the other two mechanisms; it is estimated at 79%. Key threats: (1) The gender productivity gap is measured from PatentsView, which uses name-based gender attribution and citation-weighted patents — both susceptible to gender bias (women are documented to receive 30% fewer citations than men with common names, and are 59% less likely to be credited with authorship on patents they contributed to), so the paper uses stock market valuation and textual similarity of patents as bias-resistant alternatives. (2) The exposure distortion is a residual and could capture other forces not in the model, including occupational preferences, gendered barriers to human capital retention, or mismeasurement of the female researcher share. (3) The model abstracts from the direction of innovation (unlike Einïo, Feng, and Jaravel 2022), so welfare effects through consumption-cost inequality across groups are not captured.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The three mechanisms operate through distinct theoretical channels, which allows moment-based identification. The labor market distortion works through selection on talent: if only highly talented women choose research despite earning below their marginal product, the female researcher pool should be right-shifted in the talent distribution, implying modestly higher measured productivity for women. The empirical counterpart is the gender gap in patent output (quality-weighted patents per career year), controlling for field fixed effects and team size. The child penalty works through hours worked: a higher opportunity cost of childbearing in research (amplified by greedy-job premiums) reduces mothers&amp;rsquo; time in research. The empirical counterpart is the gender gap in hours worked between parents and non-parents in research, from the Census/ACS. The exposure distortion works through the extensive margin of talent — it is a binary probability of ever having access to research as a career path, so it can block even the most talented women, unlike the other two distortions which induce selection. It is identified as the residual after the other two are estimated. The insight that the productivity gap is small and the hours gap is modest together rule out the first two as primary drivers, placing most explanatory weight on the exposure distortion.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-semi-endogenous-growth-framework-differ-from-an-endogenous-growth-approach-and-what-are-the-implications-for-the-results"&gt;Q3. How does the semi-endogenous growth framework differ from an endogenous growth approach, and what are the implications for the results?&lt;/h3&gt;
&lt;p&gt;In semi-endogenous growth (Jones 1995), the long-run per-capita growth rate equals n/[(sigma-1)(1-phi)], determined by population growth and idea difficulty, not by the quantity or quality of researchers. A reallocation of inventive talent therefore cannot raise the long-run growth rate but can raise the level of per-capita consumption by shifting the cumulative stock of ideas and thus the entire trajectory of living standards upward. This stands in contrast to endogenous growth models where reallocating talent can permanently raise the growth rate. The author justifies the semi-endogenous approach on two grounds: (1) despite sustained researcher-population growth in most advanced economies, the per-capita growth rate has not trended up; (2) the framework is qualitatively and quantitatively consistent with the documented fact that &amp;lsquo;ideas are getting harder to find&amp;rsquo; (Bloom et al. 2020, which estimates phi = -2.1 for the aggregate U.S. economy). The implication is that the paper finds more modest effects on productivity growth than prior endogenous-growth models, with the gain materializing entirely as a level shift with a long half-life of ~76 years. Einïo, Feng, and Jaravel (2022), using an endogenous growth model, find that barriers to female innovation reduce the growth rate by 1.4 percentage points; this paper&amp;rsquo;s semi-endogenous model finds a 14.2% level gain with no permanent growth rate effect.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-in-the-gender-gap-is-documented-empirically"&gt;Q4. What heterogeneity in the gender gap is documented empirically?&lt;/h3&gt;
&lt;p&gt;Field heterogeneity: Between the 1990 and 2020 inventor cohorts, the female share in chemistry and metallurgy rose from 13% to approximately 30%, while in fixed constructions and mechanical engineering it rose from under 5% to about 10%. Despite this, male-dominated fields accounted for about 53% of total patents granted in 2023. Importantly, when the inventive productivity gender gap is plotted against the female share across technological fields and cohorts, there is no significant relationship (the slope is -0.09 with a standard error of 0.2), implying selection-based barriers are not the primary driver of field-level disparities. Cohort heterogeneity: By cohort, the female share among new inventors rose from 7.5% (1990 cohort) to 17.6% (2020 cohort). Life-cycle heterogeneity: The inventive productivity gender gap (with women slightly ahead) is primarily a cohort effect rather than a within-career pattern; more recent cohorts show a somewhat larger productivity advantage for women at career onset, but the magnitude remains modest, which argues against gendered human capital depreciation as a leading explanation. Parental status heterogeneity: The fraction of female researchers who are mothers converged to the fraction of male researchers who are fathers over time (both around 40% by 2023, down from an 80% male vs. 40% female gap in 1960), suggesting research has become more accommodating. The child penalty in research (hours worked differential between parents and non-parents) has also narrowed over time and is smaller in research than in non-research occupations.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-conducted"&gt;Q5. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;Five sets of robustness exercises are reported. (1) Degree of increasing returns to scale (gamma): Jones (2002) estimates gamma from 0.05 to 0.33; Peters (2021) estimates 0.6. Across this range, the long-run consumption gain from eliminating all distortions ranges from about 2% to almost 27% for gamma going from 0.05 to 0.6. (2) Talent signal shape parameter (theta_s): With theta_s raised to 2 from the baseline 1.26 (implying greater scarcity of superstar inventors, so fewer marginal researchers are displaced), the long-run gain falls to 8.7% from 14.2%. (3) Demographic parameters (retirement rate d and entry rate b): Setting d to match expected working lives of 20 and 40 years (versus baseline 30) shifts the transition half-life by roughly 6-8 years, leaving long-run income unchanged but moving welfare gains slightly (7.6% or 6.9% vs. baseline 7.2%). (4) Knowledge spillover parameter (phi): Values of 0.5 and -6.2 (lower bound of Bloom et al.) are tested with sigma adjusted to hold gamma constant; long-run income gains remain at 14.2%, while the half-life varies modestly and welfare gains shift by at most 24 basis points. (5) Patent quality metrics: Three alternative measures of patent quality are used — stock market valuation (Kogan et al. 2017), textual &amp;lsquo;importance&amp;rsquo; (Kelly et al. 2021), forward citations, and unweighted counts. Results are consistent across measures, with the bias-resistant metrics (stock market valuation and textual importance) ruling out citation-based bias as a confound.&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-relate-to-and-differ-from-einïo-feng-and-jaravel-2022"&gt;Q6. How does this paper relate to and differ from Einïo, Feng, and Jaravel (2022)?&lt;/h3&gt;
&lt;p&gt;Einïo et al. (2022) is the closest antecedent. That paper develops a two-sector endogenous growth model with heterogeneous consumer tastes and unequal access to innovation across sociodemographic groups including gender, finding that barriers to female innovation are responsible for an 18.2% difference in the cost of living between women and men and reduce the economic growth rate by 1.4 percentage points. Brouillette&amp;rsquo;s paper uses a semi-endogenous growth framework and arrives at a 14.2% long-run level gain in income per person and a 7.2% consumption-equivalent welfare gain, with no permanent effect on the growth rate. Beyond the growth framework, the paper extends the analysis to include labor market discrimination and a child penalty for female researchers, which Einïo et al. do not model. However, Brouillette&amp;rsquo;s model abstracts from the direction of innovation — the idea that women and men produce inventions differently tailored to different users&amp;rsquo; needs — which Einïo et al. show is quantitatively important for cost-of-living inequality. The two papers are therefore treated as providing complementary insights.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-role-model-externality-extension-and-how-does-it-change-the-results"&gt;Q7. What is the role-model externality extension, and how does it change the results?&lt;/h3&gt;
&lt;p&gt;In the baseline model, the exposure distortion is a fixed parameter representing the probability of ever encountering inventive career opportunities. In the extension, this probability is multiplied by a technology friction that depends on the fraction of same-gender and opposite-gender inventors in prior generations, with elasticities calibrated from Bell et al. (2018): own-gender elasticity 0.24 for girls, cross-gender elasticity approximately 0 (statistically insignificant in the underlying regression). This creates a positive externality: current inventors increase exposure probabilities for future cohorts of the same gender, but they are not compensated for this spillover, constituting a market failure. In the extended model, some of what was previously captured as the exposure distortion is now attributed to the technological friction from role model scarcity, and the residual exposure distortion is smaller. The counterfactual elimination of all distortions yields a more modest long-run income gain of 10.6% and a consumption-equivalent welfare gain of 3.8% (compared to 14.2% and 7.2% in the baseline). The role model externality also opens a rationale for temporarily gender-differentiated wage subsidies for female researchers as transitional optimal policy: a welfare-maximizing planner might accept a slightly worse talent allocation today in order to accelerate the expansion of the female role model base, reaching the efficient allocation sooner.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s central policy implication is that interventions targeting exposure to innovation for girls earlier in the pipeline — before entry into the labor market — offer far larger aggregate productivity returns than either conventional R&amp;amp;D subsidies or policies aimed at reducing workplace discrimination or the child penalty in isolation. A 30% R&amp;amp;D subsidy yields only 1.5% long-run income per capita growth versus 14.2% from full elimination of female research barriers. Within those barriers, the exposure distortion alone accounts for the bulk of the gain: if the underrepresentation were entirely due to the labor market or child penalty distortions (selection-based mechanisms), long-run gains would be only 3.6%. Scope conditions and caveats: (1) The framework is calibrated to the U.S. and to patent-based inventors plus Census-classified researchers, so generalization to other settings requires re-estimation of distortions. (2) The semi-endogenous structure implies that gains are level effects, not growth rate effects, and the half-life of ~76 years means that most gains accrue to future rather than current generations. (3) Distributional effects are asymmetric: the current generation of male inventors suffers a 1.7% consumption loss, while future cohorts broadly gain 8.6%; this temporal and demographic incidence complicates implementation. (4) The model abstracts from the direction of innovation, so welfare effects through differential cost-of-living impacts on men and women are not captured. (5) The role model externality extension suggests that affirmative action policies for female researchers may be warranted on efficiency grounds, but the exact form of optimal transitional policy is not fully characterized.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-greedy-job-mechanism-and-how-is-it-quantified"&gt;Q9. What is the &amp;lsquo;greedy job&amp;rsquo; mechanism and how is it quantified?&lt;/h3&gt;
&lt;p&gt;The &amp;lsquo;greedy job&amp;rsquo; concept (Goldin 2021) refers to occupations where extended, inflexible hours are compensated at a premium, making it suboptimal for couples to share labor supply equally and thus imposing a larger effective cost of parenthood on whoever reduces hours (in practice, more often women). In the model, an individual researcher&amp;rsquo;s effective labor supply is proportional to alpha^(1+delta) when they have children (where alpha = 0.93 is the fraction of time parents spend working and delta &amp;gt; 0 governs the additional return to hours in research). This magnifies the talent threshold required for a parent to prefer research over production. The parameter delta is estimated empirically by regressing log hourly wages on log hours worked, an indicator for research occupation, and their interaction (plus controls for age, experience, education, occupation, state, race, marital status, year, gender, and occupation-by-gender fixed effects), using the Census/ACS with over 11.8 million observations. The estimated delta for researchers is 0.004, statistically significant but modest — implying research is a &amp;lsquo;modestly greedy job,&amp;rsquo; less so than law (0.011) or medicine (0.006). This small value of delta constrains the child penalty distortion&amp;rsquo;s aggregate impact and helps explain why the exposure distortion dominates empirically.&lt;/p&gt;
&lt;h3 id="q10-how-is-research-productivity-measured-and-what-biases-are-addressed"&gt;Q10. How is research productivity measured, and what biases are addressed?&lt;/h3&gt;
&lt;p&gt;Research productivity is measured as average quality-weighted patents granted per year over an inventor&amp;rsquo;s career, with experience fixed effects removed before averaging across years. Three patent quality metrics are used: (1) stock market valuation (Kogan et al. 2017), inferred from abnormal stock returns around patent grant announcements — chosen for its resistance to gender bias because it reflects market assessments rather than subjective citation choices; (2) &amp;lsquo;importance&amp;rsquo; (Kelly et al. 2021), measured from textual similarity between patent pairs, rewarding novelty relative to prior patents and influence on subsequent ones, and also robust to citation bias because it would require precise paraphrase rather than mere omission; (3) forward citation counts, acknowledged as potentially biased (Jensen et al. 2018 show women with common names receive 30% fewer citations, while women with rare names receive 20% more); (4) unweighted patent counts. All metrics are adjusted for 3-digit CPC class fixed effects and co-inventorship team size. The results are consistent across all four measures, with women slightly ahead in all cases, suggesting that citation bias does not qualitatively alter the productivity comparison. A further concern is attribution bias: Ross et al. (2022) show women are 59% less likely to be credited with authorship on patents they contributed to, meaning PatentsView may undercount the true female inventor population.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-paper-say-about-the-stem-education-gender-gap-specifically"&gt;Q11. What does the paper say about the STEM education gender gap specifically?&lt;/h3&gt;
&lt;p&gt;Women account for approximately 35% of employed STEM graduates aged 25 to 45 in the Census/ACS data (and less than 20% of engineering graduates). However, this STEM gap alone explains only 7% of the patenting gender gap (Hunt et al. 2013, using the 2003 NSCG which recorded patenting in the prior five years); a substantial 78% of the gap stems from differences in patenting behavior among STEM graduates themselves. Furthermore, since the early 2000s, female researchers have been more likely than male researchers to hold a college degree, ruling out educational attainment differences as the primary driver. The model addresses STEM underrepresentation not through a gendered STEM education cost but through the exposure distortion, on the grounds that: (1) exposure to role models is well-documented as influencing girls&amp;rsquo; decisions to pursue STEM (Carrell et al. 2010; Breda et al. 2023; Bell et al. 2018); and (2) if a higher STEM cost were the primary barrier, the model would predict women to be substantially more productive than men (strong positive selection), which the data does not support.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Semi-endogenous growth&lt;/strong&gt;: A growth framework in which the long-run per-capita growth rate is determined by population growth and the difficulty of finding new ideas (the knowledge spillover parameter phi), not by the quantity or quality of researchers. Reallocating inventive talent shifts the level of living standards permanently but cannot alter the long-run growth rate; &amp;lsquo;ideas are getting harder to find&amp;rsquo; (phi &amp;lt; 0 in the paper&amp;rsquo;s calibration, phi = -2.1) is an integral feature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exposure distortion&lt;/strong&gt;: A Bernoulli random variable with mean (1 - tau_E_gk) governing whether an individual of gender g and cohort k ever encounters inventive career opportunities, regardless of their talent. In the baseline model it captures the aggregate probability of not having relevant role models or other enabling conditions during formative years; it is estimated at 79% for women (meaning only 21% of women are exposed to research as a potential career path). Unlike selection-based distortions, it blocks access to the innovation system even for the most talented women.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor market distortion&lt;/strong&gt;: A proportional tax tau_L on the research earnings of female inventors, representing discrimination in compensation, credit attribution, promotions, and rent-sharing from intellectual property. It induces positive selection: under this tax, only sufficiently talented women prefer research over production, making the average female researcher marginally more productive than the average male researcher. Estimated at 3.3%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Child penalty distortion&lt;/strong&gt;: A proportional reduction tau_C in the effective research hours of mothers, capturing the disproportionate burden of childcare and household responsibilities on women&amp;rsquo;s research careers. Combined with the &amp;lsquo;greedy work&amp;rsquo; parameter delta (the premium on long hours in research), it raises the talent threshold above which a woman who wants children will still choose a research career. Estimated at 7%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Greedy job&lt;/strong&gt;: An occupation, in the sense of Goldin (2021), where working long and inflexible hours is rewarded at a premium over and above what a simple proportional-hours model would predict. In the model, captured by the parameter delta &amp;gt; 0 in the research labor supply function. Estimated at delta = 0.004 for researchers (modest relative to lawyers at 0.011 or doctors at 0.006), implying that research is a modestly greedy job, amplifying the child penalty but not dominating the exposure distortion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive vs. extensive margin of research labor&lt;/strong&gt;: The extensive margin refers to the number (fraction) of people who choose research careers; the intensive margin refers to the average quality (talent-weighted hours) of researchers. The paper&amp;rsquo;s key finding is that the 14.2% long-run income gain from eliminating gender barriers is achieved almost entirely on the intensive margin: the aggregate share of inventors barely rises, but average researcher quality increases substantially because exposure barriers had been blocking the most talented women entirely.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption-equivalent welfare variation&lt;/strong&gt;: The permanent proportional adjustment lambda to every person&amp;rsquo;s consumption in the distorted economy that would make utilitarian social welfare equal to that in the undistorted economy. A lambda of 1.072 (7.2% gain) means permanently raising everyone&amp;rsquo;s consumption by 7.2% would compensate for remaining in the distorted equilibrium rather than transitioning to the undistorted one. It is lower than the 14.2% long-run income gain because the slow transition and the discounting of future population growth reduce the present value of future gains.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inventive productivity gender gap&lt;/strong&gt;: The difference in average quality-weighted patents per year between female and male inventors, after controlling for technological field fixed effects, experience, and co-inventorship team size. Measured across multiple patent quality metrics (stock market valuation, textual importance, forward citations, unweighted counts). In the paper&amp;rsquo;s data, the gap is positive but small — women are slightly more productive — which is the key empirical moment used to identify the (small) labor market distortion and to rule out large selection-based barriers as the primary driver of underrepresentation.&lt;/p&gt;</description></item><item><title>The macroeconomics of automation</title><link>https://macropaperwarehouse.com/papers/the-macroeconomics-of-automation/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-macroeconomics-of-automation/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks a foundational question: can the economy-wide degree of automation be measured coherently from standard macroeconomic data, without relying on technology-specific proxies such as robot counts or AI investment surveys? Existing micro-level proxies are fragmented across technologies and difficult to aggregate, leaving it unclear how automation evolves at the macro level or how it relates to capital deepening, factor shares, and productivity growth. The authors, Hideki Nakamura, Masakatsu Nakamura, and Shota Moriwaki, address this by developing a task-based general equilibrium framework in which the aggregate degree of automation emerges endogenously and is fully identified from observable macroeconomic aggregates.&lt;/p&gt;
&lt;p&gt;The theoretical architecture begins with a continuum of tasks, each exhibiting Leontief technology at the task level. Within each task, capital and labor are perfectly substitutable, but firms choose the least-cost input given factor prices. Tasks are ordered by the relative efficiency of capital to labor; as the wage-to-capital-service-price ratio rises with capital deepening, capital performs an expanding range of tasks. Aggregating task-level Leontief decisions over a firm generates a global (envelope) production function. The paper&amp;rsquo;s first main theorem shows that under a mild regularity condition on task efficiency orderings, this aggregation delivers a standard neoclassical production function. Its second set of results identifies the precise efficiency structure under which the aggregate function takes the CES form: that structure corresponds to a Pareto cumulative distribution of input efficiencies. This Pareto structure yields a clean closed-form relationship: the degree of automation is determined entirely by the capital-labor ratio (in efficiency units) and the elasticity of substitution. When the elasticity exceeds one, the degree of automation equals the capital income share; when the elasticity falls below one, it equals the labor income share. Neutral technical progress leaves the degree of automation unchanged at a given capital-labor ratio; capital-augmenting progress raises it; labor-augmenting progress lowers it.&lt;/p&gt;
&lt;p&gt;The empirical application uses panel data from the 2023 Japan Industrial Productivity (JIP) database covering 52 manufacturing industries from 1994 to 2020 (N = 1,404 industry-year observations; two industries excluded for data quality). The CES production function is estimated via GMM using first-differenced factor-share equations derived from the normalized CES system (de La Grandville 1989 normalization), with five sets of instrumental variables drawn from lagged factor prices, information stock and its price, trade openness, workforce age composition, and part-time employment shares.&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. Under the assumption of neutral technical progress, the elasticity of substitution sigma is significantly above one but close to one, ranging from 1.049 to 1.102 across the five IV sets (all significant at least at the 10 percent level). Under the assumption of capital-augmenting technical progress (gK &amp;gt; 0, gL = 0), sigma ranges from 1.035 to 1.068, again robustly greater than one. Capital-augmenting technical progress is statistically significant across all specifications; labor-augmenting technical progress cannot be confirmed in any specification. The average estimated degree of automation across the 52 industries over the full sample period is 0.417 (standard deviation 0.171, minimum 0.138, maximum 0.811). The average rises steadily from 0.407 in 1994 to 0.426 in 2020, temporarily declining around the 2008 financial crisis before recovering. Substantial heterogeneity persists across industries throughout the sample. The distribution shifts rightward over time but retains a fat left tail, with the mode just above 0.3 and several industries exceeding 0.7.&lt;/p&gt;
&lt;p&gt;The two-level CES extension decomposes aggregate capital into industrial robots and other capital, exploiting a purpose-built robot capital stock constructed via the RAS and perpetual inventory methods (initial year 1985). Industrial robots account for only 0.44 percent of aggregate capital stock on average. The two-level estimation yields higher elasticities (sigma-a between 1.191 and 1.346 across IV sets for the composite-labor margin; sigma-b between 1.049 and 1.096 for the robots-other-capital margin). The degree of automation for the composite rises from 0.398 to 0.430 over the sample, a more pronounced increase than the standard CES estimate, reflecting robots&amp;rsquo; amplifying role in automation.&lt;/p&gt;
&lt;p&gt;The paper benchmarks three automation measures against an internal consistency criterion: the squared distance between the automation degree inferred from the capital-labor ratio and that inferred from output per worker, given the same CES structure. The Pareto-based measure (the paper&amp;rsquo;s preferred measure) achieves a distance of 0.0000319, far below the Cobb-Douglas alternative (0.002484) and the continuity-preserving alternative (0.00999), validating the Pareto efficiency-distribution assumption. The Cobb-Douglas alternative yields a mean automation of 0.500 rising from 0.454 to 0.529; the continuity alternative rises more sharply from 0.208 to 0.589 but is discontinuous and sometimes falls outside the unit interval.&lt;/p&gt;
&lt;p&gt;For policy and theory, the paper&amp;rsquo;s framework implies that Japan&amp;rsquo;s sustained capital accumulation during its prolonged stagnation after 1990 translated into rising automation even without commensurate TFP growth, connecting automation dynamics to the &amp;ldquo;productivity paradox.&amp;rdquo; The model also shows that automation can rise alongside an increasing labor income share when sigma is below one, caution against interpreting a stable or rising labor share as evidence against ongoing automation. The degree of automation provides a unified lens connecting capital deepening, factor shares, and productivity in a single theory-consistent measure.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-and-what-observables-are-used-to-infer-the-degree-of-automation"&gt;Q1. What is the core identification strategy and what observables are used to infer the degree of automation?&lt;/h3&gt;
&lt;p&gt;The degree of automation is identified from the first-order conditions of the CES production function. Under the Pareto efficiency-distribution assumption, the CES structure implies a one-to-one mapping from the aggregate capital-labor ratio (in efficiency units), the share parameter s, and the elasticity of substitution rho to the degree of automation (Theorem 4, Eq. 25 and 31). In practice, the authors estimate the CES production function via GMM on first-differenced factor-share equations, recover rho and gK, and plug those into the formula for the degree of automation. No direct observation of tasks, robots (in the standard CES step), or technology-specific adoption decisions is required.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-threats-to-identification-and-how-do-the-authors-address-them"&gt;Q2. What are the main threats to identification and how do the authors address them?&lt;/h3&gt;
&lt;p&gt;The main threats are endogeneity of the output-to-labor and output-to-capital ratios (both simultaneously determined with factor prices) and measurement error in the capital-labor ratio (arising from industry classification changes and the RAS procedure used to construct robot data). The authors address endogeneity via GMM estimation using five distinct IV sets that include lagged factor prices, information stock and its price, trade openness, and workforce composition variables. They report that elasticity estimates are stable across all five IV sets and across alternative sample windows (including a longer 1973-2011 sample from pre-SNA-revision data), and conclude that measurement error is unlikely to drive the results. The overidentification test is not rejected for any IV set in the baseline CES specification (and for most in the two-level specification).&lt;/p&gt;
&lt;h3 id="q3-what-theoretical-result-connects-the-degree-of-automation-to-factor-income-shares"&gt;Q3. What theoretical result connects the degree of automation to factor income shares?&lt;/h3&gt;
&lt;p&gt;Corollary 1 establishes that under the Pareto efficiency structure (Eq. 22) with competitive factor markets, the degree of automation equals the capital income share when sigma &amp;gt; 1, and equals the labor income share when sigma &amp;lt; 1. This makes the degree of automation directly readable from income-share data in the theoretically preferred case (sigma &amp;gt; 1 for Japan). The empirical results are consistent with this: the average degree of automation across manufacturing industries is close to the average capital income share over the sample, providing a cross-check for Corollary 1.&lt;/p&gt;
&lt;h3 id="q4-why-does-the-paper-use-a-leontief-production-function-at-the-task-level-while-obtaining-a-ces-function-at-the-aggregate-level"&gt;Q4. Why does the paper use a Leontief production function at the task level while obtaining a CES function at the aggregate level?&lt;/h3&gt;
&lt;p&gt;The Leontief specification at the task level reflects the idea of a bottleneck in production: within a single narrowly-defined task, only capital or labor is used (once a task is automated, capital fully replaces labor in that task). Perfect substitutability between capital and labor operates at the extensive margin (which tasks are automated) rather than within a task. The aggregate (envelope) function, formed by varying the automation cutoff as the capital-labor ratio changes, generates any elasticity of substitution from zero to infinity. The Pareto efficiency-distribution assumption pins down the specific case of a CES aggregate.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-two-level-ces-extension-work-and-what-does-it-add"&gt;Q5. How does the two-level CES extension work, and what does it add?&lt;/h3&gt;
&lt;p&gt;The two-level CES nests industrial robots and other capital into a capital composite at the inner level (robots vs. other capital, with elasticity sigma-b), then combines that composite with labor at the outer level (composite vs. labor, with elasticity sigma-a). Robot data for 52 industries are constructed via the RAS and perpetual inventory methods with an initial year of 1985. Because robots account for only 0.44 percent of aggregate capital on average, they have a small direct weight, but the two-level decomposition isolates their specific contribution to the automation margin. The two-level CES estimates sigma-a between 1.191 and 1.346 (higher than the standard CES estimates), and finds that the test of equality between sigma-a and sigma-b is rejected for three of five IV sets, suggesting the two elasticities genuinely differ. The average degree of automation rises more steeply under the two-level estimate (0.398 to 0.430) than under the standard CES estimate (0.407 to 0.426), indicating that explicitly accounting for robots reveals a more pronounced automation trend.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-papers-internal-consistency-criterion-and-how-does-it-rank-alternative-automation-measures"&gt;Q6. What is the paper&amp;rsquo;s internal consistency criterion, and how does it rank alternative automation measures?&lt;/h3&gt;
&lt;p&gt;Internal consistency is defined as the mean squared gap between the degree of automation inferred from the capital-labor ratio (Eq. 37, the paper&amp;rsquo;s preferred measure) and the degree of automation implied by observed output per worker given the same CES structure (Eq. 41). A smaller gap means the measure is more coherent with the CES framework from which it is derived. The Pareto-based measure achieves a distance of 0.0000319, more than seventy times smaller than the Cobb-Douglas alternative (0.002484) and over three hundred times smaller than the continuity-preserving alternative (0.00999). The authors therefore select the Pareto-based measure as most internally consistent with CES production.&lt;/p&gt;
&lt;h3 id="q7-what-is-documented-about-heterogeneity-in-automation-across-industries"&gt;Q7. What is documented about heterogeneity in automation across industries?&lt;/h3&gt;
&lt;p&gt;The degree of automation varies substantially across the 52 manufacturing industries, with a standard deviation of 0.171 and a range from 0.138 to 0.811 in the standard CES estimation. The kernel density in 1994 has a fat left tail with a mode just above 0.3, and several industries already exceed 0.7. The distribution shifts rightward by 2020 but remains dispersed. The authors split industries into those with an increasing capital income share (34 industries) and those with a decreasing share (18 industries) and test whether the elasticity of substitution differs between groups; they find no statistically significant difference for any IV set, implying the CES structure is uniform across industries even though automation levels differ.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-connect-automation-to-tfp-and-the-productivity-paradox"&gt;Q8. How does the paper connect automation to TFP and the productivity paradox?&lt;/h3&gt;
&lt;p&gt;The theoretical framework shows that automation via task reallocation shifts the production function in a northeast direction in (k, y) space but does not shift it upward in a way that registers as TFP growth. Formally, increasing automation does not appear to impact TFP growth (citing Nakamura and Nakamura, 2008). The empirical finding that the degree of automation rose from 0.407 to 0.426 during Japan&amp;rsquo;s prolonged stagnation (1994-2020), a period of slow output-per-worker growth, is consistent with this: capital accumulation drove automation forward even though measured TFP growth was subdued. The paper thus links automation dynamics to Japan&amp;rsquo;s productivity paradox and implies that standard TFP accounting may understate the technological transformation underway.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-relationship-between-the-elasticity-of-substitution-and-the-direction-of-factor-share-changes-under-automation"&gt;Q9. What is the relationship between the elasticity of substitution and the direction of factor share changes under automation?&lt;/h3&gt;
&lt;p&gt;The CES framework implies that when sigma &amp;gt; 1 (capital and labor more substitutable), capital accumulation raises the capital income share and lowers the labor share; the degree of automation equals the capital income share. When sigma &amp;lt; 1, capital accumulation raises the wage-to-rental ratio by more, increasing the labor income share; the degree of automation equals the labor income share. In both cases automation rises with capital deepening. A key implication is that observing a stable or rising labor income share does not rule out rising automation when sigma is below one or close to one. The authors&amp;rsquo; estimate of sigma slightly above one for Japanese manufacturing implies a slightly rising capital share, consistent with the panel-estimated trend (b-hat = 0.00102, t-value = 6.84).&lt;/p&gt;
&lt;h3 id="q10-what-are-the-robustness-checks-and-how-stable-are-the-estimates"&gt;Q10. What are the robustness checks and how stable are the estimates?&lt;/h3&gt;
&lt;p&gt;Robustness checks include: (1) five distinct IV sets spanning different combinations of lagged wages, capital rental prices, information stock, trade openness, and workforce composition; (2) estimation under both neutral and capital-augmenting technical progress assumptions; (3) estimation using a longer sample (1973-2011 using pre-SNA-revision data), which yields a sigma still significantly above one and close to one, with slightly larger capital-augmenting technical progress reflecting higher growth in that period; (4) estimation of the full CES production function equation simultaneously with the two FOC equations (Appendix E.2), yielding similar elasticity estimates; (5) a structural change test splitting industries by capital-share trend, finding no significant difference in elasticity between subgroups. Unit root tests (Harris-Tzavalis and augmented Dickey-Fuller) confirm stationarity of all key variables except the part-time ratio, which also passes the ADF test.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-caveats-and-acknowledged-limitations"&gt;Q11. What are the caveats and acknowledged limitations?&lt;/h3&gt;
&lt;p&gt;The authors acknowledge several limitations. First, three conditions cannot be simultaneously satisfied: a CES aggregate, the degree of automation lying in the unit interval, and continuity of the automation measure at unit elasticity (sigma = 1). The preferred measure prioritizes the unit-interval restriction and sacrifices continuity at sigma = 1, making direct comparisons across the sigma &amp;lt; 1 and sigma &amp;gt; 1 cases problematic (an alternative continuous measure is derived in Appendix C but may fall outside the unit interval). Second, the framework abstracts from the creation of new tasks; changes in the total number of tasks over time would affect the automation measure. Third, the paper does not decompose automation by skill level; the observed differences between skilled and unskilled labor in automation suggest a need for nested CES structures in future work. Fourth, the two-level CES nesting (robots within capital composite) is dictated by data availability; alternative nestings, such as grouping robots and labor at the first level, are not separately identifiable.&lt;/p&gt;
&lt;h3 id="q12-how-does-this-paper-differ-from-and-improve-upon-the-prior-literature"&gt;Q12. How does this paper differ from and improve upon the prior literature?&lt;/h3&gt;
&lt;p&gt;The paper improves on micro-proxy approaches (robot counts, AI investment, task-exposure indices from Acemoglu-Restrepo 2020, Adachi 2025, etc.) by providing an aggregate, theory-consistent measure that does not require technology-specific data. It extends prior CES microfoundation work (Jones 2005 Pareto-Cobb-Douglas result, Growiec 2008 Weibull-CES results) by deriving the Pareto efficiency structure that yields CES specifically from task-level automation decisions. It improves on the authors&amp;rsquo; own prior work (Nakamura and Nakamura 2008, Nakamura 2009, 2010) by providing a complete theoretical justification for input efficiencies, a full treatment of the elasticity of substitution, and an empirical implementation. Relative to Artuc et al. (2023) and Adachi (2025), which use Frechet distributions for task productivity, this paper uses a deterministic framework with Pareto-distributed input efficiencies and emphasizes aggregate-level identification rather than cross-occupational substitution.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-policy-implications"&gt;Q13. What are the policy implications?&lt;/h3&gt;
&lt;p&gt;The paper does not make direct policy prescriptions, but its framework has several implications. First, policymakers tracking automation can use standard national accounts data (capital stock, labor input, output, factor shares) rather than waiting for technology-specific surveys, enabling faster and more comprehensive monitoring. Second, the result that automation can advance during periods of slow TFP growth suggests that technology policy focused solely on productivity metrics may underestimate the pace of labor displacement. Third, the finding that Japan&amp;rsquo;s capital accumulation drove automation even through prolonged stagnation implies that capital subsidies or policies encouraging investment could accelerate automation independent of TFP. Fourth, the model&amp;rsquo;s prediction that automation rises alongside increasing labor shares under low substitutability (sigma &amp;lt; 1) warns against complacency: labor-income gains and technology-driven labor displacement can coexist. Fifth, the need for future work on skill heterogeneity and task creation suggests that the framework can be extended to inform distributional policies.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Degree of automation&lt;/strong&gt;: In this paper, the share of the unit task continuum performed by capital rather than labor, denoted a_t, ranging from 0 to 1. It is determined endogenously in equilibrium by relative factor prices and increases with the capital-labor ratio. It is distinct from any technology-specific proxy and emerges as a function of aggregate macroeconomic observables.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Task-based production framework&lt;/strong&gt;: A model in which output requires completing a continuum of tasks, each exhibiting Leontief technology at the task level (capital and labor are perfectly substitutable within a task, but the firm either fully automates a task or uses labor exclusively). Tasks are ordered by the relative efficiency of capital to labor, and firms choose the automation cutoff that minimizes cost given factor prices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pareto efficiency distribution&lt;/strong&gt;: The specific parametric form of aggregate capital- and labor-input efficiency functions (Eq. 22) under which the task-level aggregation yields a CES production function at the macro level. The relationship between the degree of automation and aggregate input efficiencies follows a Pareto cumulative distribution, which also delivers the highest internal consistency among automation measures tested.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Internal consistency criterion&lt;/strong&gt;: A criterion for selecting among automation measures, defined as the mean squared gap between the automation degree inferred from the capital-labor relationship and the automation degree implied by the output-per-worker relationship, within the same CES structure (Eq. 42). A smaller gap indicates that the measure is more coherent with the CES production framework from which it is derived.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital-augmenting technical progress&lt;/strong&gt;: An exogenous shift in the efficiency of capital inputs (A_K,t) that raises the effective capital-labor ratio and therefore the degree of automation at any given physical capital-labor ratio. Distinguished from labor-augmenting and neutral technical progress. In the empirical estimation, capital-augmenting technical progress is statistically significant across all specifications, while labor-augmenting technical progress cannot be confirmed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Two-level CES production function&lt;/strong&gt;: An extension of the standard CES that nests industrial robots and other capital into a capital composite at the inner level (with substitution elasticity sigma-b), then combines the composite with labor at the outer level (with elasticity sigma-a). Allows separate identification of the automation role of robots versus other capital, yielding a more pronounced increase in the degree of automation than the standard CES when robots are explicitly accounted for.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Automation frontier&lt;/strong&gt;: The marginal task at which the cost of capital use exactly equals the cost of labor use, i.e., the task a_t at which lambda(a_t)/theta(a_t) = w_t/R_t. Tasks with indices below this frontier are automated; tasks above are performed by labor. As the wage-to-rental ratio rises, the frontier expands (more tasks become automated), capturing the central mechanism by which capital deepening drives automation.&lt;/p&gt;</description></item><item><title>Train to Opportunity: the Effect of Infrastructure on Intergenerational Mobility</title><link>https://macropaperwarehouse.com/papers/train-to-opportunity-the-effect-of-infrastructure-on-intergenerational-mobility/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/train-to-opportunity-the-effect-of-infrastructure-on-intergenerational-mobility/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether proximity to transport infrastructure can sever the occupational tie between parents and children — a question with direct bearing on the debate over place-based versus people-based policies. The authors exploit the nineteenth-century expansion of the railroad network across England and Wales, a setting where the First and Second Industrial Revolutions were remaking the occupational structure at the same time that the railroad was knitting together local labor markets and enabling geographic mobility.&lt;/p&gt;
&lt;p&gt;The empirical strategy centers on a novel dataset of close to 980,848 father-son pairs constructed from the full digitized population censuses of England and Wales in 1851, 1881, and 1911 (I-CeM project). Individuals are tracked across consecutive censuses using the Abramitzky-Mill-Perez (2019) linking procedure, which achieves match rates of 43–50% for men aged 40–52. Crucially, each individual is geolocated to the street level by matching census addresses to the GB1900 gazetteer, allowing railroad access to be measured as the straight-line distance from the childhood residence to the nearest train station — a finer measure than the district-level presence indicators used in prior work. Sons&amp;rsquo; occupations are observed at ages 40–52; fathers&amp;rsquo; occupations are measured 30 years earlier when sons were aged 10–22. Occupational mobility uses two complementary scales: HISCO categories (farming, laborer, services, sales, clerical, managerial, professional) and the continuous HISCAM social-interaction-distance ranking (scores 28–99, mean 50, SD 10).&lt;/p&gt;
&lt;p&gt;The key endogeneity problem is that railroad companies targeted low-density, cheap land, and that wealthy landowners and local politicians influenced station placement. To isolate exogenous variation, the authors construct a dynamic least-cost path (DLCP) network connecting 53 major towns identified by their 1801 populations (top 10% of the population distribution, threshold 9,172 inhabitants). The DLCP assigns slope costs to 50x50 meter grid cells and finds the minimum-cost path between every town pair. Lines are ranked by betweenness centrality to separate &amp;ldquo;early&amp;rdquo; 1851 lines from &amp;ldquo;late&amp;rdquo; 1881 lines, giving a time-varying instrument. Proximity to the nearest DLCP line is used as the instrument for proximity to the nearest actual train station, with standard errors clustered at the parish level. Controls include county and census-year fixed effects, distance to the nearest 1801 major town and its population, distance to Roman roads, ancient ports, and navigable waterways, plus household characteristics (number of servants as a wealth proxy, household size, and father&amp;rsquo;s foreign birth).&lt;/p&gt;
&lt;p&gt;Main results (preferred IV specification with full controls): sons who grew up one standard deviation — approximately 5 km, or about one hour&amp;rsquo;s walk — closer to a train station were 11 percentage points more likely to work in an occupation category different from their father&amp;rsquo;s. They were 5 percentage points more likely to be upwardly mobile, defined as a son&amp;rsquo;s HISCAM score exceeding his father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s distribution. The downward mobility estimate is 3 percentage points — positive but smaller in magnitude — indicating that railroad access raises occupational churn asymmetrically, predominantly upward. First-stage F-statistics exceed the Staiger-Stock threshold comfortably (135–414 across specifications). OLS estimates are uniformly smaller than IV estimates, consistent with historical evidence that the railroad targeted areas with weaker growth trajectories.&lt;/p&gt;
&lt;p&gt;The occupational transitions underlying these results run strongly out of farming and into professional, clerical, sales, and services categories, regardless of the father&amp;rsquo;s own occupation (Table IV). Sons growing up closer to the railroad were 19 percentage points less likely to work in a declining occupation and 16 percentage points more likely to work in a growing occupation. The distributional pattern shows an inverted-U relationship with father&amp;rsquo;s occupational decile for occupation-category switching and rank divergence, with the greatest gains concentrated among sons of middle-ranking fathers. For upward mobility specifically, the benefits diminish monotonically as father&amp;rsquo;s rank rises — sons from blue-collar backgrounds gained more (upward mobility coefficient 0.064) than sons from white-collar backgrounds (0.031).&lt;/p&gt;
&lt;p&gt;The authors decompose the total railroad effect on intergenerational mobility into three channels using a structural decomposition applied to a sample of 342,715 brothers: (1) changes in local labor-market opportunities, estimated as the effect on mobility for stayers; (2) changes in the returns to spatial mobility, estimated via a within-family comparison of brothers who moved versus stayed; and (3) changes in the rate of spatial mobility itself. Better railroad access raised the probability of moving away from the birth county by 15 percentage points. However, the estimated return to spatial mobility — the extra boost from actually moving — was reduced by railroad access (negative interaction between proximity and mover status), meaning the railroad decreased the relative advantage of leaving. The decomposition (Table C.6) shows that changes in local opportunities account for the great majority of the total mobility effect. Parish-level evidence confirms the local opportunity mechanism: better-connected parishes saw population growth, more industrial chimneys, more entrepreneurs, higher shares of skilled and literate workers, higher Gini coefficients, and higher median occupational ranks — consistent with agglomeration, industrialization, and skill-biased structural change.&lt;/p&gt;
&lt;p&gt;The policy implication is that transport infrastructure investment can reduce intergenerational persistence in occupational status, primarily by restructuring the local labor market rather than by enabling workers to exit. The caveat is that these gains were unevenly distributed — middle- and lower-ranking families benefited most, and the railroad simultaneously raised local inequality alongside local mobility.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-and-what-are-the-main-threats-it-addresses"&gt;Q1. What is the core identification strategy and what are the main threats it addresses?&lt;/h3&gt;
&lt;p&gt;The authors use a &amp;lsquo;dynamic least-cost path&amp;rsquo; (DLCP) instrument. They connect 53 major English and Welsh towns (defined as the top 10% of the 1801 population distribution, with at least 9,172 inhabitants) via least-cost routes computed over a 50×50 meter terrain grid that assigns slope-based costs to each cell. The instrument is proximity from the childhood residence to the nearest line in this DLCP network. The logic is that individuals incidentally located near the geographic route between major historical towns are more likely to be near an actual railroad — but the DLCP route is based purely on terrain costs, not on local demand, local resources, or the political lobbying that shaped where stations were actually placed. The strategy addresses: (a) reverse causality from high-growth areas attracting railroad placement; (b) sorting of ambitious or wealthy households toward connected parishes; (c) railroad companies&amp;rsquo; demand-driven routing choices. The exclusion restriction could be violated if location along least-cost paths between 1801 major towns is directly correlated with intergenerational mobility for reasons other than the railroad. The paper addresses this by controlling for distance to the nearest 1801 major town and its population (proximity to nodes), proximity to Roman roads, ancient ports, and navigable waterways (pre-existing trade routes), and household wealth proxies.&lt;/p&gt;
&lt;h3 id="q2-how-is-the-instrument-made-dynamic-and-why-does-this-matter"&gt;Q2. How is the instrument made dynamic, and why does this matter?&lt;/h3&gt;
&lt;p&gt;The authors divide the hypothetical network into &amp;rsquo;early&amp;rsquo; (1851) and &amp;rsquo;late&amp;rsquo; (1881) lines by ranking lines in decreasing order of betweenness centrality — the number of times a line connects major towns via shortest paths — until the total cost of the 1851 observed network is exhausted. This dynamic structure means the instrument varies across both space and census cohorts (sons measured in 1851-1881 versus 1881-1911). Without the dynamic feature, the instrument could conflate the effects of lines that were built early (and thus had decades to affect local economies) with lines built later. The temporal variation bolsters the plausibility of the exclusion restriction and is shown to be robust in alternative specifications using static least-cost paths and slope-free least-cost paths.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-four-dependent-variables-and-how-is-intergenerational-mobility-defined"&gt;Q3. What are the four dependent variables and how is intergenerational mobility defined?&lt;/h3&gt;
&lt;p&gt;The paper uses four measures: (1) an indicator equal to one if the son works in a different HISCO occupation category than his father; (2) the absolute value of the difference in HISCAM scores between son and father; (3) &amp;lsquo;upward mobility,&amp;rsquo; an indicator equal to one if the son&amp;rsquo;s HISCAM score exceeds his father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s score distribution; (4) &amp;lsquo;downward mobility,&amp;rsquo; the symmetric indicator for a decline greater than one standard deviation. Sons&amp;rsquo; occupations are observed when sons are 40–52 years old; fathers&amp;rsquo; occupations are measured 30 years earlier when sons were 10–22. The HISCAM scale is held constant over the period (national GB scale, 1800–1938) so that rankings reflect fixed social stratification positions rather than period-specific prestige. The paper also uses time-varying HISCAM, HISCLASS, Woollard, and Armstrong classifications as robustness checks.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-first-stage-performance-of-the-instrument"&gt;Q4. What is the first-stage performance of the instrument?&lt;/h3&gt;
&lt;p&gt;The first-stage relationship between proximity to the nearest DLCP line and proximity to the nearest actual train station is positive and statistically significant across all specifications. The Sanderson-Windmeijer F-statistic is 414 in the specification without controls, 136 with county and year fixed effects and full controls, and remains well above the conventional threshold of 10. The first-stage coefficient drops from 0.640 to 0.339 when full controls are added, indicating that a portion of the geographic correlation between the DLCP and the actual network reflects the pre-existing economic importance of towns and travel routes — which is precisely what the controls absorb.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q5. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The paper decomposes the total IV effect on intergenerational mobility using a three-part decomposition: (1) Changes in local opportunities, measured as the effect of proximity on mobility for sons who stayed in their birth county (stayers); (2) Changes in the returns to spatial mobility, estimated by comparing brothers who moved with brothers who stayed (using family fixed effects), and interacting this comparison with railroad proximity; (3) Changes in the rate of spatial mobility itself, estimated from the effect of proximity on the probability of county-to-county migration. Table C.6 shows that local opportunities account for the dominant share of the total effect. The railroad raised the migration probability by 15 percentage points (Table VI), so spatial mobility channels exist — but the railroad decreased the relative advantage of actually moving (negative interaction term in Table V), meaning the local opportunity channel more than offsets the spatial channel. Supporting evidence from parish-level regressions (Table VII) shows that better-connected parishes experienced significantly higher population growth, more industrial chimneys, more entrepreneurs per 100 square meters, higher shares of skilled and literate workers, higher Gini coefficients, and higher median occupational ranks — consistent with agglomeration and skill-biased industrialization.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-is-documented-by-fathers-occupation-and-position-in-the-distribution"&gt;Q6. What heterogeneity is documented by father&amp;rsquo;s occupation and position in the distribution?&lt;/h3&gt;
&lt;p&gt;The effects are heterogeneous by the father&amp;rsquo;s occupational position. Figure 6 shows an inverted-U pattern for occupation-category switching and absolute rank divergence: sons of middle-ranking fathers benefit most from railroad access. For upward mobility (Figure 6c), the benefits diminish monotonically from the lower end of the father&amp;rsquo;s distribution — sons of low-ranking fathers are most likely to move up. Sons of white-collar fathers see smaller (and sometimes statistically insignificant) upward mobility gains (0.031) compared with sons of blue-collar fathers (0.064), while the occupation-category switching benefit is also larger for blue-collar sons (0.108 vs. 0.057) (Table C.1). Separate transition matrices by HISCO category (Table IV) show that railroad access reduces the probability of farming for sons of all father types, and raises probabilities of clerical, sales, and services occupations. Effects on becoming a laborer are heterogeneous: for sons of farmers, proximity raises the probability of becoming a laborer; for sons in service occupations, it decreases it.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-run"&gt;Q7. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper performs an extensive battery. (1) Alternative connectivity measures: distance to the nearest railroad line, indicator variables for train station within 5, 10, and 15 km, and parish-level station presence. (2) Alternative mobility thresholds: 0.5, 1.5, and 2 standard deviations for upward and downward mobility; time-varying HISCAM to account for changing occupational prestige. (3) Removing railroad-specific occupations (train conductors, controllers) to check for mechanical effects. (4) Alternative specifications: second-order polynomials, parish fixed effects (10,419 parishes), and fully nonparametric covariate controls via k-means clustering (500 clusters). (5) Alternative instruments: a slope-free DLCP and a static (non-dynamic) least-cost path. (6) Geolocation robustness: using parish centroids instead of street-level addresses. (7) Linking bias: controlling for the individual probability of being linked using cubic polynomials on linkage probability and surname-frequency dummies; also checking that the railroad network explains little of the share of linked individuals at the parish level. (8) Subsamples: by census year (1851-1881 vs. 1881-1911), by county (leave-one-out), by rural/urban status, by father&amp;rsquo;s age, by son&amp;rsquo;s age, by birth order, by native/first-/second-generation immigrant status, by whether the son was born in the same county he grew up in, and by whether the father was in farming. (9) Causal response weighting: the Loken-Mogstad-Wiswall decomposition shows positive IV weights across the entire proximity distribution, consistent with a LATE interpretation. Results are stable across all checks.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-handle-the-selection-into-migration-problem-in-estimating-returns-to-spatial-mobility"&gt;Q8. How does the paper handle the selection-into-migration problem in estimating returns to spatial mobility?&lt;/h3&gt;
&lt;p&gt;The authors follow Abramitzky, Boustan, and Eriksson (2012) and use a within-family comparison of brothers — a subsample of 342,715 sons from 157,369 households who grew up in the same household but one moved county while the other stayed. Family fixed effects absorb the shared household characteristics (wealth, motivation, family networks, financial constraints) that jointly determine the propensity to migrate and the baseline mobility trajectory. The railroad-proximity interaction with mover status is instrumented using the interaction of the DLCP instrument with the mover indicator, via a control function approach. The estimated baseline return to spatial mobility (the mover premium) is positive and significant — movers have higher occupation-category divergence and shift more in both directions — but the railroad-induced change in return to mobility is negative, meaning that proximity to the railroad reduced the additional mobility benefit of actually migrating. This finding is the core of the conclusion that local opportunities, not spatial mobility, dominate.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-paper-document-about-local-labor-market-changes-induced-by-the-railroad"&gt;Q9. What does the paper document about local labor market changes induced by the railroad?&lt;/h3&gt;
&lt;p&gt;Parish-level IV regressions (Table VII) show that better proximity to the 1851 network (instrumented by the DLCP) is associated with: significantly higher population growth between 1851 and 1881; a significantly larger number of industrial chimneys (proxying factory concentration, sourced from Heblich-Trew-Zylberberg (2021)); more entrepreneurs per 100 square meters (from the British Business Census of Entrepreneurs); higher shares of high-skilled and literate workers; a higher Gini coefficient over occupational ranks; and a higher median occupational rank. Additionally, sons in better-connected parishes were 19 percentage points less likely to work in a declining occupation and 16 percentage points more likely to work in a growing occupation (Table C.3). Sons were also 3 percentage points more likely to be literate and 7 percentage points more likely to work in a non-manual occupation (Table C.5). These findings collectively point to agglomeration, industrialization, skill-biased technological change, and the creation of a new entrepreneur class as the mechanisms by which the railroad transformed local labor market structure.&lt;/p&gt;
&lt;h3 id="q10-what-prior-work-does-this-paper-relate-to-most-closely-and-what-distinguishes-it"&gt;Q10. What prior work does this paper relate to most closely, and what distinguishes it?&lt;/h3&gt;
&lt;p&gt;The paper sits at the intersection of the railroad-infrastructure and intergenerational-mobility literatures. In the infrastructure tradition, it relates closely to Donaldson (2018, AER) on railroads in India, Donaldson and Hornbeck (2016, QJE) on US market access, Bogart et al. (2022, JUE) on population and structural change in England and Wales, and Heblich-Redding-Sturm (2020, QJE) on London commuting and urban growth. The closest prior paper is Perez (2017) on nineteenth-century Argentina, who finds railroad access shifted children from agricultural into white-collar and skilled blue-collar occupations; this paper provides similar evidence for England and Wales at individual level and adds a full mechanism decomposition. In the intergenerational mobility tradition it relates to Long and Ferrie (2013, AER) and Long (2013, ERH) on census-based occupational mobility in Victorian Britain. The key methodological advantages of the current paper are: (a) use of the full (not 2%) census for all three years, yielding close to 1 million father-son pairs with match rates of 43–50% versus 15–33% in prior work; (b) street-level geolocation enabling individual-level rather than district-level measurement of railroad access; (c) the explicit three-way mechanism decomposition separating local opportunities, returns to migration, and migration rates; and (d) documenting rich heterogeneity by father&amp;rsquo;s occupational rank and occupation category.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-what-scope-conditions-limit-their-external-validity"&gt;Q11. What are the policy implications and what scope conditions limit their external validity?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s core policy message is that transport infrastructure investment can be an effective mechanism for reducing intergenerational occupational persistence — primarily by creating new local labor market opportunities rather than by enabling low-income workers to reach distant job centers. This provides historical support for place-based policies of the sort embodied in the Biden &amp;lsquo;Build Back Better&amp;rsquo; infrastructure proposals or the UK HS2 high-speed railway project (mentioned in the paper). The main scope conditions limiting generalizability are: (1) The setting is nineteenth-century England and Wales during the Industrial Revolution, when the occupational structure was shifting rapidly from farming to industry and commerce — the railroads arrived at a moment of latent demand for new labor market structures; (2) The benefits were not evenly distributed: middle-ranking families (by father&amp;rsquo;s occupational rank) gained most in absolute occupational switching and rank divergence, while the lowest-ranked families gained most specifically in upward mobility; (3) The railroad simultaneously raised local inequality alongside local mobility, suggesting infrastructure investment can be inequality-increasing in the cross-sectional distribution of wages even as it reduces intergenerational persistence; (4) The effects are highly localized — even 5 km of additional distance matters — implying that the placement of stations relative to where low-income families actually live is crucial for achieving distributional goals.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-paper-document-about-the-baseline-patterns-of-intergenerational-mobility-in-the-sample"&gt;Q12. What does the paper document about the baseline patterns of intergenerational mobility in the sample?&lt;/h3&gt;
&lt;p&gt;In the full sample of 980,848 father-son pairs covering 1851-1881 and 1881-1911, 80% of sons do not remain in the same HISCO occupation category as their father. The correlation between father&amp;rsquo;s and son&amp;rsquo;s HISCAM ranks is 0.28. Among sons, 18% experienced upward mobility (son&amp;rsquo;s HISCAM rank more than one SD higher than father&amp;rsquo;s) and 15% experienced downward mobility (more than one SD lower). About 31% of sons moved to a different county from where they grew up, settling on average 100 km away. Sons grew up on average 3.28 km from the nearest train station (SD 5.45 km). These descriptives reveal strong spatial clustering in intergenerational mobility patterns at the parish level.&lt;/p&gt;
&lt;h3 id="q13-does-the-late-interpretation-hold-and-what-does-the-weighting-function-show"&gt;Q13. Does the LATE interpretation hold and what does the weighting function show?&lt;/h3&gt;
&lt;p&gt;The authors verify the LATE interpretation via two approaches. First, following Loken-Mogstad-Wiswall (2012), they compute the causal response weighting function as the covariance between each discrete proximity indicator and the DLCP instrument, divided by the covariance between the proximity measure and the DLCP instrument. They find positive weights across the entire distribution of proximity to the nearest train station, concentrated most heavily for individuals residing 0.5 to 1.5 proximity units (approximately 2.7 to 8.1 km) from a train station — these are the individuals whose proximity is most affected by incidental location along the DLCP. The absence of negative weights indicates the IV estimate does not mix complier and never/always-taker effects in a sign-reversing way. Second, following Blandhol et al. (2022), a fully nonparametric specification using 500 k-means clusters for covariates yields estimates very close to the parametric baseline, consistent with a LATE interpretation of the linear IV estimator.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Dynamic Least-Cost Path (DLCP) Network&lt;/strong&gt;: The paper&amp;rsquo;s instrument for railroad access. A hypothetical railroad network connecting England and Wales&amp;rsquo;s 53 largest towns in 1801 via routes that minimize geographic cost (distance plus slope-based terrain costs), ignoring all demand-side factors. Lines are classified as &amp;rsquo;early&amp;rsquo; (1851) or &amp;rsquo;late&amp;rsquo; (1881) by betweenness centrality until the cost budget of the actual 1851 network is exhausted. Proximity from childhood residence to the nearest DLCP line instruments proximity to the nearest actual train station.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intergenerational Occupational Mobility&lt;/strong&gt;: In this paper, the degree to which a son&amp;rsquo;s adult occupation differs from his father&amp;rsquo;s, measured both categorically (same versus different HISCO category) and cardinally (difference in HISCAM scores). Upward (downward) mobility is specifically defined as the son&amp;rsquo;s HISCAM score exceeding (falling below) the father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s HISCAM distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HISCAM Score&lt;/strong&gt;: A continuous occupational ranking (range 28–99, mean 50, SD 10) derived from the frequency of social interactions — marriages, friendships, parent-child links — between occupations in historical data. Higher scores indicate a more advantageous position in the social stratification structure. The paper uses the national Great Britain scale, held constant for 1800–1938, to make rankings comparable across census years.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Local Opportunities Channel&lt;/strong&gt;: The mechanism by which railroad access improved intergenerational mobility through restructuring the local labor market — enabling commuting, attracting factories and entrepreneurs, spurring urbanization and industrialization, and creating new occupations requiring new skills — without requiring sons to migrate away from their birth county. Identified empirically as the effect of railroad proximity on mobility outcomes for sons who stayed in their birth county (stayers).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Returns to Spatial Mobility&lt;/strong&gt;: The additional intergenerational mobility benefit (or penalty) associated with actually migrating to another county, estimated using within-family variation among brothers — one who moved and one who stayed — to net out shared household-level determinants of mobility. The paper finds that railroad access reduced (made more negative) the returns to spatial mobility, meaning that the relative advantage of leaving shrank as local opportunities expanded.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inconsequential Place IV Approach&lt;/strong&gt;: An identification strategy (following Chandra-Thompson 2000 and Michaels 2008) in which the instrument for infrastructure access is constructed from the geographic convenience of locations lying between endpoints of a planned network, rather than from demand-side factors at those locations. The DLCP instrument in this paper is a specific implementation: individuals living between 1801 major towns incidentally receive railroad access because the low-cost route between towns passes near their residence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Occupational Tie (Father-Son)&lt;/strong&gt;: The tendency for sons to remain in the same occupation category or same position in the occupational ranking as their father. In this paper, severing the occupational tie means a son moves to a different HISCO category and/or achieves a HISCAM score meaningfully different from his father&amp;rsquo;s. The railroad&amp;rsquo;s main effect is framed as reducing this tie, with upward mobility being the dominant direction of change.&lt;/p&gt;</description></item><item><title>Universal Daycare and Mothers' Working Lifetime</title><link>https://macropaperwarehouse.com/papers/universal-daycare-and-mothers-working-lifetime/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/universal-daycare-and-mothers-working-lifetime/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper estimates the causal effects of universal daycare access on mothers&amp;rsquo; labor force participation, full-time employment, hours worked, and earnings across 34 years after the birth of their first child — the longest window examined in this literature. The motivation is twofold: the existing evidence base is overwhelmingly short-run, and the human capital channel (reduced depreciation of skills, accumulation of experience) implies that early labor market attachment during child-rearing years could compound over decades in ways that short-run estimates miss entirely.&lt;/p&gt;
&lt;p&gt;The identification exploits Denmark&amp;rsquo;s 1964 reform that converted a targeted (means-tested) childcare system into a universal one, which triggered a staggered geographic roll-out of daycare centers from 1966 onward across the country&amp;rsquo;s 2,033 neighborhoods nested in 277 municipalities. The paper combines digitized historical daycare yearbooks (1964–1975), the 1970 census, and administrative registers from Statistics Denmark covering 370,602 mothers who had their first child between 1964 and 1975. Employment is measured via annual contributions to the Supplementary Pension Fund (ATP); earnings from tax records are available from 1980 through 2015, adjusted to 2016 USD. The empirical strategy is a difference-in-differences design comparing mothers in neighborhoods with versus without daycare within the same municipality over time. Daycare availability when the first-born child turns four is used as the fixed treatment indicator for the long-run regressions. Municipality fixed effects absorb cross-sectional confounders; year-of-first-birth dummies capture macro trends.&lt;/p&gt;
&lt;p&gt;The contemporaneous effects are already substantial. Once year and municipality fixed effects and covariates are included, daycare availability raises the probability of participation by 1.5 percentage points when the child is two, rising to 5.3–5.7 percentage points for years three through six — translating to roughly 9 percent more likely to participate relative to the mean. Full-time employment rises by 9–12 percent relative to the mean for years three through six; hours worked increase by 0.27 hours per week (1.8 percent) when the child is four.&lt;/p&gt;
&lt;p&gt;The long-run effects persist throughout the entire working life. Relative to the sample mean, mothers with daycare access are 9.7 percent more likely to participate when the first child turns four, declining to 5.7 percent at child age 14, 3.1 percent at child age 22, and still 1.2 percent at child age 34 (when the average mother is approximately 57.7 years old). Full-time employment effects follow a parallel trajectory: 11 percent higher at child age four, 8.2 percent at child age 14, and 4.4 percent at child age 34. Log earnings (conditional on employment) range between 3 and 6 percent higher throughout the observation window; mothers earn 5.3 percent more when the child is 16 and 4.2 percent more when the child is 34.&lt;/p&gt;
&lt;p&gt;Heterogeneity by education is a central finding. For low-educated mothers (no post-secondary education, 50 percent of the sample), participation effects are 10.1 percent at child age 10, 5.1 percent at child age 17, and remain statistically significant through 32 years. For higher-educated mothers, participation effects are 3.9 percent at child age 10, fall below 1 percent by child age 17, and become statistically indistinguishable from zero by child age 23. Employment effects are thus larger and more persistent for low-educated mothers. Earnings effects, however, are more closely aligned across education groups and show a distinctive pattern for higher-educated mothers: earnings effects persist and remain significant long after employment effects have faded, suggesting that sustained attachment during child-rearing years translates into qualitative career advancement (not just more years worked) for the more educated group.&lt;/p&gt;
&lt;p&gt;Potential mediators include reduced secondary fertility and increased parental separation. Daycare for children aged three to six reduces the total number of children by 0.036 (1.6 percent relative to the mean of 2.2), reduces the probability of having more than two children by 1.8 percentage points (6.0 percent), and increases birth spacing by 0.137 years, making mothers 2.2 percentage points less likely to have a second child within two years. Additionally, mothers with daycare access are 2 percentage points more likely to live apart from the first-born child&amp;rsquo;s father when that child turns 16 — consistent with greater female economic independence. These mediator effects do not vary systematically by education level. Daycare access does not affect additional educational attainment after first birth, ruling out re-skilling as a channel.&lt;/p&gt;
&lt;p&gt;The policy implication is that subsidized universal daycare is not merely a short-run labor supply intervention but a persistent investment in female human capital accumulation, with effects that compound over careers and remain economically meaningful into near-retirement ages.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-key-threats-to-it"&gt;Q1. What is the identification strategy and what are the key threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses a staggered difference-in-differences design. The key variation is the timing of daycare center openings across neighborhoods within municipalities following the 1964/1966 Danish reform. Daycare availability in the year the first-born child turns four is the fixed treatment indicator for long-run regressions; current-year daycare availability is used for contemporaneous regressions. Municipality fixed effects absorb time-invariant local differences; year-of-first-birth dummies absorb aggregate time trends. The main threat is non-random placement of daycare centers — if centers opened in areas where female labor force participation was already rising, the estimates would be upward biased. The paper addresses this with (1) an event study at the neighborhood level using data from 1960 through 2003 showing no pre-reform differential trends between neighborhoods that later received daycare and those that did not (compared against placebo neighborhoods assigned fictitious opening dates mimicking the actual distribution), and (2) a selective migration check showing that mothers who moved longer distances from their birthplace were no more likely to reside in a neighborhood with daycare once the full conditioning set is included. A residual concern is that for mothers having their first child before 1970, neighborhood assignment is measured post-birth (1970 census), which is addressed by a robustness check excluding the pre-1970 first-birth cohort.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-deal-with-heterogeneous-treatment-effects-and-two-way-fixed-effects-bias"&gt;Q2. How does the paper deal with heterogeneous treatment effects and two-way fixed effects bias?&lt;/h3&gt;
&lt;p&gt;The paper acknowledges the recent literature on TWFE bias under treatment effect heterogeneity (De Chaisemartin and d&amp;rsquo;Haultfoeuille 2020; Callaway and Sant&amp;rsquo;Anna 2021; Sun and Abraham 2021; Borusyak et al. 2024). It replicates the pre-reform event study using the Borusyak et al. (2024) imputation estimator, which is robust to heterogeneous treatment effects and allows for covariates, and finds similar results to the standard TWFE event study (Appendix Figure A.2). The main long-run regressions fix the treatment indicator to daycare availability when the child is four, so there is no variation in treatment timing within a regression, limiting but not eliminating TWFE concerns for the long-run estimates.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-main-mechanism-behind-the-persistent-effects"&gt;Q3. What is the main mechanism behind the persistent effects?&lt;/h3&gt;
&lt;p&gt;The paper attributes the persistence to human capital dynamics: labor force participation during the child-rearing years reduces depreciation of previously accumulated human capital (from education and prior work experience) and enables new on-the-job human capital accumulation through the current job. For low-educated mothers, the primary channel appears to be the extensive margin — daycare moves mothers who would otherwise become homemakers into paid employment, and the employment effects persist because once labor market attachment is established, it is durable. For higher-educated mothers, the earnings-employment gap is the key signal: employment effects fade within roughly 23 years (consistent with convergence once children are no longer preschool age and informal care becomes feasible), yet earnings remain elevated for decades, suggesting that the women who maintained employment during child-rearing years accrued qualitatively better positions — more experience, better job-match, more promotions — compared to those who did not.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-main-mediators-and-how-are-they-distinguished"&gt;Q4. What are the main mediators and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;Three mediators are examined. First, secondary fertility: daycare for children aged 3–6 reduces number of children by 0.036, probability of a third child by 1.8 percentage points, and probability of a fourth child by 0.5 percentage points. The effect operates through daycare for children 3–6 (not 0–2), consistent with the main employment effects operating when the child is three or older. The fertility reduction increases the opportunity cost interpretation — daycare raises the effective wage, making additional children more costly in terms of foregone earnings. Second, birth spacing: mothers with daycare access wait 0.137 more years between first and second child, and are 2.2 percentage points less likely to have the second child within two years, allowing longer uninterrupted work spells. Third, parental separation: mothers with daycare access are 2 percentage points more likely to live apart from the child&amp;rsquo;s father at child age 16, consistent with greater economic independence from labor market participation reducing barriers to separation. Additional educational attainment after first birth is tested and found to be an insignificant channel (no significant effect overall, a marginal effect only for low-educated mothers), ruling out re-skilling as a mediator.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented-beyond-the-education-split"&gt;Q5. What heterogeneity is documented beyond the education split?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s primary heterogeneity analysis is by maternal education level (low: no post-secondary education versus higher: any post-secondary education including vocational training, college, or university). The education split produces the most substantive finding: employment effects are larger and more persistent for low-educated mothers, while the earnings-employment divergence is the distinctive feature for higher-educated mothers. No other dimensions of heterogeneity (by birth cohort, by municipality type beyond the urban indicator, by parity) are formally reported in the main results, though geographic robustness checks (exclusion of three largest cities, exclusion of suburbs) implicitly test whether effects are concentrated in particular settings and find they are not.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Four main sets of robustness checks are reported. First, selective migration: regressions of daycare availability on distance moved from birthplace (linear, quadratic, and IHST-transformed) with the full conditioning set show no significant relationship, ruling out systematic sorting into daycare neighborhoods. Second, pre-1970 cohort exclusion: restricting to mothers with first birth after 1970 (for whom the 1970 census address is predetermined relative to birth) yields qualitatively similar results, though participation effect sizes are somewhat smaller. Third, urban geography: excluding the three largest municipalities (Copenhagen, Frederiksberg, Aarhus, Odense) and separately excluding suburbs of Copenhagen and Aarhus both leave the main results intact. Fourth, differential time trends: allowing the most populous neighborhood within each municipality to have its own set of time dummies (to capture potentially faster urban trend evolution) does not change the finding that participation and earnings effects persist beyond 30 years. The paper also shows that results are robust to an alternative participation definition based solely on ATP contributions for all years (versus mixing ATP pre-1980 and earnings post-1980).&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-prior-work-and-what-is-its-main-contribution"&gt;Q7. How does this paper relate to prior work and what is its main contribution?&lt;/h3&gt;
&lt;p&gt;The prior literature falls into two camps. The short-run camp (Havnes and Mogstad 2011 for Norway; Carta and Rizzica 2018 for Italy; Bettendorf et al. 2015 for Netherlands; Cascio 2009 and Fitzpatrick 2012 for the US) documents modest to moderate employment effects during the preschool years. The medium-run camp (Lefebvre et al. 2009 and Haeck et al. 2015 for Quebec; Nollenberger and Rodriguez-Planas 2015 for Spain; Herbst 2017 for the US Lanham Act) tracks effects up to about 11–17 years. This paper&amp;rsquo;s first contribution is extending the window to 34 years — covering the majority of the working life — using Danish administrative data that allow continuous observation rather than decennial census snapshots. The second contribution is documenting the earnings-employment divergence for higher-educated mothers specifically, which was not visible in shorter windows. The third contribution is the simultaneous analysis of fertility, spacing, and parental separation as mediators using the same administrative data and identification strategy, rather than treating these as separate exercises in different papers.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-scope-conditions-and-policy-implications"&gt;Q8. What are the scope conditions and policy implications?&lt;/h3&gt;
&lt;p&gt;Several scope conditions qualify the policy implications. First, the context is a universal reform in a Nordic welfare state with strong labor market institutions and universal access; the results may not directly generalize to settings with low baseline female employment or weak formal sector employment. Second, the relevant margin for the 1960s–70s cohorts was daycare for children aged three to six; the paper notes that by recent decades the relevant margin has shifted to children under two (consistent with Simonsen 2010 finding effects for younger children in 2001 data), possibly reflecting changing cultural norms or the fact that 1960s–70s mothers had multiple children before returning to work. Third, the employment effects are larger for low-educated mothers, so the labor market attachment argument applies most forcefully to this group. Fourth, the negative fertility effects mean that the total welfare calculation must weigh labor market gains against reductions in desired family size. The policy implication the paper emphasizes is that universal daycare is an investment in long-run economic output, not merely a short-run participation subsidy, because the labor market attachment it induces during child-rearing years compounds over careers through human capital accumulation.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-sample-and-data-structure"&gt;Q9. What is the sample and data structure?&lt;/h3&gt;
&lt;p&gt;The sample consists of 370,602 mothers who had their first child between 1964 and 1975 and were resident in Denmark in 1970 (from the census), after excluding women with immigrant backgrounds (2.2 percent) and those who died or emigrated before the first child turned 16 (0.6 percent). Employment is observed from the birth of the first child through 34 years after (1964–2009 approximately); earnings from 1980 through 2015. The daycare panel is constructed from historical yearbooks (1964–1975) and administrative registers (1976–1993) and provides yearly neighborhood-level data on daycare availability. The average mother in the sample was born in 1945, was 23.7 years old at first birth, had 10.8 years of education, and had 2.2 children total. The sample is split roughly 50/50 between low-educated and higher-educated mothers.&lt;/p&gt;
&lt;h3 id="q10-why-do-effects-appear-only-when-the-child-is-three-not-earlier"&gt;Q10. Why do effects appear only when the child is three, not earlier?&lt;/h3&gt;
&lt;p&gt;The paper finds that contemporary participation effects are small and statistically insignificant for years zero through two, then jump sharply at year three. The paper attributes this to two factors: (1) the universal daycare reform primarily expanded slots for children aged three to six, with nurseries for children under three expanding much more slowly through the 1980s and 1990s (Figure A.1 in the paper); and (2) cultural norms and the multi-child fertility pattern of this cohort — mothers in the 1960s–70s were more likely to have multiple children before returning to work, implying that the eldest child often reached age three or four before the mother re-entered employment. This contrasts with more recent periods (Simonsen 2010 uses 2001 data) where the relevant margin has shifted to children under two.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Universal daycare&lt;/strong&gt;: In the paper&amp;rsquo;s sense, daycare centers open to children from all socioeconomic backgrounds (not means-tested), with building costs fully publicly funded and operating costs split among state, municipality, and parents (with parents paying 30 percent), following the 1964 Danish reform. Contrasted with the pre-reform &amp;rsquo;targeted&amp;rsquo; system that only subsidized institutions where two-thirds of children came from low-income families.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Working lifetime effects&lt;/strong&gt;: The paper&amp;rsquo;s central object of analysis: the causal impact of early daycare access on maternal labor outcomes measured annually across 34 years after the birth of the first child, covering the majority of the working life. Distinguished from short-run (0–7 year) and medium-run (up to 11–17 year) effects documented in prior work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor market attachment&lt;/strong&gt;: As used in the paper, the sustained connection to paid employment during the child-rearing years (when children are of preschool age). The paper argues that attachment during this period is the mechanism for long-run effects because it reduces human capital depreciation and enables on-the-job accumulation of experience and job-specific skills.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;ATP (Supplementary Pension Fund) contributions&lt;/strong&gt;: The paper&amp;rsquo;s primary employment measure for years before 1980. Annual ATP contributions are proportional to hours worked: one-third contribution corresponds to 10–19 hours/week, two-thirds to 20–29 hours/week, and full contribution to 30 or more hours/week. Used to construct both a participation dummy and a full-time employment dummy (full ATP contribution = at least 30 hours/week). Crucially, the unemployed, self-employed, and those outside the labor force made no ATP contributions during this period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Human capital depreciation channel&lt;/strong&gt;: The mechanism by which absence from the labor market during child-rearing years erodes previously accumulated skills (from education and prior work). The paper uses this concept, following Adda et al. (2017) and Lefebvre et al. (2009), to explain why participation effects on earnings can persist long after direct employment effects have diminished: mothers who worked during preschool years entered subsequent career phases with a larger, less-depreciated human capital stock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Secondary fertility decisions&lt;/strong&gt;: The paper&amp;rsquo;s term for fertility choices conditional on already having a first child, i.e., the decision to have additional children. Examined on the intensive margin (number of additional children, spacing between births) rather than extensive margin (whether to have any children), because the sample consists entirely of women who already have at least one child.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Daycare for 3–6 year olds vs. 0–2 year olds&lt;/strong&gt;: The paper distinguishes between two types of daycare that expanded at different speeds: daycare for children aged 3–6 expanded rapidly from 1966, while nurseries for children under 3 (crèches) expanded only from the 1980s–1990s. All significant effects in the paper — on employment, fertility, and parental separation — load onto access to daycare for children aged 3–6, not 0–2, consistent with the historical timing of the expansion.&lt;/p&gt;</description></item><item><title>University Research and the Market for Higher Education</title><link>https://macropaperwarehouse.com/papers/university-research-and-the-market-for-higher-education/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/university-research-and-the-market-for-higher-education/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper proposes that university R&amp;amp;D is determined endogenously by competition for tuition and talented students in the market for higher education, and asks why universities fund research internally with tuition despite negligible returns to patenting. Motivation: between 2000 and 2018 U.S. universities accounted for 13% of aggregate R&amp;amp;D spending and 53% of all basic-research spending, yet in 2018 over 25% of university research was internally funded (25.54% in 2018; federal government 52.97%) while between 1991 and 2018 the median university earned patent licensing revenue totaling less than 2% of its R&amp;amp;D expenditure. Internal funds therefore come essentially from tuition.&lt;/p&gt;
&lt;p&gt;Approach: (1) four stylized facts from administrative microdata (IPEDS, NSF HERD survey covering 916 universities / 99.1% of sector R&amp;amp;D, AUTM patent-licensing survey, Web of Science / Leiden bibliometrics); (2) a causal natural experiment; (3) a general-equilibrium model of the higher-education sector with heterogeneous universities choosing teaching and research, calibrated to U.S. data; and (4) policy counterfactuals.&lt;/p&gt;
&lt;p&gt;Causal evidence: the authors exploit the 1998-2003 doubling of the NIH budget (from $13.6bn to $27.1bn) using a Bartik shift-share instrument built from each university&amp;rsquo;s pre-period (1993-1997) share of federal life-science grants, regressing the change in net tuition (1993-1997 to 2004-2008) on the instrumented change in R&amp;amp;D per student, with state-clustered standard errors and state-specific trends. The benchmark estimate is that a $1.00 increase in R&amp;amp;D spending per student raises tuition by $0.15 (s.e. 0.05) — universities recoup up to 15% of R&amp;amp;D through higher tuition. Across specifications the effect ranges $0.10-$0.15; it is driven by research universities (non-liberal-arts), is statistically insignificant for liberal arts colleges, and a placebo using student-amenities spending shows no significant effect. The point estimate is about 60% larger at private non-profits than publics, but that difference is not statistically significant.&lt;/p&gt;
&lt;p&gt;Model and mechanism: education quality q = k^ωk * z̄^ωz * eT^ωe depends on intangible knowledge capital k (accumulated via research, k&amp;rsquo; = k^γk * eR^γe), peer ability z̄, and teaching spending. Universities maximize discounted education quality, funding research from tuition. Equilibrium features an endogenous college hierarchy with two-dimensional sorting by ability and family income. The research share sR rises with the steepness of the college quality-ladder Σq/Σk; when students are highly stratified or tuition rises sharply with rank, universities invest in research even if the direct contribution to teaching (ωk) is small — research persists even as ωk→0 (acting as a pure signal). Incentives fall when intangible capital is highly dispersed across colleges.&lt;/p&gt;
&lt;p&gt;Calibration matches the joint distribution of research, tuition, and student ability, plus untargeted R&amp;amp;D dispersion; simulated NIH expansion yields $0.18 per $1 in steady state and $0.11 along the transition, bracketing the empirical $0.10-$0.15.&lt;/p&gt;
&lt;p&gt;Policy findings (long-run, vs baseline): removing all need-based federal tuition subsidies cuts university research by 8.1% (replacing progressive with revenue-neutral flat tuition subsidy: -2.2%); progressive aid compresses revenue dispersion, steepens the quality-ladder, and raises the research share (+0.8 pp). Removing all federal research grants cuts research by 69.1% — only 6.9 pp below the government&amp;rsquo;s 76% funding share, implying crowding-out: the meritocratic grant structure concentrates funds at top schools, flattening the ladder and cutting the research share by 16.4 pp. A revenue-neutral flat research subsidy would instead raise research by 14.8%, human capital by 9.6%, and output by 11.1%.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;A Bartik/shift-share IV exploiting the 1998-2003 NIH budget doubling. Each university&amp;rsquo;s change in R&amp;amp;D is instrumented by its pre-period (1993-1997) share of all federal life-science research grants. Relevance: NIH was the bulk of federal life-science funding before the shock and did not substantially change award criteria, so high-share schools received mechanically larger funding increases. Exogeneity requires that universities did not systematically invest in life-science research in the pre-period in anticipation of the expansion. The estimation is in long-differences comparing steady states; standard errors are clustered at the state level with state-specific tuition trends. Threats: the NIH expansion occurs at a common point in time, so it may correlate with other contemporaneous market changes; initially larger or higher-quality research universities might have raised tuition for reasons unrelated to R&amp;amp;D. The authors address this with group-specific time trends (public/private, pre-existing life-science status, school size, initial quality via faculty-student ratio) and pre-trend controls (1987-1992 faculty-student ratio, FTE size, life-science status). A limitation the authors acknowledge: they cannot test the effect on subsequent student ability because ability proxies are only available after the intervention.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished"&gt;Q2. What are the main mechanisms and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;The college quality-ladder Σq/Σk (the cross-sectional elasticity of education quality with respect to intangible capital) is the sufficient statistic for research incentives. Equation (14) decomposes it into three channels: (i) the direct teaching contribution of research ωk; (ii) attracting better students, ωz × Σz̄/Σk; and (iii) charging higher tuition, ωe × ΣR/Σk. Channels (ii) and (iii) flow from competition for talented students and tuition and can dominate even when ωk is tiny. Empirically, Σz̄/Σk maps to the cross-sectional elasticity of student ability w.r.t. research (Figure 3) and ΣR/Σk to the elasticity of tuition w.r.t. research (Figure 4), so the calibration disciplines these channels with observable cross-sectional relationships.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;The tuition effect is concentrated in research universities (non-liberal-arts), with a larger, highly significant point estimate; for liberal arts colleges the NIH shock has no statistically significant effect on tuition (the authors caution the LAC sample is smaller — ~32% of institutions, ~24% of FTE — and more heterogeneous, so power may be insufficient). The effect appears ~60% stronger at private non-profits than publics, but the difference is not statistically significant. Across the model, top schools and bottom schools both invest less in research when intangible capital is highly dispersed (top schools face weak incentives to improve already-secure rank; bottom schools find climbing too costly).&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Empirically: adding pre-trend controls (column 3) leaves estimates intact; splitting by NLA vs LAC; and a placebo replacing R&amp;amp;D with student-services (amenities) spending, which yields no significant effect, rejecting spurious cross-category correlation. In the model: (1) the limiting case ωk→0 where research is a pure signal — the research share falls from 8.8% to 2.4% of tuition but stays strictly positive, and policy effects retain 50% (tuition-subsidy removal: -0.4 pp vs -0.8) and 66% (research-subsidy removal: +10.8 vs +16.4 pp) of their magnitude; (2) allowing some teaching expenditure to also enter intangible-capital production (γT&amp;gt;0), where the research share falls from 8.8% to 4.7% and policy effects moderate (-0.4 pp and +7.1 pp). In both, existing tuition policies still boost research and federal research grants still crowd it out.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-relate-to-and-differ-from-prior-work"&gt;Q5. How does this relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;It builds on equilibrium higher-education models — Epple, Romano &amp;amp; Sieg (2006) (quality maximization, exogenous endowment hierarchy, finite universities with market power) and Cai &amp;amp; Heathcote (2022) (competitive, constant-returns technology) — but endogenizes university R&amp;amp;D alongside teaching. A theoretical contribution is proving existence of a unique dynamic equilibrium with quality maximization and an endogenous college-quality hierarchy with a continuum of colleges; Cai &amp;amp; Heathcote argued no quality-maximization equilibrium exists when colleges are ex-ante identical (all want to be at the top), which this paper resolves via the endogenous knowledge hierarchy. It contributes to the economics of science / university-R&amp;amp;D literature by adding market-driven incentives, and to the basic-research-subsidy literature (Akcigit et al.) by showing universities have private incentives to do basic research, implying the need for government subsidy may be smaller than the standard Nelson/Arrow/Rosenberg view holds.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Two main implications. First, a novel complementarity between equity and innovation: progressive need-based tuition aid compresses revenue dispersion across colleges, makes them more similar, steepens the quality-ladder, and raises research (+8.1% relative to a no-subsidy world; flat subsidy gives only ~one-quarter of that, +2.2%). Second, current meritocratic federal research grants partially crowd out internal research and raise educational inequality by concentrating resources at top schools; removing them cuts research by 69.1% (only 6.9 pp below the 76% federal share, the gap being the crowding-out). A revenue-neutral flat research subsidy would raise research by 14.8%, human capital 9.6%, and output 11.1%, eliminating the equity-innovation trade-off because it lowers research cost without altering market structure. Scope conditions: these are long-run steady-state comparisons in a calibrated model of 4-year public and private non-profit U.S. institutions; magnitudes depend on the hard-to-measure ωk and on the research-technology specification, as the robustness exercises show.&lt;/p&gt;
&lt;h3 id="q7-why-do-universities-fund-research-from-tuition-rather-than-patents-and-does-the-model-rationalize-it"&gt;Q7. Why do universities fund research from tuition rather than patents, and does the model rationalize it?&lt;/h3&gt;
&lt;p&gt;Because patent licensing is too small (median &amp;lt;2% of R&amp;amp;D, 1991-2018) to fund the &amp;gt;25% of R&amp;amp;D that is internal, and unrestricted operating funds are composed almost entirely of tuition (much of it from unrecovered facilities-and-administration costs on sponsored projects — roughly $7bn in 2018). The model rationalizes diverting tuition to research because research raises education quality and thus students&amp;rsquo; willingness to pay, so in a competitive sector students accept it. The model also replicates the joint pattern that higher-R&amp;amp;D universities are higher-ranked, attract wealthier and abler students, and charge higher tuition.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-sources-of-inefficiency-in-the-model"&gt;Q8. What are the sources of inefficiency in the model?&lt;/h3&gt;
&lt;p&gt;Two. First, borrowing constraints prevent efficient sorting of students by ability (a social planner would send the ablest to the best colleges, but students are limited by parental capacity to pay). Second, university knowledge has positive spillovers to the real economy (calibrated ιk = 0.1) that colleges do not internalize, causing under-investment; however, quality-maximizing colleges face extra competitive incentives to do research, so net under- or over-investment is ambiguous and depends on stratification relative to spillover strength.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;College quality-ladder (Σq/Σk)&lt;/strong&gt;: The equilibrium cross-sectional elasticity of education quality with respect to a university&amp;rsquo;s intangible knowledge capital — a sufficient statistic for a university&amp;rsquo;s private incentive to invest in research. Steeper ladder (more stratification, tuition rising more with rank) means stronger research incentives.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intangible (knowledge) capital k&lt;/strong&gt;: Institution-specific intangible capital accumulated by investing in research (k&amp;rsquo; = k^γk eR^γe). It is primarily frontier knowledge and ideas exposed to students, but also networks, recruiting, labs, and methods; it can act purely as a reputation signal in the limiting case ωk→0.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Research share (sR)&lt;/strong&gt;: The share of a university&amp;rsquo;s tuition revenue allocated to research in equilibrium (≈8.8% under existing policies). It increases with college forward-lookingness (βc) and the steepness of the quality-ladder, and decreases with the dispersion of intangible capital across colleges.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crowding-out of internal research&lt;/strong&gt;: In the paper&amp;rsquo;s sense, the phenomenon whereby federal grants, by concentrating funds at top schools, raise the dispersion of research (Σk), flatten the quality-ladder (Σq/Σk), lower the research share, and thereby reduce universities&amp;rsquo; internal research spending — so total research rises less than the government&amp;rsquo;s funding share (69.1% decline vs 76% share on removal).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Equity-innovation complementarity&lt;/strong&gt;: The model&amp;rsquo;s finding that progressive need-based tuition aid, by compressing revenue dispersion and making colleges more similar, steepens competition and raises university research — so equity-promoting policy also boosts basic research, rather than trading off against it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Education-innovation gap (ωk calibration)&lt;/strong&gt;: Biasi &amp;amp; Ma&amp;rsquo;s (2021) measure of how frontier-current a university&amp;rsquo;s curriculum is, interpreted in the model as log(k). A one-unit decrease is associated with a 0.011% rise in graduate income; normalized by its school-level standard deviation of 0.85, it is used to pin down ωk via ωk·α = .011/.85·Σk.&lt;/p&gt;</description></item><item><title>Heterogeneity in Manufacturing Growth Risk</title><link>https://macropaperwarehouse.com/papers/heterogeneity-in-manufacturing-growth-risk/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/heterogeneity-in-manufacturing-growth-risk/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; Since the Great Recession, quantifying downside risks to economic activity (rather than only expected outcomes) has become central for policymakers and investors. A large &amp;ldquo;growth-at-risk&amp;rdquo; literature documents that tightening financial conditions sharply raise downside risks to aggregate output while leaving upside potential roughly unchanged (Adrian, Boyarchenko and Giannone, 2019). This paper argues that the aggregate focus misses important structure: aggregate fluctuations can originate from industry-specific shocks, and recessions sharply raise cross-industry dispersion in growth (Bloom, 2014). The authors ask how downside output-growth risk from tight financial conditions differs across U.S. manufacturing industries, and which industry characteristics explain that heterogeneity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and method.&lt;/strong&gt; They use monthly industrial production (IP) growth for 74 U.S. manufacturing industries at the four-digit NAICS level over January 1973–July 2020 (Federal Reserve G.17; same industry selection as Chang and Hwang, 2015), and the Chicago Fed&amp;rsquo;s National Financial Conditions Index (NFCI) as the financial-conditions gauge. The method is a two-level (multi-level) quantile regression. Level 1 (following Adrian et al., 2019) regresses the τ-th quantile of average h-month-ahead IP growth on the current NFCI and current IP growth, industry by industry, focusing on h=3. Level 2 (inspired by Petersen and Strongin, 1996) regresses the estimated level-1 NFCI quantile coefficients cross-sectionally on standardized, time-invariant industry characteristics (capital, materials, energy, production-labor and overhead-labor intensities; a correlation-based labor-hoarding measure; four-firm concentration ratio; industry size measured by value-added share; and a durability dummy). Inference uses a stationary bootstrap (1,000 replications) that propagates level-1 estimation uncertainty into level 2. Industries split into 45 durables and 29 nondurables.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main quantitative findings.&lt;/strong&gt; Deteriorating financial conditions hit downside risk far harder than the center or upside of the growth distribution. On average across industries, a one-standard-deviation positive NFCI shock lowers three-month-ahead IP growth by 0.237% at the median and 0.773% at the 5% quantile, and raises the 95% quantile by 0.042%. The average 5% NFCI coefficient is -0.77 across all industries versus -0.31 (linear) and -0.24 (median); 47 of 74 industries (63.5%) have significant 5% coefficients, only 5 (6.8%) have significant 95% coefficients. Durables are about twice as sensitive in the left tail: average 5% coefficients are -0.96 (durables) versus -0.48 (nondurables), with 75.6% of durables versus 44.8% of nondurables significant at 5%. Some industries (computer, aerospace, food, dairy) are essentially unaffected across the whole distribution. The relationship is nonlinear for 46 of 74 industries (62.2%) at the 5% quantile (77.8% of durables, 37.9% of nondurables). Galvao et al. (2018) slope-homogeneity tests reject coefficient equality across industries for lower quantiles. Subsample analysis (1973-84 / 1985-2006 / 2007-2020) shows tail effects strongest in the most recent period (average 5% coefficient -1.38 vs -0.73 and -0.49), weakest during the Great Moderation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Explaining heterogeneity / implications.&lt;/strong&gt; In the all-manufacturing second level, large industries and durable-goods producers have significantly more vulnerable downside growth, while capital-intensive, overhead-labor-intensive, and labor-hoarding industries are less vulnerable. Within durables, size, materials intensity (more vulnerable) and overhead labor intensity (less vulnerable) matter; within nondurables, energy intensity (more vulnerable) and labor hoarding (less vulnerable) matter. Implication: industry-targeted stabilization policy may be more effective than nationwide policy given the heterogeneity, and investors can build industry-rotation strategies less exposed to financial-market shocks.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-empiricalidentification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the empirical/identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The strategy is descriptive-predictive rather than causal. Level 1 estimates industry-specific quantile regressions of average h-month-ahead IP growth on the current NFCI and current IP growth (Koenker-Bassett check-function minimization via the Frisch-Newton interior-point algorithm). Level 2 regresses the estimated NFCI quantile coefficients on standardized industry characteristics via OLS. The key inferential innovation is a stationary bootstrap (Politis-Romano 1994; block length via Politis-White 2004 with Patton et al. 2009 correction, expected block ~36.76 set by the NFCI series) that jointly resamples industry IP and NFCI and feeds level-1 estimation uncertainty into level-2 confidence bands. Main threats: (i) the relationship is associational, not identified as causal — the NFCI is endogenous to the macroeconomy; (ii) generated-regressor problem in level 2 (coefficients are estimates), addressed by the bootstrap; (iii) small cross-sections (45 durables, 29 nondurables, even fewer at the three-digit level) reduce power to detect characteristic effects; (iv) time-invariant characteristics are averaged over varying available windows, abstracting from time variation.&lt;/p&gt;
&lt;h3 id="q2-how-is-nonlinearity-established-and-against-what-benchmark"&gt;Q2. How is nonlinearity established, and against what benchmark?&lt;/h3&gt;
&lt;p&gt;Quantile coefficients are compared to OLS linear coefficients (constant across quantiles) using 95% bootstrap bands generated under a null that the data-generating process is a VAR(4) for the NFCI and IP growth (the Adrian et al. 2019 approach). Quantile estimates falling outside those bands are evidence of nonlinearity. 46 of 74 industries (62.2%) have a 5% coefficient significantly different from OLS; the total manufacturing sector is also nonlinear, mirroring Adrian et al. (2019) for aggregate GDP.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Three layers. (1) Durables vs nondurables: durables roughly twice as sensitive in the left tail (avg 5% coefficient -0.96 vs -0.48). (2) Within sectors: e.g. motor vehicles, motor bodies and motor parts have significant 5% coefficients below -2; resin and fiber below -1.5; while computer, aerospace and food are insignificant/unaffected. (3) Across the distribution: strong effects at low quantiles, near-zero at high quantiles (avg 95% coefficient 0.04). Industries with large negative 5% coefficients also tend to have larger positive 95% coefficients (higher conditional volatility under tight conditions), most clearly iron, motor vehicles, fiber and resin — though upside gains are generally smaller than the downside increase.&lt;/p&gt;
&lt;h3 id="q4-which-industry-characteristics-explain-the-heterogeneity-and-in-which-direction"&gt;Q4. Which industry characteristics explain the heterogeneity, and in which direction?&lt;/h3&gt;
&lt;p&gt;All-manufacturing (74 industries): negative effects on lower-quantile NFCI coefficients (i.e. more downside vulnerability) from industry size and durability; positive effects (less vulnerability) from overhead labor intensity, labor hoarding, and capital intensity. Durables: significant negative effect of materials intensity, negative (small) effect of size, positive effect of overhead labor intensity; production labor intensity significant at some higher quantiles. Nondurables: significant negative effect of energy intensity, positive effect of labor hoarding. Energy intensity, production labor intensity and concentration ratio are NOT significant for total manufacturing or durables in the way Petersen-Strongin found for cyclicality.&lt;/p&gt;
&lt;h3 id="q5-what-economic-mechanisms-are-offered-for-each-characteristic-effect"&gt;Q5. What economic mechanisms are offered for each characteristic effect?&lt;/h3&gt;
&lt;p&gt;Size: mean reversion — an industry larger than average is more likely to see growth fall (Braun-Larrain 2005). Durability: durable production is inherently more cyclical (Petersen-Strongin 1996). Labor hoarding / overhead labor: firms retain trained (especially nonproduction) workers due to sunk hiring/training costs (Becker 1962; Oi 1962; Parsons 1986), lowering the incentive to cut production in downturns. Capital intensity: higher fixed-to-variable cost ratio reduces incentive to cut output, and tangible capital provides collateral easing financing (consistent with Braun-Larrain 2005). Materials intensity (durables): higher share of variable costs raises cyclicality; also links to the negative materials-intensity/TFP relation of Baptist-Hepburn (2013).&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;(i) Additional controls (Gilchrist-Zakrajsek variables: term spread, real federal funds rate, credit spread, excess bond premium, plus extra IP lags) — qualitatively similar, wider bands. (ii) Unobserved heterogeneity via Ando-Bai (2020) interactive-fixed-effects panel quantile model (one common factor optimal) — highly similar. (iii) Alternative NAICS disaggregation: three-digit (21 industries; capital intensity dropped for multicollinearity; only labor hoarding and durability significant) and six-digit (101 industries; more characteristics significant, including production labor intensity and concentration ratio). (iv) Longer horizons h=6 and h=12 — qualitatively similar but weaker/less significant as horizon lengthens. (v) Subsample analysis of both the growth-risk coefficients and the characteristic construction windows (1973-84, 1985-2006, 2007-2020; and start dates 1958/1973/1987) — effects relatively stable; size and labor-hoarding effects weaken in recent periods while overhead labor and durability stay significant.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-relate-to-and-differ-from-petersen-and-strongin-1996-and-adrian-et-al-2019"&gt;Q7. How does this relate to and differ from Petersen and Strongin (1996) and Adrian et al. (2019)?&lt;/h3&gt;
&lt;p&gt;It extends Adrian et al. (2019) from aggregate to industry-level growth-at-risk, documenting substantial cross-industry variation that is invisible at the aggregate level — to the authors&amp;rsquo; knowledge the first disaggregate growth-at-risk study. It extends Petersen-Strongin (1996), who used a linear cyclicality framework, by allowing a flexible/nonlinear quantile relationship specifically with financial conditions. Findings broadly echo Petersen-Strongin for downside risk (materials intensity most important in durables; labor hoarding for nondurables — their only significant nondurable effect), but deviate by NOT finding energy intensity, production labor intensity, or concentration ratio significant in durables, and by adding size and capital intensity (cf. Braun-Larrain 2005) as relevant for total manufacturing. The agreement is attributed to business and financial cycles being closely intertwined (Claessens et al. 2012).&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Because vulnerability is highly heterogeneous, industry-level stabilization policy may be more effective than nationwide policy (OECD 2003), and policies can be targeted using the signalling characteristics (size, durability, materials/energy intensity vs capital/overhead-labor intensity and labor hoarding). Investors can build industry-rotation strategies less exposed to financial shocks. Scope conditions: evidence is U.S. manufacturing only, associational not causal, conditional on the NFCI as the financial-conditions measure, strongest at the three-month horizon and in the post-2007 subsample, and characteristic effects rest on relatively small cross-sections.&lt;/p&gt;
&lt;h3 id="q9-are-there-caveats-the-authors-themselves-flag"&gt;Q9. Are there caveats the authors themselves flag?&lt;/h3&gt;
&lt;p&gt;Yes: after splitting into durables/nondurables, fewer characteristic effects are significant, which the authors attribute to smaller cross-sections rather than absence of effects; the two-level model is estimated sequentially (two-step) not simultaneously; characteristics are treated as time-invariant averages (justified by stable cross-industry rankings, though production labor intensity shows a downward trend); and upside potential, while present, is generally smaller than the increased downside risk.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Growth-at-risk / downside growth risk&lt;/strong&gt;: The lower-quantile (e.g. 5%) of the conditional distribution of future output growth given current conditions; here the 5% quantile of average three-month-ahead industry IP growth conditional on the NFCI, capturing how bad growth could plausibly get under tight financial conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Multi-level quantile regression&lt;/strong&gt;: The authors&amp;rsquo; two-step procedure: level 1 estimates industry-specific quantile regressions of future IP growth on the NFCI and current IP growth; level 2 regresses the estimated NFCI quantile coefficients cross-sectionally on industry characteristics, with a bootstrap carrying level-1 uncertainty into level-2 inference.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;NFCI (National Financial Conditions Index)&lt;/strong&gt;: Chicago Fed weekly index of U.S. money, debt, equity, and (shadow) banking conditions built from a large dynamic factor model; positive values mean tighter-than-average financial conditions, negative values looser-than-average. Averaged to monthly here.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor hoarding&lt;/strong&gt;: Retention of employees during downturns because of sunk search, hiring and training costs; measured here as the negative correlation between changes in materials usage and changes in production-worker hours (a value of -1 = no hoarding), so higher values indicate more hoarding and predict less cyclical, less vulnerable growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Overhead labor intensity&lt;/strong&gt;: Cost of nonproduction (overhead) labor relative to value added. Because nonproduction workers embody more firm-specific investment, they are more subject to labor hoarding, so overhead-labor-intensive industries have less vulnerable downside growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Durable vs nondurable goods sector&lt;/strong&gt;: Federal Reserve classification (45 durable, 29 nondurable industries here). Durable-goods production is more cyclical and, in this paper, about twice as sensitive in the left tail of the growth distribution to adverse financial conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Slope homogeneity test&lt;/strong&gt;: Galvao et al. (2018) Swamy-type and standardized Swamy-type tests for a quantile-regression fixed-effects panel, used to formally reject equality of NFCI quantile slopes across industries, especially at lower quantiles.&lt;/p&gt;</description></item><item><title>Policy transition risk, carbon premiums, and asset prices</title><link>https://macropaperwarehouse.com/papers/policy-transition-risk-carbon-premiums-and-asset-prices/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/policy-transition-risk-carbon-premiums-and-asset-prices/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Central bankers, regulators, and investors increasingly worry about climate &amp;ldquo;transition risks&amp;rdquo; — abrupt shifts in climate policy, green technology breakthroughs, or consumer-preference shifts that re-price assets (Carney&amp;rsquo;s &amp;ldquo;tragedy of the horizon&amp;rdquo;). Rather than use the fixed NGFS-style stress-test scenarios, the authors ask how &lt;em&gt;policy transition risk&lt;/em&gt; — modeled as stochastic, reversible jumps between climate-policy regimes — endogenously affects carbon pricing, asset prices, risk premiums, the risk-free rate, and the speed of the green transition.&lt;/p&gt;
&lt;p&gt;Model setup: A global two-sector continuous-time DSGE macro-finance model of the climate and economy (building on Hambel, Kraft, van der Ploeg 2024). Two sectors produce perfectly substitutable final goods via Cobb-Douglas in capital and a CES energy composite of fossil fuel and renewables; sector 1 is &amp;ldquo;green&amp;rdquo; (renewables-intensive) and sector 2 is &amp;ldquo;brown&amp;rdquo; (fossil-intensive). Investment carries quadratic intertemporal adjustment costs and brown-to-green capital reallocation carries quadratic intrasectoral costs (a dollar of brown converts to less than a dollar of green). Temperature rises in cumulative emissions (TCRE specification). Households have Epstein-Zin recursive preferences; dividends are leveraged consumption (D=C^phi, phi&amp;gt;1). Capital is exposed to Brownian shocks plus Barro-style macro-disaster jumps; learning-by-doing lowers renewable costs. The core model has a two-state policy Markov chain — BAU (no carbon pricing) and CAP (carbon pricing internalizing damages and enforcing a Tcap=2C cap; if the cap is breached, fossil use is forced to zero). Policy tips with transition intensity calibrated at lambda_x = 4% per year from BAU to CAP. Model solved by finite differences; 20,000 simulated paths to 2100. Calibration: RRA gamma=2.977, EIS psi=1.5, time preference delta=0.0346, initial GDP $116tn, initial brown-capital share S0=0.876, TCRE=1.8 C/TtC, T0=1.27C.&lt;/p&gt;
&lt;p&gt;Main quantitative findings: (1) Under pure BAU, the green transition is slow and temperatures reach on average 3.9C above pre-industrial by 2100; risk-free rate and risk premiums are almost unaffected (TFP damage alone cannot generate a temperature premium). (2) With policy transition risk, by 2100 about 28% of paths stay below 1.8C, 46% land between 1.8C and 2.5C, and the rest exceed 2.5C; roughly 45% of paths adhere to the 2C cap; 94% of paths have active climate policy by 2100. On the illustrative path tipping to CAP in 2045, a carbon price of &lt;del&gt;$700/tC (&lt;/del&gt;$190/tCO2) is imposed; the green share price jumps +22% and the brown price drops -21.5% on impact. In the ~4% of paths where CAP is adopted in 2021, the carbon tax starts at ~$218/tC ($60/tCO2), about 50% larger than Pigouvian pricing without an enforced cap — because the cap forces policymakers to catch up. (3) The model generates a sizable, positive carbon premium (brown minus green risk premium) that is initially near zero but becomes large when temperature is close to or above the 2C cap and the economy is still carbon-intensive; the dominant channel is the asymmetric temperature-shock impact on the brown sector&amp;rsquo;s price-dividend ratio (third term of eq. 3.4). Without transition risk (first-best Pigouvian pricing), the carbon premium is slightly negative. (4) The mean risk-free rate starts at 0.8% and is largely stable, but its lower quantile falls sharply when temperature approaches/exceeds the cap as precautionary saving rises. (5) Extensions table: in the pure PIGOU scenario (no cap, no transition risk) climate disasters roughly double the optimal carbon tax from $45/tCO2 (2025) to $91, and adding irreversible climate tipping raises it to $121; in the core BAU-&amp;gt;CAP model the average optimal CO2 tax rises from $73 to $108 (disasters) to $134 (tipping). News effects on share prices are far larger for policy tips than for climate or technology tips (climate tipping events move prices ~3-5%; a BAU-&amp;gt;CAP tip moves the brown price ~-27% and brown price-dividend ratio ~-13%, green price +18%, green PDR +42%).&lt;/p&gt;
&lt;p&gt;Implications: Policy transition risk makes average policy more ambitious than BAU but less than first-best; it produces risk-driven carbon premiums that accelerate the green transition, raises precautionary saving, and depresses the risk-free rate near the cap. Physical risks alone (assumed symmetric across sectors) cannot generate a sizable carbon premium but do raise carbon prices and create a temperature risk premium on all assets.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;This is a calibrated structural (DSGE) model, not an empirical identification design, so &amp;lsquo;identification&amp;rsquo; here means the model mechanism that generates carbon premiums plus calibration to external sources. The carbon premium is generated purely endogenously by making the brown sector more fossil-/carbon-intensive than the green sector, with physical risks assumed to load symmetrically on both capital stocks so any premium asymmetry comes from policy transition risk and temperature exposure rather than from differential physical-risk loadings. The main threats the authors acknowledge are: (i) calibration choices for negative-emissions cost curves and transition probabilities are &amp;rsquo;tentative&amp;rsquo; and partly curve-fit/ad hoc; (ii) exogenous and stark policy states (two or three regimes with given/partly exogenous transition intensities) are a simplified representation of the political process; (iii) global-economy calibration sits uneasily with national-election interpretations of policy tipping. They argue the forward-looking households/firms make the model robust to the Lucas critique.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished"&gt;Q2. What are the main mechanisms, and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;Three channels for the carbon premium appear in equation (3.4): (1) a stochastic-discount-factor/transition-shock term scaling in transition intensities lambda_x; (2) a diffusive term from the volatility of the brown-capital share affecting the brown price-dividend ratio more (largest when S(1-S) is high, i.e. share neither very high nor very low) and from higher consumption-capital-ratio volatility in the brown sector combined with leverage; (3) a temperature-shock term that becomes large near 2C because the policy transition to CAP becomes potentially devastating (forced phase-out of fossil fuel) and hits the brown PDR much more than the green PDR. The authors state the third (temperature-near-cap) effect is quantitatively the most important. The premium is risk-driven, distinguished from preference-driven mechanisms (Pastor et al. 2021; Pedersen et al. 2021; Zerbib 2022) in which green investors accept lower returns.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Heterogeneity is across states and paths rather than across firms in data. The carbon premium and risk-free-rate response depend nonlinearly on temperature (large near/above 2C) and on the brown-capital share S (large transition effect when S is high). Across simulated paths the outcomes diverge widely: ~28% below 1.8C, ~46% between 1.8C and 2.5C, the rest above 2.5C by 2100. The price impact of news differs sharply by type: policy tips dominate climate tips and technology tips. The risk-free rate&amp;rsquo;s lower quantile falls much more in high-temperature paths.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-and-extensions-are-run"&gt;Q4. What robustness checks and extensions are run?&lt;/h3&gt;
&lt;p&gt;Extensions: (a) recurring temperature-dependent climate disasters (intensity rising linearly in T, lambda_c-hat=0.096, lambda_c(T0)=0.122, expected loss 1.5% vs 25% for macro disasters, alpha_c=65.7); (b) irreversible climate tipping via a 3-state chain raising TCRE from 1.8 to 2.1 to 2.4 C/TtC and adding permanent damages d=0,0.025,0.05; (c) a negative-emissions/technology-breakthrough state (2-state chain, ~50% chance of competitive technology by 2050, intensity 0.0224, cost curve fit to Rebonato et al. 2023); (d) a richer 3-state policy chain BAU/PIGOU/CAP with reversible and partly endogenous transition probabilities (switch to active policy rising toward 75% if T&amp;gt;1.5C; lobbying makes switches depend on brown/green capital shares), giving an 18-state (2x3x3) Markov chain. Core qualitative results (positive carbon premium driven by policy risk near the cap, precautionary saving lowering the risk-free rate) survive all extensions; the carbon premium is smaller in the 3-state model because only ~30% of paths reach CAP. A model variant with exhaustible fossil resources (cap 3000 GtC) found the exhaustibility constraint non-binding.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It extends Hambel et al. (2024), which used a two-sector economy for climate disasters/tipping and first-best carbon prices but did not study policy transition risk or carbon premiums. It builds general-equilibrium structure on the partial-equilibrium reduced-form insights of Hsu et al. (2023) on the pollution premium (who report a 4.42% annual pollution premium). It is most closely related to Barnett (2024), also a DSGE transition-risk model, but adds richer interactions among climate tipping, political risk, and technology breakthrough, imperfect energy substitution, and intrasectoral adjustment costs; Barnett instead emphasizes a climate-policy-driven &amp;lsquo;run on fossil fuel&amp;rsquo;. It provides a risk-based mechanism for the carbon premium documented empirically by Bolton and Kacperczyk (2021, 2023) and Hsu et al. (2023), while noting contrary evidence (Pastor et al. 2021; Bauer et al. 2022; Aswani et al. 2024; Zhang 2025 — who finds the premium turns negative in the U.S. after a data-lag correction; Hambel and van der Sanden 2024). Calibration of policy scenarios follows Moore et al. (2022).&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Under policy transition risk, average climate policy is more ambitious than BAU but less ambitious than first-best; policymakers may set carbon taxes even higher than first-best to &amp;lsquo;catch up&amp;rsquo; for time lost by predecessors when the economy is close to the temperature cap. Carbon premiums encourage firms to shift investment from brown to green and accelerate the transition. Scope conditions: carbon premiums are large only when the economy is still carbon-intensive (high brown-capital share) AND temperature is near or above the 2C cap; if policymakers implement first-best Pigouvian taxes while ignoring transition risk, the carbon premium is slightly negative. Physical-risk symmetry across sectors is assumed; if physical risk hit sectors differently there would be additional carbon-premium effects.&lt;/p&gt;
&lt;h3 id="q7-what-happens-to-asset-prices-at-the-moment-of-each-type-of-tipping"&gt;Q7. What happens to asset prices at the moment of each type of tipping?&lt;/h3&gt;
&lt;p&gt;At a tip to more ambitious carbon pricing, green share prices rise and brown share prices fall (and conversely when policy weakens). At a climate tip, both green and brown share prices fall (~3-5% each in the illustrative path). When negative-emissions technology becomes available, green prices jump down and brown prices jump up while the carbon price falls (because the brown sector may use fossil fuel again). The brown asset becomes worthless once the transition completes and the brown capital stock is run down; partial stranding occurs when the cap is crossed and fossil use is banned. News effects on prices are much larger for policy than for climate or technology tipping.&lt;/p&gt;
&lt;h3 id="q8-what-drives-the-risk-free-rate-dynamics"&gt;Q8. What drives the risk-free rate dynamics?&lt;/h3&gt;
&lt;p&gt;The risk-free rate (eq. 3.2) combines discounting, consumption-smoothing, standard diffusion and macro-disaster precautionary saving, an uninsurable temperature-risk term (small because consumption volatility is close to capital volatility, and it vanishes under CRRA), and a novel policy-transition-risk term that makes the rate jump with the policy state. Increased transition risk raises precautionary saving and lowers the rate, especially when temperature is close to its cap (where forced fossil phase-out makes expected consumption growth drop). As the transition completes and brown capital shrinks, precautionary saving falls and the rate stabilizes. Mean rate ~0.8%, stable; lower quantile falls over time.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Policy transition risk&lt;/strong&gt;: In this paper, the risk arising from stochastic, reversible jumps between discrete climate-policy regimes (no / modest / ambitious carbon pricing), modeled as a Markov chain with given or partly endogenous transition intensities — distinct from fixed NGFS-style scenarios. Financial markets price these regime-change risks even in the BAU state.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Carbon premium&lt;/strong&gt;: Defined as the difference between the brown and green risk premiums (r^p_2 minus r^p_1). In the model it is a purely risk-driven, endogenous object arising because policy/temperature shocks hit the carbon-intensive brown sector&amp;rsquo;s price-dividend ratio more than the green sector&amp;rsquo;s; it is large near the temperature cap and slightly negative under first-best pricing without transition risk.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;CAP policy state&lt;/strong&gt;: The &amp;lsquo;ambitious carbon pricing&amp;rsquo; regime in which policymakers set the carbon tax to internalize warming damages AND enforce a hard temperature cap Tcap=2C; if the cap is breached, fossil-fuel use is forced to zero (F1=F2=0) and carbon prices exceed the usual social cost of carbon.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PIGOU policy state&lt;/strong&gt;: The &amp;lsquo;modest carbon pricing&amp;rsquo; regime (added in the extended 3-state chain) that internalizes all global-warming externalities, including risks of climate disasters and tipping, but does NOT impose a temperature cap — yielding lower carbon taxes than CAP.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;TCRE (transient climate response to cumulative emissions)&lt;/strong&gt;: The proportionality coefficient (theta/vartheta) translating cumulative net emissions into temperature change; calibrated at 1.8 C/TtC in the core model and allowed to jump irreversibly to 2.1 and 2.4 C/TtC under climate tipping.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Temperature/transition risk premium&lt;/strong&gt;: A positive risk premium carried by all risky assets stemming from physical climate risk (disasters and tipping) that rises with the level of temperature; distinct from the carbon premium, which is the brown-minus-green differential and is driven mainly by asymmetric policy-transition exposure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Partial asset stranding&lt;/strong&gt;: The situation when the temperature cap is crossed and fossil fuel may no longer be burned, so the brown sector — though still operable with renewables — loses the value of its fossil-based capital, causing the brown share price to fall.&lt;/p&gt;</description></item><item><title>Does a Financial Crisis Impair Corporate Innovation?</title><link>https://macropaperwarehouse.com/papers/does-a-financial-crisis-impair-corporate-innovation/</link><pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/does-a-financial-crisis-impair-corporate-innovation/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Why do financial crises leave such deep and protracted economic wounds, with crisis-stricken economies failing to revert to pre-crisis growth trends even a decade later? Imai and Sawada test one specific channel: that crisis-induced disruptions in financial intermediation impair firms&amp;rsquo; ability to fund innovation projects, stalling technological progress and thereby pushing the economy onto a permanently lower growth path. They study this in the context of Japan&amp;rsquo;s 1997-1998 financial crisis, which featured a sharp decline in bank credit, the collapse of three major banks (Hokkaido Takushoku Bank, Long-Term Credit Bank, Nippon Credit Bank), and a failure to recover the pre-crisis growth trend. Laeven and Valencia (2020) estimate the crisis&amp;rsquo;s fiscal cost to Japanese taxpayers at 8.5% of GDP and its economic cost (GDP deviation from trend, 1997-2001) at 45% of GDP.&lt;/p&gt;
&lt;p&gt;Data and strategy: The authors link three firm-level longitudinal datasets. Innovation output is measured from the Institute of Intellectual Property (IIP) Patent Database (Japan Patent Office data): patent applications, granted patents (only ~30% of Japanese applications are granted, taking 7-8 years), and citation-weighted patents using forward citations accumulated in a 17-year window after application. The core sample period is 1994-2003 (a 10-year window around the crisis), with forward citations tracked up to 2018; this long post-crisis window is a deliberate design choice that lets truncation-prone citation data mature. Bank dependence is proxied by the ratio of total loans to total assets (drawn from Nikkei Financial Quest financial statements). Bank-failure exposure is identified from the Corporate Borrowings Database: firms borrowing more than 10% of total bank loans from a failed bank in the year before its failure are coded as client firms. Patent applicants are matched to financial data via NISTEP company-name identification codes, covering roughly 75% of patents by NISTEP-ID firms and 58% of all applications.&lt;/p&gt;
&lt;p&gt;Two empirical designs: (1) A DiD interacting the loan-to-assets ratio with a Crisis dummy (=1 for 1997-2001), with firm, industry-year, and prefecture-year fixed effects, firm controls (log sales, log age, ROA, cash-to-assets, tangible-to-assets) lagged one year and also interacted with the crisis dummy. (2) A bank-failure DiD adding a Bank Failure dummy (=1 for HTB clients 1997-2001, LTCB/NCB clients 1998-2001).&lt;/p&gt;
&lt;p&gt;Main findings with magnitudes: Bank-dependent firms cut both the quantity and quality of innovation more sharply and persistently after the crisis; the loan-ratio-x-crisis interaction is negative and significant for applications, grants, and citations, and robust to the fully saturated fixed-effects model. In the event-study, high bank-dependence (top quartile) firms gained roughly 50% fewer patents over 1997-2003 relative to low-dependence firms (marginally significant), with no pre-trend in 1994-1995. The effect is concentrated in small and medium firms (insignificant for large firms). Decomposing loan maturity, the short-term-loans-x-crisis interaction is negative and robustly significant while the long-term-loans interaction is not, pointing to rollover risk as the main mechanism. For bank failures, the average effect across all firms is small and insignificant, but for small firms it is negative and significant: bank failures are associated with declines of about 12% in granted patents and 17% in cited-weighted patents; the dynamic counterfactual implies small firms whose main bank failed would have been granted about 50% more patents absent the failure, with effects peaking ~2 years after failure and recovering to pre-failure levels within about 4 years.&lt;/p&gt;
&lt;p&gt;Implications: Post-crisis innovation performance depends on the degree to which firms rely on monitored, difficult-to-replace relationship lending. The crisis-induced decline in innovation among opaque, bank-dependent firms is offered as a plausible explanation for Japan&amp;rsquo;s long-term post-1990s productivity and growth stagnation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-two-identification-strategies-and-what-is-the-key-identifying-assumption"&gt;Q1. What are the two identification strategies, and what is the key identifying assumption?&lt;/h3&gt;
&lt;p&gt;First, a difference-in-differences design interacting a continuous bank-dependence proxy (loan-to-assets ratio) with a Crisis dummy (=1 for 1997-2001), identifying off differential responses of more- vs. less-bank-dependent firms. Second, a bank-failure DiD interacting a Bank Failure dummy (for clients borrowing &amp;gt;10% of bank loans from HTB/LTCB/NCB before failure) with the crisis period. The key identifying assumption is parallel trends: clients of failed banks and clients of surviving banks would have followed the same innovation path absent the failures. The authors support this with event-study coefficients showing no significant pre-trends (1994-1995 for bank dependence; 3-4 and 2 years before failure for bank failures).&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-threats-to-identification-and-how-are-they-addressed"&gt;Q2. What are the main threats to identification and how are they addressed?&lt;/h3&gt;
&lt;p&gt;(1) Bank-dependent firms might be concentrated in declining or cyclically sensitive industries or worse regions — addressed by adding industry-year and prefecture-year fixed effects, so estimates come from firms in the same industry and prefecture; results are insensitive. (2) The decline might reflect poor financial performance or other firm correlates — addressed by interacting the crisis dummy with firm-level controls (size, age, ROA, tangible-to-assets, cash-to-assets); results hold. (3) Exposure to the late-1990s East Asian crisis via exports — addressed by interacting an overseas-sales-to-total-sales ratio with the crisis dummy (losing over half the sample); results robust (Table A2). (4) &amp;lsquo;Cleansing&amp;rsquo;/zombie-lending selection (failed banks served unviable firms) — addressed by dropping non-innovative firms and restricting to manufacturing (least affected by zombie lending); effects persist. (5) Omitted-variable bias for bank failure — assessed via coefficient-stability arguments (Altonji et al. 2005, Oster 2019); estimates stable to inclusion/exclusion of controls.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-main-mechanism-and-how-is-it-distinguished-empirically"&gt;Q3. What is the main mechanism and how is it distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The bank lending channel: crises raise the cost of intermediated funds, disproportionately hurting firms reliant on bank finance. The authors further pin down rollover risk by decomposing loans into short-term (residual maturity &amp;lt;=1 year) and long-term relative to assets and interacting each with the crisis. The short-term-loan interaction is negative and robustly significant; the long-term-loan interaction is negative but not robustly significant and becomes insignificant when both are included. This indicates the impairment operates mainly through firms&amp;rsquo; exposure to short-term rollover risk rather than long-term debt levels.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented"&gt;Q4. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Effects are concentrated in small and medium-sized firms (terciles by 1996 sales). For large firms the bank-dependence-x-crisis interaction is insignificant. Bank-failure effects are insignificant on average but negative and significant for small firms (about -12% granted patents, -17% cited-weighted patents), and small/insignificant for medium and large firms. The interpretation is that smaller, opaque firms face more severe asymmetric-information problems and find it hardest to replace an informed relationship lender when their main bank fails.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Progressive fixed effects (firm+year; +industry-year; +prefecture-year); crisis-dummy interactions with firm controls; dropping non-innovative firms (never applied/granted patents); restricting to manufacturing (least zombie-affected); R&amp;amp;D-intensity-based industry exclusions; an alternative small-firm definition (first quartile vs first tercile — application results similar, citation results weaken since these firms&amp;rsquo; patents are rarely cited); using R&amp;amp;D expenditure (Toyo Keizai self-reported) as an alternative outcome (bank-dependent firms cut R&amp;amp;D more, Table A1); interacting overseas-sales ratio with crisis (Table A2); separating loans from other debts (loans interaction more robust than other-debt interaction, Table A3); and an industry-linear-trend specification (qualitatively unchanged, unreported).&lt;/p&gt;
&lt;h3 id="q6-did-the-financial-health-of-the-main-bank-matter-beyond-the-binary-failure-event"&gt;Q6. Did the financial health of the main bank matter, beyond the binary failure event?&lt;/h3&gt;
&lt;p&gt;No robustly. Using percentage change in main banks&amp;rsquo; share prices from 1993-1998 (interacted with the crisis dummy) to proxy bank weakness, the authors find no robust evidence that clients of weaker-but-surviving banks innovated differently. They conclude differences in main-bank financial health are second-order relative to firm-level heterogeneity in bank dependence (Table A4).&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It builds on Japanese bank-health-to-real-activity studies (Peek and Rosengren, Gibson, Amiti-Weinstein, etc.) but tracks much longer-horizon, persistent effects on innovation rather than short-term investment/employment. Relative to Nanda and Nicholas (2014, Great Depression patenting), it uses linked bank-firm data with industry-year and region-year fixed effects to control for demand shocks, and argues 1990s Japan (scarcer breakthrough opportunities) may be more relevant to contemporary settings than the technologically fertile 1930s US. Unlike Hardy and Sever (2021), which uses only US-office patents granted to foreign firms (selection concerns) at industry level, this paper uses all domestically granted Japanese patents at the firm level. It follows Duval, Hong, and Timmer (2020) on balance-sheet heterogeneity and Huber (2018) on bank failures, but adds invention-quality measurement via long forward-citation windows that the 2008-crisis literature cannot yet exploit. It complements Hombert and Matray (2017) on relationship lending and small-firm innovation.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-dynamics-of-the-bank-failure-effect-on-small-firms"&gt;Q8. What are the dynamics of the bank-failure effect on small firms?&lt;/h3&gt;
&lt;p&gt;In the event study, pre-failure coefficients (3-4 and 2 years before) are small and insignificant. Post-failure coefficients are largely negative, with the largest, significant declines about 2 years after failure (consistent with lags in producing innovation). Innovation performance recovers to pre-failure levels within about 4 years, but cumulative losses are large — implying small firms would have received roughly 50% more patents absent the failure. Effects are qualitatively similar excluding non-innovative firms or non-manufacturing firms.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policytheoretical-implications-and-their-scope-conditions"&gt;Q9. What are the policy/theoretical implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The adverse real effects of a systemic banking crisis can linger because opaque, bank-dependent firms&amp;rsquo; innovation declines persistently, plausibly contributing to Japan&amp;rsquo;s long-run post-crisis productivity and growth stagnation. Scope conditions: the effect is specific to small, opaque, bank-dependent firms reliant on relationship and especially short-term bank finance; it does not generalize to large firms; the mechanism is loss of monitored, difficult-to-replace relationship lending plus rollover risk, not generic financial weakness or main-bank fragility; and the setting (heavily bank-centered Japanese financial system, scarce breakthrough opportunities) shapes external validity.&lt;/p&gt;
&lt;h3 id="q10-what-are-notable-caveats-and-data-limitations"&gt;Q10. What are notable caveats and data limitations?&lt;/h3&gt;
&lt;p&gt;Bank dependence is proxied by total loans (including loans from non-financial parents/affiliates) over assets rather than pure bank borrowings, because the cleaner Corporate Borrowings Database omits pre-1996 OTC firms; the authors verify total loans only slightly exceed bank borrowings and results hold on the cleaner sub-sample. Patent-financial matching covers ~58% of all applications. Cumulative bank-dependence effects (~50%) are only marginally significant. R&amp;amp;D-based outcomes are hampered by a 2000 Japanese accounting-standard change and inconsistent firm reporting. Citation data are truncated, motivating the long 17-year (and 15-year for 1994-2003) windows.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item></channel></rss>